- Add nvidia-container-toolkit repo:
curl -s -L https://nvidia.github.io/libnvidia-container/stable/rpm/nvidia-container-toolkit.repo | \ sudo tee /etc/yum.repos.d/nvidia-container-toolkit.repo
-
Install and configure
nvidia-container-toolkit:sudo dnf install nvidia-container-toolkit sudo nvidia-ctk runtime configure --runtime=containerd sudo nvidia-ctk runtime configure --runtime=docker sudo systemctl restart containerd sudo systemctl restart docker -
Create docker-compose file for vLLM with content:
services: vllm-openai: image: vllm/vllm-openai:latest container_name: vllm-openai restart: unless-stopped ports: - "8000:8000" volumes: - ./cache:/root/.cache/huggingface environment: - HF_TOKEN=your_huggingface_token - VLLM_API_KEY=sk-nothing ipc: host deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu] command: > --model Qwen/Qwen3-0.6B --enable-auto-tool-choice --tool-call-parser hermes
The docker command above will start vLLM with Qwen3-0.6B model, you can change it to other models.
- Start vLLM:
docker compose up
Reference
- https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html#with-dnf-rhel-centos-fedora-amazon-linux
- https://stackoverflow.com/questions/52865988/nvidia-docker-unknown-runtime-specified-nvidia
Comments