My Quick Notes

Ain't Nobody Got Time to Read

March 12, 2026

Run vLLM with Docker on Fedora

  • Add nvidia-container-toolkit repo:
    curl -s -L https://nvidia.github.io/libnvidia-container/stable/rpm/nvidia-container-toolkit.repo | \
        sudo tee /etc/yum.repos.d/nvidia-container-toolkit.repo
  • Install and configure nvidia-container-toolkit:

    sudo dnf install nvidia-container-toolkit
    
    sudo nvidia-ctk runtime configure --runtime=containerd
    sudo nvidia-ctk runtime configure --runtime=docker
    sudo systemctl restart containerd
    sudo systemctl restart docker
  • Create docker-compose file for vLLM with content:

    services:
      vllm-openai:
        image: vllm/vllm-openai:latest
        container_name: vllm-openai
        restart: unless-stopped
        ports:
          - "8000:8000"
        volumes:
          - ./cache:/root/.cache/huggingface
        environment:
          - HF_TOKEN=your_huggingface_token
          - VLLM_API_KEY=sk-nothing
        ipc: host
        deploy:
          resources:
            reservations:
              devices:
                - driver: nvidia
                  count: all
                  capabilities: [gpu]
        command: >
          --model Qwen/Qwen3-0.6B
          --enable-auto-tool-choice
          --tool-call-parser hermes

The docker command above will start vLLM with Qwen3-0.6B model, you can change it to other models.

  • Start vLLM:
    docker compose up

Reference

  • https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html#with-dnf-rhel-centos-fedora-amazon-linux
  • https://stackoverflow.com/questions/52865988/nvidia-docker-unknown-runtime-specified-nvidia
PreviousUpgrade Fedora 43 to Fedora 44
NextHyprland - Pixelated your screen when locking

Comments