Setup llama.cpp
llama.cpp
├── compose.yaml
└── models
└── sweep-next-edit-1.5b.q8_0.v2.gguf
The model can be downloaded from: https://huggingface.co/sweepai/sweep-next-edit-1.5B.
Content of compose.yaml:
-
For Nvidia with CUDA:
services: fim: # Change image to llama.cpp:server-cuda if your card does not support cuda13 image: ghcr.io/ggml-org/llama.cpp:server-cuda13 ipc: host deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu] ports: - "8080:8080" volumes: - ./models:/models tty: true command: - -m - /models/qwen2.5-coder-7b-instruct-q8_0.gguf - --port - "8080" -
For AMD with rocm:
services: fim: image: ghcr.io/ggml-org/llama.cpp:server-rocm ipc: host ports: - "8080:8080" volumes: - ./models:/models tty: true devices: - /dev/kfd:/dev/kfd - /dev/dri:/dev/dri group_add: - video security_opt: - seccomp:unconfined cap_add: - SYS_PTRACE environment: - AMD_VISIBLE_DEVICES=all - HSA_OVERRIDE_GFX_VERSION=10.3.0 # For RDNA2 GPU like 6700XT command: - -m - /models/sweep-next-edit-1.5b.q8_0.v2.gguf - --port - "8080"
You can increase context size by adding option --ctx-size to command field, for example:
command:
- -m
- /models/sweep-next-edit-1.5b.q8_0.v2.gguf
- --port
- "8080"
- --host
- 0.0.0.0
- --ctx-size
- "147000"
Start the container:
docker compose up -dConfigure Zed
In your Zed's settings.json, configure edit_predictions to your local LLM server which just setup above:
{
"show_edit_predictions": true,
"edit_predictions": {
"open_ai_compatible_api": {
"prompt_format": "qwen",
"model": "sweep-next-edit-1.5b.q8_0.v2.gguf",
"api_url": "http://localhost:8080/v1/completions"
},
"allow_data_collection": "no",
"mode": "eager",
"provider": "open_ai_compatible_api"
},
}
Done.
Comments