Cài llama.cpp
llama.cpp
├── compose.yaml
└── models
└── sweep-next-edit-1.5b.q8_0.v2.gguf
Tải edit prediction model từ địa chỉ: https://huggingface.co/sweepai/sweep-next-edit-1.5B.
Tạo file compose.yaml với nội dung sau:
-
Nếu bạn dùng card Nvidia:
services: fim: # Change image to llama.cpp:server-cuda if your card does not support cuda13 image: ghcr.io/ggml-org/llama.cpp:server-cuda13 ipc: host deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu] ports: - "8080:8080" volumes: - ./models:/models tty: true command: - -m - /models/qwen2.5-coder-7b-instruct-q8_0.gguf - --port - "8080" -
Nếu bạn dùng card AMD:
services: fim: image: ghcr.io/ggml-org/llama.cpp:server-rocm ipc: host ports: - "8080:8080" volumes: - ./models:/models tty: true devices: - /dev/kfd:/dev/kfd - /dev/dri:/dev/dri group_add: - video security_opt: - seccomp:unconfined cap_add: - SYS_PTRACE environment: - AMD_VISIBLE_DEVICES=all - HSA_OVERRIDE_GFX_VERSION=10.3.0 # For RDNA2 GPU like 6700XT command: - -m - /models/sweep-next-edit-1.5b.q8_0.v2.gguf - --port - "8080"
Context có thể được cấu hình qua tùy chọn --ctx-size trong phần command, ví dụ:
command:
- -m
- /models/sweep-next-edit-1.5b.q8_0.v2.gguf
- --port
- "8080"
- --host
- 0.0.0.0
- --ctx-size
- "147000"
Chạy container:
docker compose up -dCấu hình Zed
Trong file settings.json của Zed, chỉ định edit_predictions sang local LLM server bên trên:
{
"show_edit_predictions": true,
"edit_predictions": {
"open_ai_compatible_api": {
"prompt_format": "qwen",
"model": "sweep-next-edit-1.5b.q8_0.v2.gguf",
"api_url": "http://localhost:8080/v1/completions"
},
"allow_data_collection": "no",
"mode": "eager",
"provider": "open_ai_compatible_api"
},
}
Xong.
Bình luận