Note
This is adderek/llama.cpp, a fork of ggml-org/llama.cpp tuned for AMD Radeon RX 7900 XTX (gfx1100, RDNA3) on ROCm 7.2.4 / Linux, and tested only there. Upstream is merged in regularly. The most important changes:
KV cache compression
- TurboQuant KV cache (
--cache-type-k/-v turbo2|turbo3|turbo4): 2.1-4.25 bits per value, 3.8x smaller than f16 at turbo4. Works on dense, GQA, sparse (QSA, Qwen3.8-Flash-Next), MLA and DSA attention; K and V may use different types. - Correctness fixes on top of it: the QSA path read the cache in the wrong basis and silently answered from the wrong part of the context; GQA models with padded heads crashed.
Speed on RDNA3
- FlashAttention for turbo caches: 2.6-5.4x faster prefill at 16-64k context; decode up to 34% faster, on Ornith-35B faster than an f16 cache at long context.
- MoE models larger than VRAM: offloaded expert matmuls go to the GPU you choose
(
GGML_CUDA_OP_OFFLOAD_DEVICES), 2x prefill here. - Models larger than VRAM + RAM on branch
moe-tier(not yet inmaster): hot experts on the GPU, cold ones streamed from NVMe with O_DIRECT, 2.4 instead of 0.9 tokens/s on a 244 GB 397B model.
Server and runtime
--reasoning autodoes not enable thinking (upstream does); use--reasoning on.- Speculative decoding parameters can be set per request; logs show wall-clock time and
busy/total slots; a stall watchdog dumps backtraces (
LLAMA_STALL_WATCHDOG_SECS). - ROCm robustness: fixed a GPU fault on prompt-cache restore, a turbo decode hang and an abort under CUDA-graph capture.
--hugepagesfor model weights; K2-Horizon model support.
Regression tests for every fork feature run in ctest (test-turbo-kv-*, test-fattn-turbo4, ...).
Details, measurements, environment variables and which GPUs it can run on (no turbo on Metal / Vulkan / SYCL; CUDA never built): FORK.md. Everything below is the upstream README, unchanged.
LLM inference in C/C++
ggml / ops / maintainer PRs / dev stats / lib llama API / llama-server REST API
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with llama cli
|
Built-in web UI against llama serve
|
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - nothings/stb - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- mackron/miniaudio - Single-header audio format decoder, used by multimodal subsystem - Public domain
- sheredom/subprocess.h - Single-header process launching solution for C and C++ - Public domain

