8ba4db150f2ba8dbedd8eb64e243072079033f41
Wires the dual-plane ternary type into the Vulkan backend through three paths only: - get_rows and the generic scalar mul_mat_vec use per-element decode in dequant_funcs.glsl: the byte and base-3 digit are located from the element index (regions qs[0..16), qs[16..24), qh[0..2)), the byte is multiplied by 3^n mod 256 and the top digit taken. w = d1*t1 + d2*t2 is accumulated in fp32; both products are exact so the sum carries a single float rounding and reproduces the CPU reference bit by bit (verified: 0 mismatches over the 112-block synthetic test and over 214,695,936 elements of real model tensors). - larger matmuls fall back to dequant_dt3.comp (decode each byte once with q <- q*3 mod 256, fp32 sum, one rounding at the f16 write) plus the existing f16 matmul pipelines. The qh bytes hold only 4 trits; their 5th base-3 digit is packing padding that decodes to -1, so both decoders stop at 4 digits. Deliberately NOT implemented, and declined instead of half-supported: no coopmat/coopmat2/MMQ shaders are generated for DT3, and supports_op answers false for MUL_MAT_ID (mul_mat_vec_id shaders are not generated either). GET_ROWS and MUL_MAT answer true. test-dt3-gpu accepts the Vulkan backend (and IGPU-type devices) and passes on RADV STRIX1: dequant 0 mismatches, mul_mat n<=8 norm rel err ~1e-8, GEMM fallback bit-identical to an F16 GEMM on fp16-rounded weights.
tool-call: fix Qwen 2.5 Coder support, add micro benchmarks, support trigger patterns for lazy grammars (#12034)
llama.cpp
LLM inference in C/C++
manifesto / ggml / ops / maintainer PRs / compile times / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - stb-image - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- miniaudio.h - Single-header audio format decoder, used by multimodal subsystem - Public domain
- subprocess.h - Single-header process launching solution for C and C++ - Public domain
Languages
C++
55%
C
16.2%
Python
7.1%
Cuda
5.6%
TypeScript
4.4%
Other
11.5%