a52077c4cabb4f3c0298329c9d2dd1324d5604cb
CI (cpu) / ubuntu (x64, ubuntu-22.04) (push) Failing after 30s
CI (CUDA, ubuntu) / hip (push) Failing after 55s
CI (android) / ndk (push) Failing after 1m27s
CI (android) / arm64 (push) Failing after 1m33s
CI (android) / default (push) Failing after 2m39s
CI (CUDA, ubuntu) / cuda (push) Failing after 1m14s
CI (sanitize) / ctest (ubuntu-24.04, THREAD) (push) Failing after 8s
CI (CUDA, ubuntu) / musa (push) Failing after 5m44s
CI (sycl) / ubuntu-24-sycl (fp16, ON) (push) Failing after 3m44s
CI (sycl) / ubuntu-24-sycl (fp32, OFF) (push) Failing after 2m56s
CI (vulkan) / ubuntu-llvmpipe (push) Failing after 1m42s
CI (webgpu) / format (push) Successful in 19s
CI (webgpu) / ubuntu (push) Failing after 8s
CI (cpu) / windows (x64, x64-openblas, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_OPENMP=OFF -DGGML_BLAS=ON -DGGML_BLA… (push) Canceled after 0s
CI (cpu) / windows (x64, x64-vulkan, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_VULKAN=ON) (push) Canceled after 0s
CI (3rd-party) / ubuntu-24-llguidance (push) Canceled after 0s
CI (apple) / macos-latest-arm64 (push) Canceled after 0s
CI (apple) / macos-latest-x64 (push) Canceled after 0s
CI (apple) / macos-latest-ios-xcode (push) Canceled after 0s
CI (apple) / macos-latest-tvos (push) Canceled after 0s
CI (apple) / macos-latest-visionos (push) Canceled after 0s
Build relocatable cmake package / linux (push) Canceled after 0s
CI (sanitize) / ctest ([self-hosted X64 Linux], UNDEFINED) (push) Canceled after 0s
CI (cpu) / ubuntu (arm64, ubuntu-24.04-arm) (push) Canceled after 0s
CI (cpu) / windows (arm64, arm64, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON) (push) Canceled after 0s
CI (cpu) / windows (x64, x64-cpu-static, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DBUILD_SHARED_LIBS=OFF) (push) Canceled after 0s
CI (ibm) / ubuntu-24-s390x (push) Canceled after 0s
CI (ibm) / ubuntu-24-ppc64le (push) Canceled after 0s
CI (opencl) / windows-2025-opencl-adreno (push) Canceled after 0s
CI (openvino) / ubuntu-24-openvino (push) Canceled after 0s
CI (openvino) / openvino-windows-2022 (push) Canceled after 0s
CI (riscv) / ubuntu-cpu-riscv64-native (push) Canceled after 0s
CI (cpu) / build-cmake-pkg (push) Canceled after 0s
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, ADDRESS) (push) Canceled after 0s
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, THREAD) (push) Canceled after 0s
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, UNDEFINED) (push) Canceled after 0s
CI (rpc) / ubuntu-24-rpc (push) Canceled after 0s
CI (sanitize) / ctest ([self-hosted X64 Linux], ADDRESS) (push) Canceled after 0s
CI (self-hosted) / gpu-cuda (push) Canceled after 0s
CI (self-hosted) / gpu-rocm (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-nvidia-cm (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-nvidia-cm2 (push) Canceled after 0s
CI (self-hosted) / gpu-webgpu-nvidia (push) Canceled after 0s
CI (self-hosted) / gpu-metal (push) Canceled after 0s
CI (self-hosted) / gpu-webgpu-apple (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-apple (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-intel-linux (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-intel-windows (push) Canceled after 0s
CI (self-hosted) / gpu-openvino-low-perf (push) Canceled after 0s
CI (self-hosted) / cpu-x64-high-perf (push) Canceled after 0s
CI (self-hosted) / cpu-arm64-high-perf-graviton4 (push) Canceled after 0s
CI (self-hosted) / cpu-arm64-graviton4-kleidiai (push) Canceled after 0s
CI (sycl) / windows-latest-sycl (push) Canceled after 0s
CI (virtgpu) / ubuntu-24-virtgpu (push) Canceled after 0s
CI (vulkan) / ubuntu-arm64 (push) Canceled after 0s
CI (wasm) / ubuntu-webgpu (push) Canceled after 0s
CI (webgpu) / macos (push) Canceled after 0s
Code Style Checker / model-naming (push) Canceled after 0s
EditorConfig Checker / editorconfig (push) Canceled after 0s
Release / check-release (push) Canceled after 0s
Release / get-version (push) Canceled after 0s
Release / windows-cuda (13.4, arm64) (push) Canceled after 0s
Release / windows-cuda (12.4, x64) (push) Canceled after 0s
Release / windows-cuda (13.3, x64) (push) Canceled after 0s
Release / windows-sycl (push) Canceled after 0s
Release / ubuntu-24-sycl (fp16, ON) (push) Canceled after 0s
Release / ubuntu-24-sycl (fp32, OFF) (push) Canceled after 0s
Release / ubuntu-22-rocm (7.2.1, x64, gfx908;gfx90a;gfx942;gfx1030;gfx1100;gfx1101;gfx1102;gfx1151;gfx1150;gfx1200;gfx1201) (push) Canceled after 0s
Release / windows-hip (gfx1150;gfx1151;gfx1200;gfx1201;gfx1100;gfx1101;gfx1102;gfx1030;gfx1031;gfx1032, radeon) (push) Canceled after 0s
Release / macos-cpu (arm64, arm64, -DGGML_METAL_EMBED_LIBRARY=ON -DCMAKE_OSX_DEPLOYMENT_TARGET=13.3, macos-26) (push) Canceled after 0s
Release / macos-cpu (x64, x64, -DGGML_METAL=OFF -DCMAKE_OSX_DEPLOYMENT_TARGET=13.3, macos-15-intel) (push) Canceled after 0s
Release / ubuntu-cpu (arm64, ubuntu-24.04-arm) (push) Canceled after 0s
Release / ubuntu-cpu (s390x, ubuntu-24.04-s390x) (push) Canceled after 0s
Server (sanitize) / server (RelWithDebInfo, UNDEFINED) (push) Canceled after 0s
Server (sanitize) / server (RelWithDebInfo, ADDRESS) (push) Canceled after 0s
Server (self-hosted) / server-metal (push) Canceled after 0s
Server (self-hosted) / server-cuda (push) Canceled after 0s
Server (self-hosted) / server-kleidiai (push) Canceled after 0s
Server / ubuntu (push) Canceled after 0s
Server / windows (push) Canceled after 0s
CI (apple) / macos-latest-swift (generic/platform=iOS) (push) Canceled after 0s
CI (apple) / macos-latest-swift (generic/platform=macOS) (push) Canceled after 0s
CI (apple) / macos-latest-swift (generic/platform=tvOS) (push) Canceled after 0s
Release / ubuntu-cpu (x64, ubuntu-22.04) (push) Canceled after 0s
Release / ubuntu-vulkan (arm64, ubuntu-24.04-arm) (push) Canceled after 0s
Release / ubuntu-vulkan (x64, ubuntu-22.04) (push) Canceled after 0s
Release / android-arm64 (push) Canceled after 0s
Release / ubuntu-24-openvino (push) Canceled after 0s
Release / windows-openvino (push) Canceled after 0s
Release / windows-cpu (arm64) (push) Canceled after 0s
Release / windows-cpu (x64) (push) Canceled after 0s
Release / windows (arm64, opencl-adreno, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DCMAKE_PREFIX_PATH="$env:RUNNER_TEMP/opencl-arm64-release" -DGGML_OPENCL=ON -DGGML_OPENCL_USE_ADRENO_KERNELS=ON, ggml-opencl) (push) Canceled after 0s
Release / windows (x64, vulkan, -DGGML_VULKAN=ON, ggml-vulkan) (push) Canceled after 0s
Release / ios-xcode (push) Canceled after 0s
Release / ui-build (push) Canceled after 0s
Release / release (push) Canceled after 0s
Release / ui-publish (push) Canceled after 0s
tool-call: fix Qwen 2.5 Coder support, add micro benchmarks, support trigger patterns for lazy grammars (#12034)
llama.cpp
LLM inference in C/C++
manifesto / ggml / ops / maintainer PRs / compile times / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - stb-image - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- miniaudio.h - Single-header audio format decoder, used by multimodal subsystem - Public domain
- subprocess.h - Single-header process launching solution for C and C++ - Public domain
Languages
C++
55%
C
16.2%
Python
7.1%
Cuda
5.6%
TypeScript
4.4%
Other
11.5%