13 Commits
Author SHA1 Message Date
Guido Imperiale a52077c4ca chat : Align Laguna-S-2.1 chat template to huggingface (#26232)
CI (cpu) / ubuntu (x64, ubuntu-22.04) (push) Failing after 30s
CI (CUDA, ubuntu) / hip (push) Failing after 55s
CI (android) / ndk (push) Failing after 1m27s
CI (android) / arm64 (push) Failing after 1m33s
CI (android) / default (push) Failing after 2m39s
CI (CUDA, ubuntu) / cuda (push) Failing after 1m14s
CI (sanitize) / ctest (ubuntu-24.04, THREAD) (push) Failing after 8s
CI (CUDA, ubuntu) / musa (push) Failing after 5m44s
CI (sycl) / ubuntu-24-sycl (fp16, ON) (push) Failing after 3m44s
CI (sycl) / ubuntu-24-sycl (fp32, OFF) (push) Failing after 2m56s
CI (vulkan) / ubuntu-llvmpipe (push) Failing after 1m42s
CI (webgpu) / format (push) Successful in 19s
CI (webgpu) / ubuntu (push) Failing after 8s
CI (cpu) / windows (x64, x64-openblas, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_OPENMP=OFF -DGGML_BLAS=ON -DGGML_BLA… (push) Canceled after 0s
CI (cpu) / windows (x64, x64-vulkan, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_VULKAN=ON) (push) Canceled after 0s
CI (3rd-party) / ubuntu-24-llguidance (push) Canceled after 0s
CI (apple) / macos-latest-arm64 (push) Canceled after 0s
CI (apple) / macos-latest-x64 (push) Canceled after 0s
CI (apple) / macos-latest-ios-xcode (push) Canceled after 0s
CI (apple) / macos-latest-tvos (push) Canceled after 0s
CI (apple) / macos-latest-visionos (push) Canceled after 0s
Build relocatable cmake package / linux (push) Canceled after 0s
CI (sanitize) / ctest ([self-hosted X64 Linux], UNDEFINED) (push) Canceled after 0s
CI (cpu) / ubuntu (arm64, ubuntu-24.04-arm) (push) Canceled after 0s
CI (cpu) / windows (arm64, arm64, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON) (push) Canceled after 0s
CI (cpu) / windows (x64, x64-cpu-static, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DBUILD_SHARED_LIBS=OFF) (push) Canceled after 0s
CI (ibm) / ubuntu-24-s390x (push) Canceled after 0s
CI (ibm) / ubuntu-24-ppc64le (push) Canceled after 0s
CI (opencl) / windows-2025-opencl-adreno (push) Canceled after 0s
CI (openvino) / ubuntu-24-openvino (push) Canceled after 0s
CI (openvino) / openvino-windows-2022 (push) Canceled after 0s
CI (riscv) / ubuntu-cpu-riscv64-native (push) Canceled after 0s
CI (cpu) / build-cmake-pkg (push) Canceled after 0s
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, ADDRESS) (push) Canceled after 0s
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, THREAD) (push) Canceled after 0s
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, UNDEFINED) (push) Canceled after 0s
CI (rpc) / ubuntu-24-rpc (push) Canceled after 0s
CI (sanitize) / ctest ([self-hosted X64 Linux], ADDRESS) (push) Canceled after 0s
CI (self-hosted) / gpu-cuda (push) Canceled after 0s
CI (self-hosted) / gpu-rocm (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-nvidia-cm (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-nvidia-cm2 (push) Canceled after 0s
CI (self-hosted) / gpu-webgpu-nvidia (push) Canceled after 0s
CI (self-hosted) / gpu-metal (push) Canceled after 0s
CI (self-hosted) / gpu-webgpu-apple (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-apple (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-intel-linux (push) Canceled after 0s
CI (self-hosted) / gpu-vulkan-intel-windows (push) Canceled after 0s
CI (self-hosted) / gpu-openvino-low-perf (push) Canceled after 0s
CI (self-hosted) / cpu-x64-high-perf (push) Canceled after 0s
CI (self-hosted) / cpu-arm64-high-perf-graviton4 (push) Canceled after 0s
CI (self-hosted) / cpu-arm64-graviton4-kleidiai (push) Canceled after 0s
CI (sycl) / windows-latest-sycl (push) Canceled after 0s
CI (virtgpu) / ubuntu-24-virtgpu (push) Canceled after 0s
CI (vulkan) / ubuntu-arm64 (push) Canceled after 0s
CI (wasm) / ubuntu-webgpu (push) Canceled after 0s
CI (webgpu) / macos (push) Canceled after 0s
Code Style Checker / model-naming (push) Canceled after 0s
EditorConfig Checker / editorconfig (push) Canceled after 0s
Release / check-release (push) Canceled after 0s
Release / get-version (push) Canceled after 0s
Release / windows-cuda (13.4, arm64) (push) Canceled after 0s
Release / windows-cuda (12.4, x64) (push) Canceled after 0s
Release / windows-cuda (13.3, x64) (push) Canceled after 0s
Release / windows-sycl (push) Canceled after 0s
Release / ubuntu-24-sycl (fp16, ON) (push) Canceled after 0s
Release / ubuntu-24-sycl (fp32, OFF) (push) Canceled after 0s
Release / ubuntu-22-rocm (7.2.1, x64, gfx908;gfx90a;gfx942;gfx1030;gfx1100;gfx1101;gfx1102;gfx1151;gfx1150;gfx1200;gfx1201) (push) Canceled after 0s
Release / windows-hip (gfx1150;gfx1151;gfx1200;gfx1201;gfx1100;gfx1101;gfx1102;gfx1030;gfx1031;gfx1032, radeon) (push) Canceled after 0s
Release / macos-cpu (arm64, arm64, -DGGML_METAL_EMBED_LIBRARY=ON -DCMAKE_OSX_DEPLOYMENT_TARGET=13.3, macos-26) (push) Canceled after 0s
Release / macos-cpu (x64, x64, -DGGML_METAL=OFF -DCMAKE_OSX_DEPLOYMENT_TARGET=13.3, macos-15-intel) (push) Canceled after 0s
Release / ubuntu-cpu (arm64, ubuntu-24.04-arm) (push) Canceled after 0s
Release / ubuntu-cpu (s390x, ubuntu-24.04-s390x) (push) Canceled after 0s
Server (sanitize) / server (RelWithDebInfo, UNDEFINED) (push) Canceled after 0s
Server (sanitize) / server (RelWithDebInfo, ADDRESS) (push) Canceled after 0s
Server (self-hosted) / server-metal (push) Canceled after 0s
Server (self-hosted) / server-cuda (push) Canceled after 0s
Server (self-hosted) / server-kleidiai (push) Canceled after 0s
Server / ubuntu (push) Canceled after 0s
Server / windows (push) Canceled after 0s
CI (apple) / macos-latest-swift (generic/platform=iOS) (push) Canceled after 0s
CI (apple) / macos-latest-swift (generic/platform=macOS) (push) Canceled after 0s
CI (apple) / macos-latest-swift (generic/platform=tvOS) (push) Canceled after 0s
Release / ubuntu-cpu (x64, ubuntu-22.04) (push) Canceled after 0s
Release / ubuntu-vulkan (arm64, ubuntu-24.04-arm) (push) Canceled after 0s
Release / ubuntu-vulkan (x64, ubuntu-22.04) (push) Canceled after 0s
Release / android-arm64 (push) Canceled after 0s
Release / ubuntu-24-openvino (push) Canceled after 0s
Release / windows-openvino (push) Canceled after 0s
Release / windows-cpu (arm64) (push) Canceled after 0s
Release / windows-cpu (x64) (push) Canceled after 0s
Release / windows (arm64, opencl-adreno, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DCMAKE_PREFIX_PATH="$env:RUNNER_TEMP/opencl-arm64-release" -DGGML_OPENCL=ON -DGGML_OPENCL_USE_ADRENO_KERNELS=ON, ggml-opencl) (push) Canceled after 0s
Release / windows (x64, vulkan, -DGGML_VULKAN=ON, ggml-vulkan) (push) Canceled after 0s
Release / ios-xcode (push) Canceled after 0s
Release / ui-build (push) Canceled after 0s
Release / release (push) Canceled after 0s
Release / ui-publish (push) Canceled after 0s
2026-08-10 05:20:59 -05:00
Joe Rowell 1f66c3ce1c Add support for Laguna XS.2 & M.1 (#25165) 2026-07-22 09:54:08 +08:00
2969d6d15d model: add Hy3 (hy_v3) support with MTP speculative decoding (#25395)
* model: add Hy3 (hy_v3) architecture support

Adds Tencent Hunyuan 3 (HF architecture HYV3ForCausalLM, GGUF arch
hy_v3): a MoE decoder stack with per-head Q/K RMSNorm, a sigmoid
router with expert selection bias, an always-active ungated shared
expert, and leading dense block(s) (first_k_dense_replace).

The base implementation is ported from charlie12345's fork
(https://github.com/charlie12345/ROCmFPX, src/models/hyv3.cpp),
adapted to current mainline APIs (hparams.n_layer(), build_qkv,
build_moe_ffn with fused gate_up + scale tensors, output_s).

Note: blk.N.exp_probs_b is stored without a .bias suffix for
compatibility with existing hy_v3 GGUFs produced by that fork.

Co-Authored-By: charlie12345 <charlie12345@users.noreply.github.com>
Co-authored-by: Piotr Wilkin <ilintar@gmail.com>
Assisted-by: Claude Fable 5
2026-07-14 00:31:04 +02:00
2d973636e2 chat: trim messages sent to StepFun parser (fixes long reasoning loops) (#25238)
* chat: trim messages sent to StepFun parser (fixes long reasoning loops)

* add regression test; remove duplicate template

* chat: trim StepFun content parts before rendering

The StepFun trim workaround ran on the already-rendered messages, where
typed content parts have been concatenated into a single string, so the
per-part whitespace could no longer be reached. Move the trim ahead of
rendering and apply it to content_parts text as well as the string
content and reasoning_content. Adds a content-parts regression test.

Co-Authored-By: Piotr Wilkin <ilintar@gmail.com>
Assisted-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: tarruda <tpadilha84@gmail.com>
2026-07-03 23:12:11 +02:00
Piotr Wilkin (ilintar) a6dff71270 chat: fix whitespace problems once and for all (#24624)
* chat: fix whitespace problems once and for all

* Purge trailing spaces from grammar generation

* Revert "Purge trailing spaces from grammar generation"

This reverts commit b0827ecb7d4767f37cefd751b3646f98d5303891.
2026-06-15 08:27:10 +02:00
e2ef8fe42c server: fix checkpoints creation (#22929)
* common : add common_chat_split_by_role

* cont : fix spans to reach end of message

* server: fix checkpoints creation

- extract message_spans from chat templates
- find the prompt token position before the latest user message
- split prompt batching at that position
- create a context checkpoint before the latest user input
- avoid periodic mid-prompt checkpoints when that position is known
- handle multimodal prompts when mapping text/template positions to server prompt tokens
- add --checkpoint-min-step to control minimum spacing between checkpoints

* cont : clean-up

* Support autoparser detection for message barriers

* server: fix message span delimiter and update docs

---------

Co-authored-by: Alde Rojas <hello@alde.dev>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>
2026-05-25 08:56:18 +03:00
Piotr Wilkin (ilintar) dcad77cc3b chat: fix handling of space in reasoning markers (#22353)
* chat: fix handling of space in reasoning markers

* fix tests

* whitespace
2026-04-25 21:24:13 +02:00
Piotr Wilkin (ilintar) 1f5d15e665 common/parser: fix reasoning whitespace bugs + extra parser tests (#21085)
* fix whitespace reasoning issues + add reconstruction tests

* Proper fix

* fix Nemotron autoparser test expectations to include newline in marker
2026-03-28 07:29:26 +01:00
Jhen-Jie Hong 7a0b6a635e common/autoparser : detect reasoning markers when enable_thinking changes system prompt (#20859) 2026-03-23 08:35:27 +01:00
Piotr Wilkin (ilintar) b1c70e2e54 common/parser: fix nasty bug causing subtle corruption of generation prompt (#20825) 2026-03-21 00:19:04 +01:00
Piotr Wilkin (ilintar)andGeorgi Gerganov 5e54d51b19 common/parser: add proper reasoning tag prefill reading (#20424)
* Implement proper prefill extraction

* Refactor cli parameters, update docs, move reasoning budget sampler part to common/reasoning-budget.cpp

* Update tools/server/server-task.cpp

* refactor: move grammars to variant, remove grammar_external, handle exception internally

* Make code less C++y

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2026-03-19 16:58:21 +01:00
Aldehir Rojas 451ef08432 common : gracefully handle incomplete output (#20191)
* common : handle incomplete UTF-8 at end of input in PEG parser

* cont : if reached end prematurely, emit needs_more_input to propagate partial output

* cont: refactor peg parse context to add lenient flag

* cont : remove partial flag, keep lenient flag
2026-03-08 17:17:02 +01:00
Piotr Wilkin (ilintar) 566059a26b Autoparser - complete refactoring of parser architecture (#18675)
* Autoparser - full single commit squish

* Final pre-merge changes: minor fixes, Kimi 2.5 model parser
2026-03-06 21:01:00 +01:00