llama.cpp b10603: GLM-4.5-Air MTP is live
Pre-release b10603 ships model : support MTP in GLM-4.5-Air (#26534). Author-reported mean ~1.19× on 4×3090 (74.55 → 88.73 tok/s at --spec-draft-n-max 1). A community Strix Halo Vulkan note says 21.80 → 27.16 tok/s. No Metal / Mac numbers in the PR or release notes. The release includes llama-b10603-bin-macos-arm64.tar.gz. Same-day follow-on: GLM-4.5V conversion discussion in the PR; Gerganov later points at ggml-org/GLM-4.5V-GGUF. b10604 (24 Aug) is DeepSeek 4 -sm tensor on multi-GPU CUDA/ROCm, not a Mac-local story. vendor