← Back to this issue

Local 3 stories

llama.cpp b10603: GLM-4.5-Air MTP is live

Pre-release b10603 ships model : support MTP in GLM-4.5-Air (#26534). Author-reported mean ~1.19× on 4×3090 (74.55 → 88.73 tok/s at --spec-draft-n-max 1). A community Strix Halo Vulkan note says 21.80 → 27.16 tok/s. No Metal / Mac numbers in the PR or release notes. The release includes llama-b10603-bin-macos-arm64.tar.gz. Same-day follow-on: GLM-4.5V conversion discussion in the PR; Gerganov later points at ggml-org/GLM-4.5V-GGUF. b10604 (24 Aug) is DeepSeek 4 -sm tensor on multi-GPU CUDA/ROCm, not a Mac-local story. vendor

Ollama 0.33.0-rc2 and 0.32.15 TTFT

Pre-release 0.33.0-rc2: Claude Desktop per-model on/off from the menu bar, Apps view, cache restore-point fixes, and a fix that disabled Claude Code’s “tokens left” system message which Ollama had moved to the front of the prompt and broke the KV cache on every request. Stable 0.32.15 (19 Aug) claims TTFT dropped from ~995 ms to ~524 ms in Ollama’s benches, plus Qwen 3.8 system-message normalization and MLX/llama.cpp dependency updates. vendor

mlx-lm still has no official Muse Glimmer release

Latest tagged mlx-lm remains v0.31.3 (2026-04-22). PR #1710 “Add Muse Glimmer (Meta) text model support” is still open. A contributor-reported generate on M5 Max 128 GB, mlx 0.32.0, mlx-community/Muse-Glimmer-30B-bf16: prompt 64 tokens at 14.150 tok/s, generation 700 tokens at 10.098 tok/s, peak 55.777 GB. That is a PR comment, not a shipped mlx-lm release. Saturday already covered the model and mlx-community 4-bit.