Glimmer ships on-device as frontier APIs drop price and EU Article 50 goes live

Grok 4.6, Gemini 3.7 Flash, and DeepSeek-V4-Pro hit APIs in the same two weeks Muse Glimmer 30B and Qwen3.8-27B became runnable on Apple Silicon. Enterprise evals, inference DLP, and model routing moved closer to production. The EU began enforcing AI Act transparency on 2 August 2026.

Window ~2026-07-25 to 2026-08-22. Primary sources only. Vendor evals and press extras flagged.

Breaking AI News

xAI Grok 4.6

Long-running agents + visual work. Same-day Cursor, Grok Build, SpaceXAI API, OpenRouter/Vercel/Cloudflare. $2 / $6 per 1M in/out (fast variant 2×). Docs: 500k context, cutoff 2026-02-01. Copilot Aug 14; Bedrock GA Aug 19 at $2 / $0.50 cached / $6. vendor

Enterprise

7 stories

Anthropic Inference Hooks (inline DLP, beta)

Claude Enterprise can POST the conversation to a customer/vendor HTTPS security server and wait for allow/deny before inference. One hook covers claude.ai, Cowork, and Claude Code.

7 stories Open section

Local

5 stories

Ollama 0.32.7–0.32.15

0.32.7: ollama run muse-glimmer:30b-mlx (MLX-first on Apple Silicon). 0.32.12: qwen3.8:27b and mlx variant.

5 stories Open section

MLX

4 stories

MLX 0.32.1

Metal gemv_wide; NAX attention unroll; head dim 96/72; higher qmv batch on M5; GGUF metadata/offset fixes; zero-copy CPU import on unified memory.

4 stories Open section

Frontier

5 stories

Google Gemini 3.7 Flash

Workhorse for coding/agents, three weeks after 3.6 Flash. Intro $0.75 / $3.75 per 1M through 2026-12-31; then $1.50 / $7.50.

5 stories Open section

OSS tools

6 stories

vLLM 0.27.0

Day-0 Kimi K3; Qwen3.5 dense/MoE; PyTorch 2.13 / Triton 3.7.1 (breaking). FlashAttention 4 on SM100; DeepSeek-V4 kernel/TTFT; Rust frontend gRPC.

SGLang 0.5.17

Day-0 Kimi K3 (2.8T LatentMoE, 1M context, native MXFP4) and MiniMax-H3 video+stereo-audio. Initial Rust frontend.

6 stories Open section

OSS models

7 stories

Kimi K3 (~2.8T, custom license)

LMSYS/SGLang: first open ~3T-class multimodal; LatentMoE 896 experts top-16; 1M context; native MXFP4. License is Kimi K3, not MIT/Apache.

7 stories Open section

Other

2 stories
2 stories Open section

Try this on a Mac

5 stories

mlx-vlm screenshot-to-tools loop

pip install -U mlx-vlm (≥0.6.12). python -m mlx_vlm.generate --model mlx-community/Muse-Glimmer-30B-4bit. 19.4 GB weights; 32 GB comfortable.

5 stories Open section