← Back to this issue

MLX 3 stories

SenseNova-U1.5-8B-MoT MLX-Swift artifacts

Four mlx-community repos were created 2026-08-24 ~04:14–04:26 UTC (4-bit/8-bit 8-step distill, 8-bit, bf16). Runtime is xocialize/sensenova-u1-swift, not mlx-lm. The 8-bit card: ~19.9 GB MLX; “Performance (M5 Max): peak 22.9 GB · 1024² ≈ 7.4 s (8-step, cfg 4)”; “8-bit reproduces the bf16 image at fixed seed (cos 0.998, 33 dB)”. The GitHub listing says “1024² in 3.2 s, 14.8 GB peak”. Those are different figures on two primary pages — do not merge them. Upstream: sensenova/SenseNova-U1.5-8B-MoT, Apache-2.0.

mlx-vlm 0.6.15 (and 0.6.14 dump)

0.6.15: stop a batched row’s answer depending on its neighbour’s length, plus tests with mlx 0.32.1. 0.6.14 is the large feature dump (PLaMo 2.1 VL, Nemotron-VoiceChat-11B, Nemotron-Parse, GOT-OCR 2.0, rerank endpoints, /v1/settings). This is the VLM follow-through for already-covered MLX 0.32.1.

LTX-2.5 MLX encoder (video, not LLM)

int8 Gemma-4 text encoder for LTX-2.5. The card claims swapping the encoder drops whole-run peak to 14.6 GB / 15.4 GB vs ~24.8 GB bf16 encoder floor. Pair with mlx-community/ltx-2.5-mlx-ditq8. LTX-2 Community License with a revenue gate. Video stack, not a chat model.