Local LLM Inference

2 tools reviewed

7
AirLLM Review 2026 — Run a 70B Model on a Single 4GB GPU (No Quantization, No Distillation)

AirLLM (27k stars, Apache-2.0) runs 70B LLMs on a single 4GB GPU with no quantization, distillation, or pruning — by streaming layers and per-expert shards from disk. It now handles DeepSeek-V3 671B on ~12GB and Kimi K3 2.8T on under 4GB. Full review of how it works, real measured speeds (including 292 s/token on Kimi K3), who it's actually for, and the HN skepticism.

7.6
slotstream Review 2026 — Stream a 104GB Qwen3.8-Flash-Next MoE From SSD and Run It in 32GB on a 48GB Mac

slotstream is a single-binary Swift + MLX engine (MIT, created 2026-08-28) that runs Qwen3.8-Flash-Next — a 125B-parameter MoE, 105GB on disk at 4-bit — on Apple Silicon Macs that can't hold it in RAM, by keeping the 3.8GB dense trunk resident and streaming routed experts from SSD through a fixed pool of cache slots. Measured on a 48GB M5 Pro: ~12 tok/s warm decode, ~2s engine start, 32GB auto-sized peak. It speaks the Ollama and OpenAI chat APIs, supports an optional speculative-decoding draft head (~×1.24), and ships byte-identical output across cache sizes as a standing test. This review covers the memory design, the tier table, real measurements, the honest limits (Apple Silicon only, one model only, 105GB download), and the HN reception.