← Back to Reviews | Local LLM Inference

slotstream Review 2026 — Stream a 104GB Qwen3.8-Flash-Next MoE From SSD and Run It in 32GB on a 48GB Mac

Marcus Webb · · Rated 7.6/10 · Free, MIT-licensed, open source (Swift/MLX). The binary is free; weights come from a Hugging Face mirror of pipenetwork/Qwen3.8-Flash-Next-MLX-4bit (105.3GB across 25 files, one-time, sha256-verified) and remain under the Qwen community license
7.6 / 10
Ease of Use 6.5
Features 7
Value for Money 8
Performance 8.5
Support & Ecosystem 6.5

✅ Pros

  • • Genuinely runs a 104GB-at-4-bit 125B MoE on a 48GB Mac: ~12 tok/s warm decode measured on an M5 Pro, ~2s engine start because only the 3.8GB dense trunk loads up front, ~32GB auto-sized peak
  • • The memory design is the product: the 3.8GB trunk stays resident, routed experts (512 per layer, 10 active per token) are pread from SSD into a fixed pool of slots shared across all 48 layers, and the pool is auto-sized to the machine and resized live every 15s when other apps want memory
  • • Byte-identical output across cache sizes is a standing test — greedy decoding with a 4GB cache equals a 24GB cache, so shrinking the cache costs speed, never correctness
  • • Honest, measurement-first engineering: a tier table generated by `slotstream doctor --sim-ram N`, a MEASUREMENTS.md that documents failed experiments, sha256-verified resumable downloads with 8 parallel TCP connections, and signed CI-built releases verifiable via `gh attestation verify`
  • • Drop-in API compatibility: `slotstream serve` on port 11434 implements the Ollama/OpenAI chat subset that curl, Open WebUI and the OpenAI SDKs use, with clear 400s for unsupported features (tools, images, JSON-schema, logprobs)
  • • Optional multi-token prediction: a 1.5GB draft head predicts the next token (~86% correct), verified in a two-token pass for measured ×1.24 decode (10.3 → 12.8 tok/s) and ×1.33 on a code prompt

⚠️ Cons

  • • Apple Silicon only, and not by accident: the engine is MLX + Metal and the design assumes unified memory — Linux/Windows with discrete GPUs would be a second engine, and it's explicitly not on the roadmap
  • • One model only: it runs exactly `qwen3.8-flash-next:4bit`, built around that model's geometry; Qwen3.8-27B is a different (dense) model and won't work — for anything else use llama.cpp or Ollama
  • • The 105GB weight download is the real barrier: ~16 minutes on a 1 Gbit/s datacenter link, ~2h20 at 100 Mbps, ~9h at 25 Mbps — disk space (~110GB free) bites before RAM does, so 512GB Macs are the realistic minimum
  • • Only the 48GB row of the tier table is measured on real hardware; 8/16/24/32GB rows are estimates from a curve, and smaller Macs also have slower SSDs
  • • The slow axis is prompt processing: every token is read before the first token appears, so an 8k prompt waits ~39s on a 48GB Mac and ~90s+ on 16GB; context is capped at 32,768 tokens per request (the largest measured, not a memory limit)
  • • Young project (created 2026-08-28): 0.2.0 broke the Ollama CLI (fixed on main, not yet released), macOS 14/15 have only had the installer exercised, and power consumption wasn't measured at review time
Best For

Owners of 48GB-class Apple Silicon Macs (M-series Pro/Max, 512GB+ disk) who want the full 125B Qwen3.8-Flash-Next running locally with an Ollama-compatible API and are willing to spend ~16 minutes downloading 105GB — plus 16-32GB Mac owners who want the best measured-effort MoE experience their machine can stream, and anyone evaluating SSD expert-streaming as a way to run models bigger than their RAM

Pricing

Free, MIT-licensed, open source (Swift/MLX). The binary is free; weights come from a Hugging Face mirror of pipenetwork/Qwen3.8-Flash-Next-MLX-4bit (105.3GB across 25 files, one-time, sha256-verified) and remain under the Qwen community license

The Pitch: A 125B MoE That Fits in 32GB

On August 28, 2026, Carlos Galarza released carloslfu/slotstream with a deceptively simple promise: run Qwen3.8-Flash-Next — a 125B-parameter mixture-of-experts model, 105GB on disk at 4-bit — on a Mac that can’t hold it. The stock loader took his 48GB MacBook Pro into 48GB of swap before producing a single token. His fix wasn’t a faster kernel; it was a memory design: keep the 3.8GB dense trunk resident, stream the routed experts from SSD through a fixed pool of cache slots, size that pool to what the machine actually has, and give memory back when other apps want it.

The Show HN (“Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s”) drew 225 points and 106 comments — one of the bigger local-LLM launches of the week — and the star count climbed past 240 in six days. The reception wasn’t all praise: commenters pointed at a crowded field of similar projects (mlx-moe-offload, streamlx, mlx-flash, Mference, SwiftLM), and the author’s response was notable — rather than argue, he committed to adding a benchmark/comparison table, cleaning up the README after writing-style criticism, and measuring power consumption. That responsiveness, plus the README’s insistence that every number ship with its method, is the project’s real differentiator.

The Memory Design: Why mmap Fails and Streaming Works

Almost all of the model’s bytes sit in two places: 68GB of routed experts (512 per layer, 10 active per token) and a 32GB n-gram table. The dense trunk is only 3.8GB and stays resident. Experts are read with pread into a fixed pool of cache slots shared by all 48 layers, so hot layers borrow slots from cold ones.

Why not just mmap the file? Because MLX can’t materialize part of a memory-mapped tensor: a top-10 expert gather evaluates all 512 experts of that layer, so an mmap path loads ~100GB and dies. The stock mlx_lm.load() route is exactly what took the dev machine into swap. slotstream’s answer is a sweep: a pass of 256+ tokens streams each layer’s experts through staging in contiguous reads and runs the expert math as grouped matrix multiplies instead of one small matvec per token over the cache. Measured against the previous release at a 16GB target: an 8k prompt went 91 → 184 tok/s, ordinary prose 66 → 140, and at the 8.1GB floor 51 → 93, at a peak 1.5GB lower.

The Tier Table: What Your Mac Actually Gets

The tier rows come straight from slotstream doctor --sim-ram N, so you can reproduce them:

Your Macslotstream takesWarm decode
8GB8.1GB (the floor)~3 tok/s, and doctor warns it will page
16GB10GB~4 tok/s
24GB16GB~8 tok/s
32GB22GB~9 tok/s
48GB and up33GB (more buys nothing)~12 tok/s

Auto-sizing takes the lowest of three limits (33GB, 70% of RAM, and 2GB under the Metal working-set limit) and shrinks further while other apps hold memory. The 33GB cap is the knee of the measured curve — in a GB-at-a-time sweep, nothing between 34 and 84GB decoded any faster, so a 128GB Mac gets the same plan as a 48GB one. While running, slotstream re-checks every 15s and resizes the cache between requests; output is byte-identical across resizes. You can override with --memory-gb, --max-ram-percent, or --experts-per-layer/--pool-gb.

Speeds, the Draft Head, and the Honest Numbers

Decode is the easy part: ~12 tok/s warm on 48GB, and follow-up turns only prefill what’s new, so time-to-first-token stays flat as a chat grows (6.0s on the eighth turn instead of 25.8s). The model ships a draft head that predicts the token after next; with --mtp (default auto), slotstream drafts and verifies it in one two-token pass — the draft is right 86% of the time, and auto turns it on only where the expert cache can still reach 120 experts per layer after the head’s 1.6GB (a ~28GB target). Measured: ×1.24 decode (10.3 → 12.8 tok/s), ×1.33 on code, ×1.18 with default sampling.

The limits are stated with the same precision. Context is capped at 32,768 tokens per request — the largest slotstream has measured, not a memory limit (the model is trained for 262,144; context state costs ~27KiB/token, so a full 32k context is under 1GB). The wait for a full 32k prompt is ~3 minutes on a 48GB Mac, ~6.4 minutes on 16GB. The 105GB download is the other gate: pull opens eight TCP connections (~112MB/s on a datacenter link, ~70MB/s bounded by round-trip to Hugging Face otherwise), all 25 files are checked against sha256 hashes compiled into the binary, and interrupted pulls resume safely.

Honest Limitations and the Crowded Field

slotstream is Apple Silicon only by design — MLX and Metal, with the cache sized against the single unified memory the OS, apps, and GPU share. A discrete card has two budgets with a bus between them, so Linux/Windows would be a second engine; the README says it isn’t on the roadmap and points you to llama.cpp instead. It runs exactly one model (qwen3.8-flash-next:4bit), and the per-user lock allows one process at a time. macOS 14/15 have only had the installer exercised, and the 0.2.0 release broke the Ollama CLI (fixed on main, verified with a real ollama run, shipping next release). Crucially, the author is upfront that the smaller tiers are estimates awaiting real measurements — the repo actively solicits measured rows with a ten-minute procedure, and its own “Related projects” section lists eight competing projects rather than pretending they don’t exist.

Verdict and Who It’s For

slotstream’s real contribution is narrower than “run big models on Macs” — it’s a rigorously measured demonstration of SSD expert-streaming done right: one model, one binary, a planner that sizes the cache to the machine and resizes it under pressure, an exact-prefix conversation cache, byte-identical output across cache sizes as a standing test, and published numbers with methods. If you own a 48GB Apple Silicon Mac and want Qwen3.8-Flash-Next locally, it’s the most honest path to ~12 tok/s without buying a 128GB machine. On 16-32GB Macs it’s a real, if slower, option; elsewhere, llama.cpp remains the mature choice. A Silver-tier pick for local-MoE enthusiasts who value measured engineering over hype.

Review based on public repository contents, README, MEASUREMENTS.md references, and HN thread 49524447 as of 2026-09-03. Star/fork counts and speed figures reflect a project six days old at review time.

slotstream Qwen3.8-Flash-Next MLX Apple-Silicon MoE Expert-Streaming SSD Local-LLM Swift Ollama Open-Source Speculative-Decoding