colibri
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
Runs a 744B-parameter MoE model on a 25GB-RAM consumer machine in pure C, streaming experts from disk.
Potential upside
- Genuine systems engineering — expert streaming makes an 'impossible' model size run locally
- Zero dependencies and pure C means it will build anywhere
Worth watching
- Disk-streamed inference is slow by nature; this is a feat, not a daily driver
A first look, not a review — this product just launched and has no user history yet. We flag what looks promising and what to check before you rely on it.
Who it's for
Local-LLM tinkerers who enjoy the frontier of what consumer hardware can host.
More AI tools