article
How we used an ensemble of on-device models to build consumer AI that works offline, costs nothing to run, and never leaves the phone
article
How we used an ensemble of on-device models to build consumer AI that works offline, costs nothing to run, and never leaves the phone
article
Conscious Engines
Running a 27B model (1-bit) on my phone and letting it perform agentic tasks
article
Conscious Engines
A 203M-parameter language model, trained for 14.76B tokens and deployed as one fixed specialist.
article
Conscious Engines
Optimizing DiffusionGemma (26B-A4B-IT-4Bit) on an Apple M5 Pro, what worked, what didn't.
article
Conscious Engines
A ternary-quantized 27B model that scores like an 8B on tool-calling - but deletes active records when nobody double-checks the honest answer.
article
Conscious Engines
A walkthrough of Apple Foundation Model 3's on-device architecture — the routing changes that let a 20B-parameter model run on a phone.
article
An empirical evaluation of speculative decoding on Apple Silicon, across a 300-run parameter sweep