We’re putting $5M behind research and development of smaller, specialized AI models for real-world deployment — Read our manifesto →

    Research and Development

    Experiments in model architecture, training, evaluation, and efficient inference

    All research

    Building a tasks app that puts all iOS 27 on-device AI features to workexternal ↗

    Conscious Engines

    How we used an ensemble of on-device models to build consumer AI that works offline, costs nothing to run, and never leaves the phone

    apple silicon

    Running Bonsai 27B as an agent on my phone

    Conscious Engines

    Running a 27B model (1-bit) on my phone and letting it perform agentic tasks

    apple silicon

    How we trained a ternary mixture-of- experts model from scratch and deployed it on the Apple Watch Neural Engine

    Conscious Engines

    A 203M-parameter language model, trained for 14.76B tokens and deployed as one fixed specialist.

    apple silicon

    Squeezing a 26B diffusion LLM onto a Mac

    Conscious Engines

    Optimizing DiffusionGemma (26B-A4B-IT-4Bit) on an Apple M5 Pro, what worked, what didn't.

    apple silicon

    Does the 27B Bonsai actually deliver 27B?

    Conscious Engines

    A ternary-quantized 27B model that scores like an 8B on tool-calling - but deletes active records when nobody double-checks the honest answer.

    apple silicon

    Apple Foundation Model 3, what even is it?

    Conscious Engines

    A walkthrough of Apple Foundation Model 3's on-device architecture — the routing changes that let a 20B-parameter model run on a phone.

    apple silicon

    Speculative Decoding in MLX using DFlash

    Sabesh B

    An empirical evaluation of speculative decoding on Apple Silicon, across a 300-run parameter sweep

    apple siliconInference