Conscious Engines
Story of Attention - Part 1: Before the Transformer
From n-gram counting and RNNs to the first attention mechanisms.
guide
Read more →Experiments in model architecture, training, evaluation, and efficient inference
Conscious Engines
From n-gram counting and RNNs to the first attention mechanisms.
Conscious Engines
How scaled dot-product attention, multi-head attention, and positional encoding built the Transformer.
Conscious Engines
How sparse attention makes long-context processing more efficient.
Conscious Engines
How MQA and GQA reduce the memory cost of Transformer inference.
Conscious Engines
How kernels, decay, and gating make attention scale linearly.
Conscious Engines
How latent compression and differential attention shape modern language models.