Summary

The quanta hypothesis proposes that neural networks acquire knowledge and capabilities in discrete chunks, or quanta, rather than through wholly continuous improvement. Although each quantum produces a step-like improvement, aggregating many small quanta learned at different times can produce an apparently smooth loss curve. If common quanta are learned before rare ones, a power-law distribution of their frequencies can produce familiar neural scaling laws (1,2).

See also

1.
Michaud EJ, Liu Z, Girit U, Tegmark M. The Quantization Model of Neural Scaling. 2023; Available from: https://arxiv.org/abs/2303.13506
2.
Michaud EJ. On neural scaling and the quanta hypothesis. Blog post; 2026. Available from: https://ericjmichaud.com/quanta/