Mock Modular Attention
- 0 Devlogs
- 0 Total hours
The Idea was to create an alternative complete attention mechanism that could act as a drop in replacement for softmax. It replaces stochastic, unconstrained heuristic logits with deterministic kernel weightings derived from Ramanujan's third-order mock theta functions and classical q -series and introduces an approximate modular-symmetry inductive bias directly into the forward pass of Transformer architectures.