On Next-Gen Transformer: Loops Are Not What You Need
21 days ago
- The Loop Transformer (weight tying) entangles addressing and content by modifying Q, K, and V simultaneously, limiting fine-grained memory control.
- Memory injection points in a Transformer include the residual stream, normalization parameters, attention metrics, biases, summation range, head set, and FFN.
- Recursive attention separates Q, K, and V modifications, enabling stack operations (push/pop) on context and thinking about thinking.
- Context is organized hierarchically (token, reasoning block, summary, macro-step, rule library) with L3 summaries that must allow recovery of original reasoning.
- MEMENTO implements destructive pop via physical KV eviction, while INCEPTION (Recursive Transformer) controls KV reads with learnable recursive masks for composable context.
- The Recursive Transformer reduces active context length through block-based masking and summaries, naturally fitting sparse attention with a learnable mask before the Indexer.
- This architecture represents a third-generation LLM beyond RL-based long-CoT models (o1, R1), aiming for Recursive Self-Improvement (RSI).