The transformer has become the bottleneck: startups trying to replace attention
Attention, the mechanism that made transformers powerful, now drives up the cost of long-context and agent workloads. Three startups are attacking the problem with sparse attention, rolling summaries and liquid neural networks.
Research