0

Geometric Capacity of Transformers: A Tropical Geometry Perspective

To quantify the geometric capacity of transformers, we develop a tropical-geometric framework for analyzing the spatial partitions induced by conditioned self-attention.

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2604.14727CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

To quantify the geometric capacity of transformers, we develop a tropical-geometric framework for analyzing the spatial partitions induced by conditioned self-attention. In the zero-temperature limit, we show that fixed-key top-1 routing is exactly represented by a power diagram in query space, while an auxiliary log-lifted value parameterization yields a vector-valued tropical rational representation. For Multi-Head Self-Attention (MHSA) with sequence length N and H attention heads, the joint routing geometry is encoded by Minkowski sums of headwise Newton polytopes, giving an O(N^H) universal bound that sharpens to O((HN)^{d_{model}-1}) once the number of heads reaches the intrinsic dimension d_{model}. Extending this analysis across depth L, we derive the first tight asymptotic bounds on the number of linear regions in transformers (Θ!\left(N^{\min{H,d_{model}-1}L}\right)). We further show that finite-temperature softmax preserves the top-1 routing structure and admits exponentially decaying local approximation and differential bounds away from routing boundaries.