Understanding the fundamental mechanisms governing the production of meaning in the processing of natural language is critical for designing safe, thoughtful, engaging, and empowering human-agent interactions. If meaning is constituted rather than retrieved, then the search for context-independent features or circuits in the pursuit of mechanistic interpretability may be fundamentally limited. Experiments in cognitive science and social psychology have demonstrated that human semantic processing exhibits contextuality more consistent with quantum logical mechanisms than classical Boolean theories, and recent works have found similar results in large language models---in particular, clear violations of the Bell inequality in experiments of contextuality during interpretation of ambiguous expressions. In this work, we explore the CHSH |S| parameter---the metric associated with the inequality---across the inference parameter space of models spanning four orders of magnitude in scale and cross-reference our findings with MMLU, hallucination rate, and nonsense detection benchmarks. We find that the interquartile range of the |S| distribution is completely orthogonal to all external benchmarks, while overall violation rate shows weak anticorrelation with all three benchmarks. We investigate how |S| varies with sampling parameters and word order, and discuss the information-theoretic constraints that genuine contextuality imposes on prompt injection defenses and its human analogue, whereby careful construction and maintenance of social contextuality can be carried out at scale, shaping the space of possible interpretations before any particular one is reached. We consider the implications for mechanistic interpretability and how genuine contextuality sets an information-theoretic bound on the decomposability of semantic processing.
The production of meaning in the processing of natural language
Understanding the fundamental mechanisms governing the production of meaning in the processing of natural language is critical for designing safe, thoughtful, engaging, and empowering human-agent interactions.
- Preview

- Year
- 2026
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2603.20381CC-BY-4.0
- TL;DR
- Semantic Scholar