Probabilistic text generators supply conditional distributions over tokens and complete verbal continuations, whereas scientific use often requires a posterior over a finite state. Large language models are the leading example: phrase probabilities depend on prompt wording, and model-printed percentages are generated text rather than state posteriors. More generally, we ask when an observable language law can support a reproducible posterior over declared states. A semantic map groups meaning-equivalent continuations; held-out cases with reference posteriors identify a semiparametric inverse from grouped language probabilities to state probabilities. The language law remains nonparametric and no hidden model quantity is used. This is principally a theory and methods paper and makes several contributions. On the theoretical side, we derive conditions for existence, identification, stable recovery, and sequential updating; concentration, asymptotic, and nonparametric rates; identified sets under truncated probabilities; and a minimax boundary for uniform stability. On the empirical side, theorem- directed simulations verify recovery rates, compatible-set coverage, and stability gates, while two frozen language-model studies illustrate held-out recovery and conformal coverage. The results specify when observable language probabilities can provide an auditable state measurement without being interpreted as internal belief.
Identification and Learning of Semantic Observation Kernels: Partial Observation, Uniform Recovery, & Minimax Limits
Probabilistic text generators supply conditional distributions over tokens and complete verbal continuations, whereas scientific use often requires a posterior over a finite state.
- Preview

- Year
- 2026
- Hosting
- Full text hostedCC-BY-4.0
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2607.23130CC-BY-4.0
- TL;DR
- Semantic Scholar