0

Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has used such directions as static handles for behavioral steering.

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2605.09159CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has used such directions as static handles for behavioral steering. We instead treat them as dynamic signals: probes we can monitor and intervene on as reasoning unfolds. We use the term polylogue to denote the time series of alignments between persona vectors and hidden activations over the course of generation. Experiments across four open-weight models show that polylogue features contain predictive signal for correctness comparable to low-dimensional activation summaries, while remaining interpretable through their associated persona directions. They also suggest concrete steering targets, namely, which latent directions to modulate at different stages of a response. We instantiate this as a simple paragraph-conditioned intervention that improves accuracy on three of the four models but degrades the fourth, suggesting that stage-aware latent steering is possible but not yet robust. Together, this positions the polylogue as an interpretable tool for reasoning-time monitoring and intervention.