0

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Aligned language models routinely misreport under non-evidential pressure: they cave to a confident user, yet fail to revise when genuine evidence arrives. We cast this as a failure of internal incentive-compatibility and study the two demands, resist (ignore forbidden pressure)…

Preview
Year
2026
Hosting
Full text hostedCC-BY-4.0

Cite

Notes

Only stored in your browser.

Attribution

Abstract & full text
arxiv.org/abs/2607.12985CC-BY-4.0
TL;DR
Semantic Scholar
Attribution policy →

Abstract

Aligned language models routinely misreport under non-evidential pressure: they cave to a confident user, yet fail to revise when genuine evidence arrives. We cast this as a failure of internal incentive-compatibility and study the two demands, resist (ignore forbidden pressure) and update (follow licensed evidence), on a Bayesian-witness benchmark with known posteriors, where the same user disagreement is evidence or pressure purely by stated source reliability, removing the evidence/pressure confound by construction. Using interchange interventions rather than probes, we causally localize low-rank report coordinates for answer, confidence, and caveat, establishing causal sufficiency at a late intervention site rather than uniqueness or necessity, with a causal cross-talk matrix showing strong own-coordinate control and only small cross-effects (partial functional disentanglement). We then introduce a training-free counterfactual report-coordinate (CRC) clamp that references the model's own report under an incentive-neutralized counterfactual of the prompt. The two-pass full-window clamp attains resist and update of 1.00 jointly (Wilson 95% CI [0.99,1.00]; the rank-16 projection alone reaches 0.88/0.90), which we read as a causal certificate and upper bound under a constructible reference, not a claim of a deployed solution. Tested global decoding and fixed-direction steering trade one objective against the other, and resist-only training collapses updating to 0.01. The deployable single-pass compilation is lossy (0.73/0.97). The mechanism and the clamp reproduce across three model families and transfer to a natural sycophancy benchmark with significant paired improvements. Our contribution is the interface and certification method: activation-level counterfactual incentive-invariance as a structural primitive for internal incentive-compatibility.