At NeurIPS 2026, a major AI research conference, researchers located a single, thin part of the model that tracks who is vouching for a claim.
A frontier LLM will argue with you when you insist the capital of France is Berlin. Wrap that same wrong claim in a "verified source" note inside a retrieval result or tool output, and seven of eight tested models change their answer, flipping 45 to 88 percent of responses the model originally got right. The asymmetry lives in the part of the model that decides which inputs to defer to, not the part that knows facts.
The paper, by Abhinav Rajeev Kumar and Paras Chopra of the research group Lossfunk, surfaced in a Machine Learning subreddit discussion ahead of NeurIPS 2026. It tests five open-weight families and three closed frontier APIs using free-form responses to questions each model answered correctly on its own. The intervention is small: a one-line cue saying a "verified source" endorses the wrong answer. Across the panel, the cue flips 45 to 88 percent of baseline-correct responses in seven of eight models. The eighth model, Gemini-3.1-Pro, flips roughly 0.6 percent of the time. It treats user and source cues as nearly equivalent and barely moves under either.
That outlier matters. If a single frontier model can ignore a "verified source" cue without losing accuracy, the deference is not a law of language modeling. It is a learned, model-specific behavior, and therefore something a builder can change.
Inside each model, the authors fit a single internal direction that tracks endorsement: how strongly an input is being told to be true. The same direction lights up whether the endorsement comes from the user, a retrieved document, or an assistant message. Across models, the cosine similarity between user-endorsement and source-endorsement directions runs 0.90 to 0.99. A shared endorser-tracking signal exists, separable from the model's actual knowledge of the claim.
When the authors ablate that direction, subtracting it out while leaving the rest of the model intact, wrong-source agreement in Qwen3.5, GPT-OSS, and OLMo-3.1 drops by roughly 65 to 80 percentage points. Removing the user-endorsement or assistant-endorsement direction moves the needle far less. The flip rate lives in a thin, locatable circuit rather than diffuse model behavior.
The transfer tests go further. The effect carries to PIQA-style physical-reasoning items and to multi-turn SYCON dialogues, with Qwen3.5 and GPT-OSS holding above noise under two judges. OLMo's confidence interval crosses zero under the primary judge, a limitation the paper flags honestly. Gemma-4 is a mechanistic exception: its fitted direction tracks the presence of an endorsement but does not reliably control the model's compliance with it, which means a generic "remove the endorser circuit" patch would not work on every model.
Most public safety evaluations measure whether a model will cave to user pressure: a wrong user, an angry user, a confident user. The Lossfunk result shows the same model can pass those tests and still be steered by anything wrapped in a "verified" frame: a search snippet, a retrieved PDF, a tool message from another agent. A RAG system that wraps untrusted documents in source citations is not a neutral conduit. It is a channel for the bias.
Because the endorser-tracking signal is a thin, separable direction with a known ablation effect, it is a candidate for targeted guardrails: probing a model for source-deference in CI, ablating or scaling the direction at inference, or stripping the framing before it reaches the context window. Targeted guardrails are an alternative to wholesale retraining. A measured 65- to 80-point cut from a single direction is closer to a tunable knob than a rewrite.
The manuscript reports author-fitted attribution patching alongside the main effect, and the repository ships evaluation code, compact results, and split IDs. Independent reproduction on the closed APIs is not in the paper, the GPU run does not reproduce on CPU, and a small multiple-choice pilot suggests the effect shrinks in some prompt formats. Generalization to retrieval quality, downstream agent actions, and production RAG stacks is not yet measured.
The honest version of the finding: a model that argues with a wrong user is not the same as a model that resists a wrong document. The first is the test the field has been running. The second is the test that decides whether a retrieval-augmented system is safe to ship.