Polished evidence makes LLM agents act on unknowable questions

Across 12 frontier models, authoritative-looking displays—not their factual content—substantially increased decisions on provably unpredictable questions.

PaperIndependentcs.AIarXiv:2608.27167v1
Pranav Aggarwal
Research Digest··2 min read
Aggarwal tests whether LLM agents abstain when asked to make directional calls that cannot be inferred from the available information. Commitment rose from 6.5% to 54.0% as apparent evidence was added, while targeted fine-tuning eliminated commitment in the tested cases but proved sensitive to response format.

What they did

The study presented 12 frontier models with provably unpredictable questions, either alone or alongside increasingly elaborate market panels. It compared genuine panels with displays whose numbers were entirely fabricated, and used matched answerable questions to verify that the models could interpret the same materials when an answer was available.

The author also asked models to state probabilities and separately classify whether questions were knowable before acting. A 3B model was then supervised fine-tuned on 540 synthetic abstention cases—mostly involving dice, coins, jars, and timers—and evaluated on the original tasks and three unseen domains.

Key findings

  • Commitment on unknowable questions increased from 6.5% to 54.0% as apparent evidence was escalated.
  • Fully fabricated panels raised commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% elicited by genuine market data.
  • Stated probabilities changed little despite the 48-point swing in action and performed worse than a fixed base-rate, or climatological, baseline. When models first labeled questions irreducibly unknowable, they subsequently committed in only 0.4% of those cases.
  • Fine-tuning drove commitment to 0.0% on the original cases and transferred to three unseen domains, but the behavior failed under rigid response formats that left no room for reasoning.

Why it matters

The results distinguish recognizing uncertainty from acting appropriately under it. They suggest that some agents possess the relevant knowledge and can identify unknowability, yet still lack a reliable abstention gate when information is packaged with institutional or quantitative authority—a potentially trainable but deployment-critical failure mode.

Caveats

The effect was concentrated in a few models rather than universal, and the supplied abstract does not report model-level results or full dataset sizes. The fine-tuning experiment used one 3B model and only 540 synthetic examples; its apparent transfer remained context-fragile and broke under constrained output formats.

§

Analysis

The paper reframes calibration as two separate problems: estimating uncertainty and converting that estimate into an action or abstention decision. Its evidence suggests that improving probability estimates alone may not prevent agents from acting on unsupported claims.

The format dependence also complicates evaluation: an abstention policy that works in free-form reasoning may fail when deployed through schemas, APIs, or interfaces requiring a fixed answer. Robustness therefore needs to be tested across the actual response constraints used in production.

newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.