Unverified state text can flip calibrated model decisions

JevAdvBench shows that appending an opinion to a typed request altered jev-1.13.0 decisions about as often as the strongest injected command.

Top University
Jianyi Hu · Hangtao Zhang · Yi Liu · Yeqi Zeng · Li Zeng · Xianlong Wang · +2 more

Institute of Information Engineering, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Huazhong University of Science and Technology · Griffith University · Changsha University of Science and Technology

Research Digest··2 min read
Hu and colleagues introduce an adversarial benchmark for reinforcement learning for calibrated decisions, whose models return probabilities, choices, or scores directly to software.

The authors built JevAdvBench around 812 typed questions spanning 66 scenarios.

Why this paper

From Institute of Information Engineering, Chinese Academy of Sciences and 5 others · Part of Agent Security & Attacks, now 32 papers

In one line

Unverified opinions in the state field flip 12.1% of RLCD model decisions, making state an untrusted input.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.