103 papers this week in Safety & security12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Safety & security research

Agents

Codetta protocol enables high-capacity, keyless, undetectable steganography for LLM agents

Qi Pang, Virginia Smith, and Wenting Zheng propose Codetta, a steganographic protocol that allows independently deployed LLM agents to communicate covertly without a pre-shared key. The protocol achieves high capacity while maintaining provable undetectability. On three agent workloads, Codetta transmits up to 94 times more bits per visible token than the prior state-of-the-art asymmetric protocol.

today
Reinforcement learning

Diffusion reward models capture multimodal preference that scalar scores miss

Wang et al. introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over p(r|x,y) instead of point prediction. A lightweight Diffusion Transformer denoises Gaussian noise into a reward vector with no parametric family imposed, and N samples at inference form an empirical reward distribution that can be aggregated into a scalar, variance, or quantiles. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, and recovers multimodal structure where conventional heads collapse to a single point.

today
Safety & security

Internal confidence probe eclipses verbalized scores in calibration for reasoning models

Xi et al. show that the confidence reasoning models verbalize is systematically overconfident and uncorrelated with actual correctness. By training a linear probe on hidden states between chain-of-thought and answer, they obtain a confidence signal with 5 to 38 times lower expected calibration error (ECE) than verbalized scores across four benchmarks. However, this probe performs no better than majority voting at selecting correct answers, indicating internal states capture "how certain" rather than "which answer." They introduce probe-guided self-distillation (Probe-SD): using the probe to relabel confidence in model-generated traces and fine-tuning the base model, achieving in-domain ECE of 0.024 (from 0.178) and out-of-domain ECE of 0.113 (from 0.542) on Qwen3-14B.

today
Language models

Early answer confidence reveals when language models take reasoning shortcuts

Zhaohan Zhang and colleagues propose ConfLens, a framework that monitors how an LLM's confidence in its final answer evolves during chain-of-thought reasoning. They find that shortcut reasoning consistently shows a pattern of premature confidence, where the model commits to an answer early. Their new metric, the Distributional Answer Commitment Score (DACS), detects this reliably across math and code tasks, improving shortcut detection F1 by over 4.3% over strong baselines.

today
Safety & security

Repurposing backdoor trigger mechanisms as probes for adversarial defense in CLIP models

Wang et al. propose BaP, a test-time adversarial defense for CLIP that implants a defender-controlled backdoor probe into an MLP layer. Adversarial inputs strongly activate the probe, enabling detection and subsequent rectification via a small learned perturbation. The method raises average robust accuracy from 1.0% to 52.3% on a suite of 16 benchmarks, with up to 5.7x inference speedup over prior test-time defenses.

today

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses