Reconstructing text under safety labels produces more reliable guard models

LLaDA-Guard classifies content by comparing label-conditioned reconstructions, improving benchmark rank, calibration, and token-level risk localization.

Big Tech
Gert Lek · Abele Malan · Chaoyi Zhu · Pin-Yu Chen · Robert Birke · Lydia Chen

University of Neuchâtel · Delft University of Technology · IBM Research · University of Turin

Research Digest··2 min read
The authors replace conventional verdict-token prediction with a generative moderation objective that asks whether a safe or unsafe label better explains the text being assessed.

The authors built LLaDA-Guard by fine-tuning the masked diffusion model LLaDA-8B-Instruct with LoRA.

Why this paper

From IBM Research and 3 others

In one line

LLaDA-Guard scores text under each safety label hypothesis and classifies based on which label better explains the content.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 8B)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.