The authors built LLaDA-Guard by fine-tuning the masked diffusion model LLaDA-8B-Instruct with LoRA.
Reconstructing text under safety labels produces more reliable guard models
LLaDA-Guard classifies content by comparing label-conditioned reconstructions, improving benchmark rank, calibration, and token-level risk localization.
Big Tech
Gert Lek · Abele Malan · Chaoyi Zhu · Pin-Yu Chen · Robert Birke · Lydia Chen
University of Neuchâtel · Delft University of Technology · IBM Research · University of Turin
Research Digest··2 min read
The authors replace conventional verdict-token prediction with a generative moderation objective that asks whether a safe or unsafe label better explains the text being assessed.
Why this paper
From IBM Research and 3 others
In one line
LLaDA-Guard scores text under each safety label hypothesis and classifies based on which label better explains the content.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 8B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§