safety alignment
- Papers
- 15
- Released code
- 1
- First seen
- Sept 2026
- Latest
- Oct 2026
15 papers in the last two months, against 0 in the two before.
Who is working on it
The papers
Most central to this idea first, not most recent.
- Research Labcs.AI
Learned latent links can undermine safety in multi-agent systems
CISPA Helmholtz Center for Information Security · Oct 2026
- Research Labcs.AI
TRACE reduces multi-turn attack success while largely preserving model utility
PRADA Lab, King Abdullah University of Science and Technology · Oct 2026
- Big Techcs.LG
Lightweight constitutional adapters improve jailbreak defenses without adversarial training
Anthropic Fellows Program, Independent · Sept 2026
- Chinese Techcs.AI
Safety training works better when models ignore superficial prompt cues
Zhejiang University, Ant Group · Sept 2026
- Chinese Techcs.AI
Selective distillation improves safety alignment while halving rollout compute and preserving reasoning.
Zhejiang University, Ant Group · Sept 2026
- Research Labcs.AI
Learned latent links can undermine safety in multi-agent systems
CISPA Helmholtz Center for Information Security · Oct 2026
- Big Techcs.AIcode
Training on aligned data can cause misalignment in different contexts
University of Southern California, Carnegie Mellon University · Oct 2026
- Big Techcs.CL
Reconstructing text under safety labels produces more reliable guard models
University of Neuchâtel, Delft University of Technology · Sept 2026
- Big Techcs.AI
Safety instruction influence controlled via eigenvalue modulation in transformers
Google Research, Cambridge University · Sept 2026
- Industrycs.CL
Some harmful fine-tuning examples drive far more misalignment than others
EleutherAI · Sept 2026
- Big Techcs.CR
Sampled text can steer black-box models past safety safeguards
University of Southern California, Adobe · Sept 2026
- Chinese Techcs.CV
Inner pipeline safety checks match outer guardrails with fewer false positives
Alibaba AAIG · Sept 2026
- Big Techcs.LG
Reinforcement learning makes models’ admissions of failure hard to reproduce
Stanford University, Anthropic · Sept 2026
- Big Techcs.CL
Method to permanently erase unwanted LLM personas via weight edits
University of Macau, Mila – Québec AI Institute · Oct 2026
- Independentcs.CR
Kernel-level preemption could halt rogue agents before network escape
Independent Researcher · Sept 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.