The authors built the FIGS evaluation framework (Factual Integrity and Grounded Support).
New benchmark disentangles sycophancy from empathy in multi-turn LLM conversations
FIGS uses adaptive 10-turn dialogues to reveal that current models either drift into false agreement or overcorrect into cold detachment.
Research Lab
Sidharth Pulipaka · Ruta Binkyte · Ivaxi Sheth · Sahar Abdelnabi
DFKI · ELLIS Institute Tübingen · Max Planck Institute for Intelligent Systems · Tübingen AI Center
Research Digest··2 min read
The authors introduce FIGS, a benchmark with 500 multi-turn scenarios and an automated judge that separately scores sycophancy (factual yielding) and calibrated validation (appropriate empathy).
Why this paper
From DFKI and 3 others
In one line
Current models drift into sycophancy or cold detachment over multi-turn conversations; FIGS evaluates both factual firmness and calibrated empathy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§