The authors designed a reward model with six components targeting structural validity, extraction accuracy, groundedness, coverage, over-generation, and span precision.
Reinforcement learning with verifiable rewards improves event extraction
EAGER combines six fine-grained reward signals with schema-contrastive advantage estimation, outperforming prompting, supervised fine-tuning, and prior RL baselines across seven datasets.
German Research Center for Artificial Intelligence (DFKI) · Carl von Ossietzky Universität Oldenburg
Why this paper
From German Research Center for Artificial Intelligence (DFKI) and Carl von Ossietzky Universität Oldenburg · Part of Credit Assignment in Agentic RL, now 30 papers
In one line
Reinforcement learning with fine-grained verifiable rewards and schema-contrastive advantage estimation improves generative event extraction across seven benchmark datasets.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.