The authors designed four task families that require sustained, record-by-record processing: arithmetic, UUID sorting, variable lookup and CSV table transformation.
Long context windows do not ensure reliable sustained execution
Across seven open-weight models, performance deteriorated independently with longer contexts, harder operations and less structured input formats.
Big Tech
Jeffrey Willette · Krishna C. Puvvada · Boris Ginsburg
NVIDIA
Research Digest··2 min read
Willette, Puvvada and Ginsburg introduce Long-Transduction, a diagnostic for whether models can repeatedly read state, transform it and emit correctly aligned outputs over long generations.
Why this paper
From NVIDIA
In one line
Long-Transduction reveals that models suffer 62.8% accuracy loss when context length scales to 128K tokens in long-horizon agentic workflows.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§