Long context windows do not ensure reliable sustained execution

Across seven open-weight models, performance deteriorated independently with longer contexts, harder operations and less structured input formats.

Big Tech
Jeffrey Willette · Krishna C. Puvvada · Boris Ginsburg

NVIDIA

Research Digest··2 min read
Willette, Puvvada and Ginsburg introduce Long-Transduction, a diagnostic for whether models can repeatedly read state, transform it and emit correctly aligned outputs over long generations.

The authors designed four task families that require sustained, record-by-record processing: arithmetic, UUID sorting, variable lookup and CSV table transformation.

Why this paper

From NVIDIA

In one line

Long-Transduction reveals that models suffer 62.8% accuracy loss when context length scales to 128K tokens in long-horizon agentic workflows.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.