Nvidia paper finds AI models lose accuracy on long tasks, with steep drops at scale

3 days ago 14



AI agents are increasingly asked to handle long, multi-step jobs. New research from Nvidia suggests they get worse the longer those jobs run. In a paper released on September 30, 2026, Nvidia researchers found that average model accuracy fell by 62.8% as context length scaled up. What Nvidia actually tested The paper is titled “Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability” and is listed as arXiv:2609.38712. It introduces a diagnostic benchmark called Long-Transduction. The goal was narrow on purpose. The benchmark tries to isolate one specific failure: keeping track of state while executing a task over a long stretch of input. The researchers evaluated seven open-weight models, meaning models whose trained parameters are publicly available. The tasks were deliberately mundane: arithmetic, sorting UUIDs, looking up variables, and transforming tables. The test was whether models could keep doing them correctly as the input stretched from 4,000 tokens to 128,000 tokens. Each model processed 1,440 documents. The team used greedy sampling, which means the model always picks its single most likely next output, and exact-match scoring, which gives no partial ...

Read Entire Article