Microsoft study reveals long agent runs expose serious reliability issues

1 hour ago 1



Microsoft Research published a study using its DELEGATE-52 benchmark that quantifies something practitioners have long suspected: frontier AI models degrade significantly when tasked with extended, multi-step workflows. The average document fidelity loss hit approximately 25% after just 20 delegated iterations. Zoom out across all models and domains tested, and that degradation figure climbs to roughly 50%. The benchmark that broke the illusion DELEGATE-52 was designed to measure something most AI benchmarks conveniently ignore. Standard evaluations test models on short, isolated tasks. Microsoft’s benchmark instead evaluates how well models preserve document content across long chains of delegated work, simulating real-world scenarios where models must iteratively edit, transform, and reconstruct documents or code across 52 professional domains. The models tested include some of the most capable systems available: Gemini 3.1 Pro, Claude 4.6 Opus, and GPT 5.4. All of them stumbled. Catastrophic corruption, defined as fidelity scores dropping to 80% or lower, appeared in over 80% of model-domain combinations evaluated. One domain stood out as a rare bright spot. Python workflows mai...

Read Entire Article