New benchmark reveals AI agents follow complex instructions less than 30% of the time

1 hour ago 2



Ask a large language model to write a haiku and it’ll nail it. Ask it to follow a 1,700-word instruction manual with nearly a dozen constraints while using specific tools in a specific order, and things fall apart fast. A new benchmark called AGENTIF, published on May 22, 2025, puts a number on just how badly. Current advanced AI models achieve less than 30% perfect instruction following when tested against the kind of complex, real-world tasks that companies are actually trying to automate. What AGENTIF actually tested The benchmark was developed by researchers at Tsinghua University and Zhipu AI. It consists of 707 human-annotated instructions drawn from 50 task categories spanning industrial applications and open-source systems. Previous evaluation frameworks like IFEval typically used shorter, simpler prompts that don’t reflect the messy reality of deploying AI agents in production environments. AGENTIF’s instructions average 1,723 words each, and each instruction contains an average of 11.9 constraints that the model must satisfy simultaneously. The constraints themselves fall into three broad categories: formatting constraints (how the output should be structured), semantic c...

Read Entire Article