Meta introduces GAMUT benchmark to measure factual completeness in AI

2 weeks ago 16



Meta AI researchers have released GAMUT, a benchmark designed to measure something most AI evaluations quietly ignore: whether an AI’s answer is actually complete. Not just accurate, not just fluent, but whether it includes all the facts that matter. The benchmark, short for Grounded Assessment of Multimodal Factuality, was published as an arXiv paper on July 21, 2026. It introduces a structured rubric system that converts “did the AI say everything it should have” into binary, machine-gradable checklists. What GAMUT actually does GAMUT tackles the problem of omission with a two-level meta-rubric framework. The system organizes required content hierarchically, meaning it doesn’t just list facts an answer should include. It structures them by importance and category, then translates those requirements into yes-or-no questions that language models can reliably grade. The benchmark contains 1,813 questions rooted in real wearable imagery across ten distinct domains. Each question comes paired with expert-verified rubrics. Human experts defined what a complete answer looks like before any AI was tested against it. The results are humbling Meta evaluated 14 different AI models against t...

Read Entire Article