AT&T routes 40% of AI workloads through open-weight models, cutting coding costs by 56%

1 week ago 18



AT&T is quietly pulling off one of the more aggressive enterprise AI pivots in recent memory. The telecom giant now routes roughly 40% of its internal AI requests through open-weight models, and it’s not slowing down. The company has set a target of 60-70% within the next year. The math behind the decision is hard to argue with. AT&T has slashed AI coding costs by 56% while absorbing only a 2% decline in output quality. For certain complex workloads, the savings are even more dramatic, with cost reductions hitting 80-90% compared to closed-model alternatives. The tokenomics of telecom-scale AI AT&T’s AI consumption has grown at a pace that would make any CFO nervous. The company now processes approximately 45 billion tokens per day, up from around 8 billion just a year ago. That’s roughly a 5.6x increase in twelve months. AT&T calls its approach “tokenomics,” and the logic is straightforward: not every AI request needs the most expensive model in the room. The company uses an intelligent routing system built on LiteLLM and a custom AI gateway to match tasks with the appropriate model. Simpler requests get directed to open-weight options like Nvidia Nemotron, Meta Ll...

Read Entire Article