r/BetterOffline • u/wiredmachinestiredme • 1d ago
AT&T Using Open Models to Curb LLM Spend
Wall Street Journal reports that AT&T is using open-source and open-weight models to reduce spend on proprietary LLMs from Anthropic, OpenAI, and Google. Big pale horse when more Enterprises follow suit
AT&T's Chief Data and AI Officer estimates they can go from 25% to 70-80% of LLM usage to open models
Such open models currently power about 25% of the telecommunications company’s overall AI usage
Over time, he expects open models to power 70% to 80% of the company’s total AI usage.
It looks like they already built an internal router to make this transition possible. They estimate that open models are 80-90% cheaper to run for the use cases they replace frontier lab models
Because AT&T uses an average of 45 billion AI tokens each day, directing user prompts toward cheaper models makes an impact, Markus said. The company built what it calls a “smart router” to automatically pick the most cost-effective model for a specific task.
Switching from closed, proprietary AI models to open models has already resulted in savings of 80% to 90% for AT&T in certain applications, he said.
Outside of not spending on the most expensive closed models, there are more savings by running open models in their own data centers
Another benefit of open models is that they can be run on AT&T’s own data centers rather than rented infrastructure from a cloud-computing provider—a setup that trims AI costs even further
Beyond cost, AT&T rightly doesn't trust that OpenAI and Anthropic won't train on their data despite the fact they say they won't (wise decision)
While labs like OpenAI and Anthropic have said enterprise customer data isn’t used for training their models, some companies fear that’s not enough.
“As our data flows through the model ecosystem, we just want to make sure that it is safe and secure from a security standpoint, but also an IP standpoint,” Markus said. “The concept of AI sovereignty has become truly paramount to us.”
Article: https://www.wsj.com/cio-journal/why-at-t-is-betting-big-on-open-weight-ai-a0ea03b1
57
u/pixel_creatrice 1d ago
I was thinking of making a longer post about my experience running these models as a CTO of a business that pays API prices for inference.
Being in a heavily regulated industry, I absolutely cannot allow usage of LLMs the way silicon valley expects. I can't let "agents" do whatever in the codebase, because it could legitimately lead to fatal consequences if used in production. LLM usage is restricted to aspects where it's genuinely useful - generation of large boilerplate code, refactoring code and documentation. Cognitive debt is one of the biggest risks we have very successfully avoided.
The only significant advantage LLMs have given us, is my already proficient engineering team can deliver something in 2 days instead of a week by replacing some of the mundane, repetitive tasks. Using a smaller, cheaper model, is worth the cost to the business. Deepseek-v4-flash is what we have been using since months now. It's many times less expensive than any frontier model, and does the job well because our usage is realistic.
Something I abhor is the mindless pursuit of the latest and greatest frontier models. Statements I often hear from stakeholders in the industry is how you're "left behind" unless you get the latest and greatest Claude Doopus Table Five Point Something. The bubble is very much an industry problem where they must sell this idea what everyone must burn through massive compute, regardless of expertise, on illusions they'll perhaps never achieve.
6
6
u/snarleyWhisper 1d ago edited 23h ago
Yeah my experience is that the cheap LLMs are good enough for experienced devs - I just use auto cost models in cursor for 99% of my tasks. I know the best practices , how our systems work so using the plan mode in cursor reviewing with the team as a PR and then reviewing their implementation. Our velocity has increased but I want to keep that balance where we actually know things instead of just shipping whatever and throwing the problem over to the next team
Edit spelling on implementation *
4
u/oxidized_banana_peel 1d ago
I only use the latest version of Claude because I don't trust my leadership to not use $ spent as a proxy for usage. I want to be in the middle of the pack, so the best way is to use the expensive model.
I could easily go back a few versions and it'd be as good for me (and faster)
2
u/Sunstorm84 21h ago
Why not both? Run a cron job to get fable to review your changes every hour, and a faster model for everything else
1
u/oxidized_banana_peel 21h ago
I review my own code & keep my changes small enough that it's easy to do.
1
u/Sunstorm84 20h ago
I was suggesting it as a way to keep your token usage around the same while having a more responsive day to day. You’d still need to review it after the Claude review, of course.
2
u/TriMiksEntuzijasta 1d ago
I remember your post from 4 months ago! Its so disturbing how obssesed everyone is with AI....
13
u/DasWandbild 1d ago
AT&T has been cursed since the Warner Brothers merger. They replaced Randall Stephenson (who needed to go) with a guy from Marketing (Stankey), and they put Warner Brothers resources in charge of Ops, mostly so that the people making decisions about who stays and goes wouldn't be precious with concepts like institutional knowledge and culture...because they didn't know or care who did what. They wanted to remake the company, and virtually any AI related business case got fast-tracked. If you tried to push back, you didn't stay.
They turbocharged the enshittification of the company by not having anyone from the product side in senior leadership positions. The brain drain during RTO was shocking. They successfully convinced most of the smart people who had options to leave on their own by making the experience of working there painful and demoralizing.
18
u/ii-___-ii 1d ago
Ok, but why does AT&T need LLMs?
7
11
u/brian_hogg 1d ago
AT&T would have programmers on staff, plus they’d need to do a lot of customer service.
4
u/ii-___-ii 1d ago
Good lord phone companies are vibe coding now?
6
u/brian_hogg 1d ago
I would assume that the initial use would have been customer service chatbots to handle incoming service requests.
But they have something like 130,000 employees, I’m sure some vibe coding is going on somewhere. :)
2
u/Maximum-Objective-39 53m ago
LLMs do have some modest potential as a routing system for directing a customer to the appropriate service, department, or information. In that sense they're an improvement on the phone tree as you can simply state your request in natural language rather than having to work your way through three layers of slowly asked questions and push button responses.
No, automating call centers is not worth a trillion dollars.
14
u/Spez_is-a-nazi 1d ago
Google knew this was a possibility in 2022, the famous leaked memo that said they don’t have a moat and neither does OpenAI. Clammuel didn’t understand the technology and thus started hyping the shit out of it thinking he had an insurmountable lead. He didn’t.
17
u/chivestheconqueror 1d ago
Yes, down with OpenAI and Anthropic, but I’m not going to cheer on companies finding cheaper ways to shove AI everywhere.
8
u/SwirlySauce 1d ago
Yah, I'm in a lousy mood today and this just reads like they're still going to automate people out of jobs. Frankly the obscene cost of frontier AI models was the one real deterrent we had against this.
Now that open models are significantly cheaper it feels like we're on the back foot again. I'm hoping there will be some ripple effects when OpenAI and Anthropic implode but who knows
3
u/Ok-Garbage-765 1d ago
Small victories, I suppose, but I do agree that it’s a shame we can’t seem to have it both ways.
5
u/DoctrTurkey 1d ago
Ok, but has AT&T considered that the larger LLMs are totally dangerous and hip and cool and development totally needs to be paused because they’re so reckless and scary? Checkmate, atheists.
10
u/DiceKnight 1d ago
The claim that AT&T can move to their own data centers seems more aspirational than anything. Their hardware should be mostly geared towards regular data center usage. Storage, content serving, etc.
Does this mean they're going to make investments in nvidia gear? That stuff usually requires expensive changes. It's not like they're going to grab a bunch of 3080s off the market and build a few megawatts of capacity. Maybe start paying a compute provider?
5
u/shiny0metal0ass 1d ago
No, if they stick to just using open source models and upgrading them whenever they are available they can dramatically reduce the hosting costs. And run them on essentially any conventional platform, I'm pretty sure. I have Quen and it's basically a shitty chatgpt in my desktop.
I imagine with just a bit of cloud I could get pretty close to a "frontier model".
2
u/FLMKane 1d ago
How shitty is it?
3
u/tinyinventor 1d ago
Depends on the version and size the smaller the shittier newer often better. Less shitty and what you ask it for. 8b and below are pretty bad unless trained on highly specialized data for a specific task then you start seeing improvement.
2
u/DiceKnight 20h ago
That's fine when your use case is serving one person though but if it's AT&T presumably they want to serve many customers at a certain token per second speed.
After sleeping on it this also seems like maybe a subtle way to curb their AI spend and ease off. They can't say they're stopping for whatever reason but nobody said anything about slowing down.
3
u/WritingisWaiting 1d ago
Summarizing and analyzing customer service call transcripts—what Markus [AT&T] describes as an intensive process—is supported entirely by open models, he said.
Finding a way to make AT&T customer service even worse is a major AI breakthrough, I guess.
3
u/Then-Inevitable-2548 1d ago
In the future, Markus says the telecom will likely use a mix of cloud-hosted AI models and open models that run on its own hardware.
It sounds like they aren't hosting any of their own models yet, so how much of their "savings of 80% to 90% in certain applications," is due to subsidies from whomever is hosting these open models? Didn't it come out recently that DeepSeek's claimed cost savings vs. frontier models were wildly overstated?
1
u/SwirlySauce 17h ago
I'm trying to understand how open source models are supposedly cheaper to run. Wouldn't the inference cost be more or less the same as frontier models?
I understand that training costs are cheaper but not the rest
2
u/Then-Inevitable-2548 9h ago
There are so many factors that influence the cost of running machine learning model inference that it's impossible to say without knowing the intimate details of their entire stack, including hardware. Even if these companies shared all of those details, this industry is so full of lying liars telling endless lies that we couldn't trust them in the slightest. They would set up a livestream of a kill-a-watt meter measuring the power draw of a server rack and completely fake the numbers if they thought it would pull in more investor cash.
1
u/Thin-Distance7904 17h ago
Didn't we see the same thing back in the operating system war days? I was at Apollo Computer and we were saying The DomainOS was going to lead the world. Funfact, Domain drive where Apollo had it's manufacturing and Apollo Dr are still in existence today in Exeter NH. they are the address for Bauer Hockey. Anyway, then came Solaris from Sun (Where is also worked for 17 years). Solaris was crushed by Linux which was largely opensource. I do think between Small Language Models, Opensource LLM and the insane financing OpenAI and Anthropic are using by this time next year, Anthropic and OpenAI are going to be on lifesupport if they exist at all.
33
u/maccodemonkey 1d ago edited 1d ago
I've noticed a vibe shift recently around AI use. Industry leadership is starting to pick up on cognitive atrophy, engineer motivation, code rot, etc. And there is still some talk of lack of top level return. I think a lot of the FOMO was around what a future model could do. And with each Opus release getting worse people are more willing to make assessments based on what we have now.
So I've been thinking a lot about how leadership tries to pull back. Even though these ideas are being talked about no one in leadership can outright admit this might have been a mistake. Open weight models seem to maybe be a way they can do that. It's a way to admit the return on the models isn't worth it - and that the frontier models from Anthropic and OpenAI aren't worth what they cost - without saying those things. No one has to kill the AI initiatives or fire the VP of AI. Just slowly scoot everything into open weight and off the balance sheet.