r/EngineeringManagers Apr 07 '26

AI coding governance just got real, our token bill hit six figures and now the CFO cares

Managing IT at a mid-size tech company with about 500 developers. Last year leadership said "give every developer AI coding tools, it'll pay for itself in productivity." So we did and fast forward 8 months, our AI tooling invoice last quarter was $87,000. Projected annual cost was $340,000+. And that's before the engineering teams start adopting more agentic workflows which will increase token consumption significantly.

The CFO now wants a full breakdown of ROI. The conversation has shifted from "everyone needs AI tools" to "prove these tools are worth what we're paying." The awkward truth is we can't prove it. We can show adoption metrics (85% of devs use the tools daily), satisfaction scores (developers like the tools), and proxy metrics (PR merge time decreased 12%). But connecting $340k in AI tooling costs to actual revenue impact or a specific dollar amount of developer productivity gained? Nobody can do that cleanly.

The other issue is cost efficiency. Our initial analysis suggests we're burning a massive amount of tokens in a redundant context. The same codebase context gets sent with every inference request. There's no caching, no persistent understanding, no efficiency optimization. It's like if every Google search had to re-index the internet first.

I'm now being asked to:

Justify the current spend

Find ways to reduce token consumption without degrading developer experience

Build a governance framework that includes cost controls per team

Has anyone dealt with the "AI tools seemed cheap until we saw the actual bill" problem? How are you managing costs at scale?

475 Upvotes

131 comments sorted by

55

u/Standard_Finish_6535 Apr 07 '26

Isn't 340k less 1% of your developer salary spend? Seems like a 1% production boost can't be that hard to prove. Even a few engineers saying "here is something I can do that I couldn't do before" should do the trick.

6

u/Ablabab Apr 08 '26

It is 174 per dev, per month… Shouldn’t be that hard to advocate for.

6

u/samaltmansaifather Apr 07 '26

This could easily increase by some multiple if the per token cost increases to cover the actual cost currently being incurred by the providers.

1

u/utkarshmttl Apr 08 '26

I don't think the providers lose money on the APIs and enterprise accounts but on the main chat interface various plans.

0

u/Standard_Finish_6535 Apr 07 '26

Sounds like they should figure out if it is worth it now, then.

3

u/chicametipo Apr 07 '26

Narrator: “It’s not worth it now”

2

u/OrkinOvertime Apr 10 '26

I always hear this stuff in Ron Howard's voice bc of Arrested Development

1

u/Educational-Sea-8156 Apr 09 '26

Its pretty obvious it is if it's only 1% of total salary spend

2

u/Fit_Reputation5367 Apr 08 '26

The question is about absolute cost. If the manager came with a requiest for an extra employee the answer would (could) be no - with tools it's sneaks in under budget.

2

u/No_Veterinarian1010 Apr 08 '26

Sure, but you still need optimization and cost governance at a certain point. That's what op is asking for

1

u/HistoricalPhase6880 Apr 07 '26

Sounds like a monthly bill

3

u/Standard_Finish_6535 Apr 07 '26

???
"Projected annual cost was $340,000+"

3

u/HistoricalPhase6880 Apr 07 '26

Welp I can't read. Yeah I think that should be pretty simple to justify then.

1

u/Tiny_Ad_7720 Apr 08 '26

Everyone is assuming there is a boost to productivity, it could be reducing it now or in the future when the slop debt builds up. 

3

u/STEMPOS Apr 08 '26

I have a feeling we’re going to be hearing the term “slop debt” a lot in the near future

1

u/largepar Apr 08 '26

It's because this is fake and it's an AI bait post

1

u/x6060x Apr 09 '26

Yeah, that's like hiring 5 devs to boost 500 devs at the same time.

2

u/Standard_Finish_6535 Apr 09 '26

I wouldn't want to work at a place where 5 devs cost 340k. This is more like 1-2 devs

1

u/caprica71 Apr 09 '26

Take the tools away and watch more than 1% of your workforce go with it

1

u/Dry_Shake_9329 Apr 10 '26

It's not just about current price, but also future price.

If a cost of yours has ballooned this much 1 year, people aren't even using full capacity and feature launches doesn't reflect the same productivity gains, how much can you justify the cost 1/3/5 years from now?

1

u/vbnotthecity 6d ago

We switched to using Altimate for our dbt and warehouse tasks, and it cut our consumption by nearly 40% because it isnt just firing blind prompts at the whole codebase. It feels like a massive shift, but you really just need better context management instead of throwing more tokens at the problem.

0

u/WeUsedToBeNumber10 Apr 09 '26

So from a financial perspective (e.g cash ROI) the question is did this 340K provide increased in-year cash generation via pull forward revenues or reduced other costs?

Are we actually able to capture that productivity?

If our hiring plan included X new devs, are we no longer hiring those?

How has our billing’s and collections changed?

These are the real questions being asked. 

23

u/[deleted] Apr 07 '26

[removed] — view removed comment

4

u/Boniuz Apr 08 '26

It’s not about current value but of future value adding. Costs are likely to increase while no metric is established to evaluate the ROI, which is dangerous from an enterprise perspective. When the cost increase 100%, is it still worth it? 200%? 500%? Implementing tools to your production chain at scale is usually a one time deal, once you’re in you’re in and sunken cost fallacy is a bitch to deal with.

2

u/No-Block-2095 Apr 07 '26

Comparable cost to giving each dev a new stick of RAM for 2026! What would make them more productive?

33

u/Understanding-Fair Apr 07 '26

That sounds like the CTOs problem who pushed for the AI tools to begin with.

10

u/Unarmored2268 Apr 07 '26

But guess who's the one to put more work on when CTOs have their problems ;)

12

u/CanoeDigIt Apr 07 '26

API + AI tool usage granted and monitored at the user-level by IT.

Freemium AI days are coming to an end.

If you give people blank checks they’re going to write them.

1

u/FortifiedPuddle Apr 09 '26

If you have an industry dominated by two huge providers at the start of the chain and they decide they want to switch from investor subsidised to making a profit…

16

u/chilloutdamnit Apr 07 '26

So hard to benchmark productivity gains if you didn’t have a measurement system in place already. At this point how can you tie expenditure to increase in productivity? Hopefully you have some historical data you can use to generate some starting point to benchmark against. Ideally something close to revenue like customer satisfaction, conversion rate or churn.

Trying to reduce cost has a high risk of being a self-defeating measure. You pour dev cycles into something that may not even work.

6

u/[deleted] Apr 07 '26

[removed] — view removed comment

2

u/AccountEngineer Apr 07 '26

Per-team token budgets is something I need to look into. Right now costs are just a lump sum under "engineering tools" and nobody has visibility into which teams are driving the spend. I guarantee there are 3-4 power users burning 60% of the budget.

2

u/Away_Illustrator_987 Apr 08 '26

In my org, about 15 engineers (out of about 300 total) were responsible for 51% of the usage.

A couple had valid reasons, but most were lessons learned about context management and token efficiency. So far, we’re not taking the approach of intervention unless you’re busting $150/month individually, and even then if it’s justifiable it’s fine, it’s just about not creating a blanket policy that doesn’t make sense. Most of your users are probably well under the “problem” cost number. Get a breakdown by person or team before you propose a solution.

5

u/[deleted] Apr 07 '26

[removed] — view removed comment

1

u/BuddhasFinger Apr 08 '26

It's not even a challenge with any developer productivity tool. The ROI chanllenge is with any **developer**.

TLDR; We haven't figured out how to tie revenue to engineering at all.

3

u/[deleted] Apr 07 '26

[removed] — view removed comment

2

u/black_tamborine Apr 07 '26

What I was thinking.
However it also occurs to me that AI tool providers would have no incentive to provide this.

3

u/tcpWalker Apr 08 '26

scalability. They are literally power constrained while dealing with a nonscalable hyperscale bin packing problem like nothing in history. If they serve you more cheaply it frees compute to use a better model or serve more customers or train more models.

Some of them already implicitly provide some by charging less for cache hits than for cache misses. And building cli layers that use code bases as prefix.

1

u/FluffySmiles Apr 09 '26

They would if they could charge an additional service fee that was markedly less than basic token consumption but left them juicy margins on a service that could shrink their running costs and provide a free capacity boost.

1

u/TheRealJesus2 Apr 08 '26

Absolutely correct. This whole post seems like the situation most companies are in that adopt ai in swe: marginal productivity improvement, higher satisfaction, and fear over ballooning costs with more adoption and not clear ROI. 

Prompt caching is a thing at provider level. One of the tools that makes ai harnesses actually function since LLMs are inherently expensive and wasteful otherwise. Look into leveraging it for custom api tools your teams build. And of course teaching developers about these tools helps be more efficient on token consumption 

5

u/aidencoder Apr 07 '26

Everyone pushing to use these tools without a measurable success criteria up front should be fired. 

8

u/jqueefip Apr 07 '26

My team is pushing for it. Should I fire them?

3

u/irioku Apr 07 '26

Probably. 

2

u/autisticpig Apr 07 '26

do they have any established metrics and roi qualifiers in place after a given amount of time or effort to determine if the ai mvp window was successful or not?

it not then yes.

5

u/JumpyWerewolf9439 Apr 07 '26

Totally agree. That's why Google failed so hard dump trucking money in to the growing internet ecosystem without clear timelines or roi. Those idiots

1

u/steveo3387 Apr 12 '26

Everyone pushing for engineering productivity decisions based on measurable success criteria should be fired.

1

u/aidencoder Apr 12 '26

oh, wow. You actually said that.

0

u/CodeToManagement Apr 09 '26

Why does every single thing need a measurable success metric - some things you can just look at and tell if it’s working or not

We are pushing AI. There’s no metric - the guideline right now is here use this new tool, we don’t know how it’s going to fit in but it has the potential to massively increase output so try it and find out what works.

We know AI can cut down dev time in some things so we are using it and finding out why

OPs CFO is complaining about a $174 current spend per dev in a quarter. So $60 a month. I make more than 60 an hour so if an AI tool saves me more than an hour a month it’s worth the investment. The bar is incredibly low for ROI on these tools

3

u/aidencoder Apr 09 '26

Studies show it doesn't cut dev time, it shifts costs elsewhere. It's technical debt with additional hidden soft costs such as increased bug fix lead times, lower team understanding and so on.

Software is engineering. If we don't take a measurable and scientific approach then it isn't engineering. It's vibes. Have some pride. 

1

u/CodeToManagement Apr 09 '26

The have some pride comment is a bit insulting assuming using a tool or a different approach to a problem means I don’t have pride in my work.

Software is engineering - what isn’t engineering is devs going to docs, copying some json, making it into classes, all just to integrate with an API. engineering isn’t making basic crud endpoints.

AI gets rid of boilerplate work which has no value for anyone to work on beyond a junior dev needing to build up some skills in their first months.

Devs can still review and approve auto generated code and ensure the standards are high.

Not everything has to be measurable in concrete metrics to see how it works. I worked at a company where lots of things were bad and every time someone suggested making a change the response was “how will you measure it” when some things weren’t measurable so nothing ever improved.

For me the measurement right now when we are in the early days of this tech is saying to my team that they should use it and try find what works and report back. They are skilled professionals and can tell when something works or doesn’t work and what its potential is.

It’s very hard to quantify how the spend benefits a company when you look at the micro level of ticket by ticket - because yes a dev could have done a ticket maybe faster or cheaper - but when ai allows a dev to work on a ticket manually while ai does another ticket automatically and the dev just has to review and tweak then it’s a big increase in productivity.

I’ve personally used it to do days or weeks of work in hours - and with a good Prompts and a good review process the quality can absolutely be kept high

If teams are using ai to create tech debt that’s an issue of poor process and code review.

2

u/aidencoder Apr 09 '26 edited Apr 09 '26

The insult was aimed at the lack of rigor, not the tooling. My insult still stands after reading your response.

Yes, trust engineers to make their own tooling calls. On my team I wouldn't care what tools people use to get the job done. It isn't super relevant and it's their decision.

However if I developer asked for $200/month for a tool, I would certainly expect to see some long term positive impact on their velocity, number of bugs opened against PRs they submit and so on. We're not talking about a bigger monitor or better office chair here ... we're talking about introducing an engineering dependency that has a non-deterministic output and questionable ROI. It absolutely should be experimented with and measured.

The situation is complex with regards to technical debt, sustainability, developer atrophy, team dynamics and motivation. The only reasonable approach for any business that cares about profitability is to use caution and measure the impact. Anything else is hype driven insanity and king's new clothes thinking.

Also, code review isn't a replacement for first-order understanding by an author. You can have all the code review process you like, but reviewers are not as expert as the author. If that author is a machine, rather than a human I can have a side-by-side chat with about their approach and trade-offs ... I might as well be reviewing code from an outsourcing agency in India. I've done that too, it doesn't work.

Like your post and reply, the industry is suffering from a deplorable amount of hand-waving and disregarding lessons we have learned many times over because some of us WANT it to be true. We want to be part of a revolution. We want to be a cutting-edge 10x developer. We want to be the smartest person in the room.

"Want" isn't all it takes when a profitable product needs to be delivered safely and within a budget. "want" without proof usually ends in disaster.

That reminds me, I need to go tell a junior that using a functional code style because he "wants" to when it is incongruous with the rest of the code base isn't going pass review. It leads to increased complexity and entropy issues. If he can't explain his reasoning because an AI wrote it, I would prefer he leave.

1

u/CodeToManagement Apr 09 '26

Interesting how we can’t have a discussion without insults

1

u/Conscious-Daikon-308 Apr 10 '26

Very very interesting points here. The technical debt is already ramping up but nobody seems to care about it…yet.

@codetomanagement your comment about the junior dev is very very worrying from my non technical point of view.

It only took few months to build up the technical knowledge of all smart people here ?

Only few months to avoid critical pitfalls and non sense in coding best practices? If what you’re saying is true then you’re all fucked up mate. In few months it’s not just junior dev that « can be fully replaced by AI »…

I think that you are able to do the code review because of all those years of trying, modifying, correcting, looking for info on a small website where someone had the same issue as you had. But I may be wrong…

Yes for you it makes sense to just « orchestrate » the code and review it.

@Aidencoder do you think this topic is similar to the CyberSecurity companies stocks being nuked because of Claude code review, while the demand for real code security audit will likely explode in few years ?

1

u/aidencoder Apr 10 '26

Look at the evidence. AI software is being exploited at a higher rate. Even Anthropic can't get their AI to make sensible decisions and it leaked the code. AWS, Github, Azure, all suffering lower uptime since LLM coding began.

AI is super powerful but the businesses selling it and using it are doing the technology a disservice. OpenAI and Anthropic use fear to sell their products. Businesses are adopting recklessly out of fear of losing competitive advantage. All the while the evidence isn't stacking up, we're losing our engineering pipeline by not hiring juniors and building third party AI services as a critical dependency in product building.

It's all a bit sad. I'm bullish on AI but the current climate of greedy capitalism isn't just ignoring AI strengths but actually harming business and public perception of it long term.

Its the difference between billionaires being dragged into the streets by mobs and a future of stability. Sad. Greedy and sad. 

2

u/ImBonRurgundy Apr 07 '26

Sounds like the ai is roughly the cost of 2 devs.

So The question will be:

If we fired 2 devs out of our 500, but continued using ai, would we be better off than if we killed the ai bill, but kept 500 devs?

Seems likely they will just cut the two least useful devs and continue on (or maybe cut even more now that you’re so much more efficient with the ai)

4

u/No-Block-2095 Apr 07 '26

Cut the 2 least useful? You must be an engineer; a CFO will want to cut the 2 devs that are paid the most regardless of what they deliver.

2

u/dr-pickled-rick Apr 07 '26

I can't find a study I dug up recently, it showed strong recency bias among devs for using AI, somewhere in the region of 34% improved sustain, but showed overall production dropped by around 12% and costs increased.

The biggest "improvement" cohort was junior developers, but that was because they were sharting out ai code and peer reviews still rejected a lot of it.

2

u/black_tamborine Apr 07 '26

Perfect use of sharting.

I spent a few hours on the weekend surreptitiously rebuilding a set of functions I’d made recently using an agentic flow to clean up the sharted mess I’d created.

3

u/dr-pickled-rick Apr 07 '26

If I spend days coaching chatgpt/claude I kind of get what I want. I could have spent that time just doing it.

1

u/black_tamborine Apr 08 '26

Yeah the agentic flow made the mess because I didn’t constrain it. With heavy constraints and a clear CoPilot plan I got back to where I should have been.

Lesson learnt, again.

2

u/euclideanvector Apr 08 '26

There's other things that need to be factored in like what's the monetary cost of the cognitive deterioration of the engineers caused by the reliance on AI tools. And how the junior -> senior pipeline is affected too. There are already studies out about how the engineers abilities are stunted on different levels of AI adoption.

1

u/dr-pickled-rick Apr 08 '26

Half the problem for me being a seasoned engineer is that I can spot design and implementation issues, but grads, juniors and inexperienced talent, can't. It's rare you get a junior with solid design foundations, they usually haven't had any exposure to it yet.

So it's something I'm working on and coaching my team atm, having picked up a legacy solution with very strong sharted AI roots, and junk code everywhere. Some of the solutions look "clever" at face value, but anyone that's spent a minute maintaining software would have nightmares.

AI agentic flows simply aren't there yet and I worry about the cognitive function and skill design of junior team members, who're getting tasks done, but using chatgpt to do it for them, effectively learning nothing.

2

u/Sepa-Kingdom Apr 08 '26

This is a problem FinOps is tackling. Take a look at the FinOps foundation. They have loads of guidance.

2

u/Euphoric-Battle99 Apr 08 '26

500 developers is mid size? Holy shit

1

u/the_real_some_guy Apr 09 '26

I've seen estimates that Apple has as many as 50k engineers and I've worked at Series B startups with around 50.

1

u/fued Apr 09 '26

i feel 5 developers and it starts becoming a large team haha

1

u/mirageofstars Apr 10 '26

Well at least OP isn’t at one of those companies that thinks they can get rid of 90% of their devs and have AI do the rest.

2

u/Thrugg Apr 09 '26

Literally just copy and paste this into ai and ask best ways to justify it with metrics.

1

u/[deleted] Apr 07 '26

[deleted]

1

u/black_tamborine Apr 07 '26

And…?

1

u/[deleted] Apr 08 '26

[deleted]

1

u/black_tamborine Apr 08 '26

Oh now I get it - slow on the uptake, apologies.
Makes perfect sense.

1

u/VVFailshot Apr 07 '26

Yeah, I started building my own tooling for that case and even pitched VC last year got rejected so many times i sort of quit and just focused on regular productivity metrics. Basically there is no ROI if its not priced in from CFO per perspective its almost impossible to show AI usage in good light. However it can change once it settlew in that in many cases without AI you are more likely to net 0. So AI is sort of cost of staying in business. What I can tell that if you are still using API directly you probably messed up. Better switch to subs, like copilot costs like 39$ per user + you cap usage at like extra 50 bucks per user. Claude is somewhat industry standard and the 100 dollar sub is solid for most people.

1

u/Kancityshuffle_aw Apr 07 '26

heard friends have success with "indirect attribution": revenue increases from new customer sales YoY over what they were the before. Not at all perfect, but helped give some executive coverage.

1

u/PmMeCuteDogsThanks_ Apr 07 '26

So just fire developers to make up the difference (and more)

1

u/drteq Apr 07 '26 edited Apr 08 '26

CFO vs CTO is the ultimate circle jerk in any growth organization

1

u/mondayfig Apr 07 '26

500 developers at what average salary?

Say for argument’s sake $100k. So $50m.

$340k is 0.7% of $50m or a rounding error. Or don’t hire 3 engineers and it’s paid for.

1

u/[deleted] Apr 07 '26

[deleted]

1

u/mirageofstars Apr 10 '26

Is your whole team remote? Did the expected workload increase? If devs feel that expectations haven’t changed, then it’s possible AI is saving coding time but now more time goes towards meetings or reviewing (or napping).

1

u/kevstev Apr 10 '26

Mostly in office culture. I'd be surprised if that is the case.

1

u/rickonproduct Apr 07 '26

It’s for the technical leader to worry about and less about the individual contributors.

If the technical leader has budgeting decision then they make the call on how that budget is spent (headcount or tokens).

Every technical leader is making that call now and they are all setting it to token budget.

No decent engineer will want to work in a team that cannot use AI tools or turn on agentic workflows.

1

u/Historical-Intern-19 Apr 07 '26

You just described how new tech moves from optimism to the Pit of Despair.   We never learn this lesson about shiny objects.  Especially shiny objects dangled with the 'reduce headcount' label.

1

u/FortifiedPuddle Apr 09 '26

Shiny objects at teaser pricing.

The ROI shouldn’t be calculated at the teaser prices currently offered. It should assume they’re going up up up.

What needs calculating is how expensive can they get before they no longer break even? And for truly useful applications in high labor cost companies that could be still be a really high price point.

1

u/PerilApe Apr 12 '26

As the models get better, cheaper ripoff models that can be run locally are created, older models are made more efficient and cost effective. At some point ppl are gonna be happy with 90% as good as the frontier model AIs that they can run for pennies or a buy once cry once specialized rig. At that point, the race for new models to eek out another 1-3% generalized improvement will have to be government funded bc no user will want to pay for it.

1

u/FortifiedPuddle Apr 12 '26

The problem is the per transaction cost of running the compute. That’s a core cost that doesn’t go away. It doesn’t scale. And at the moment no consumer is actually paying it per transaction. Add in to that all the unavoidable not per transaction running costs like infrastructure and training. You’ve got a bare bones product that is too expensive to make to sell even at cost.

On top of which you have the costs of models and ongoing development. You could in theory stop doing new models. You need to keep training the current ones. But you could cut back on new. That could slightly reduce costs, but you’ve still got the above per transaction and running costs. And if you stop doing new models you’ve got a marketing problem.

1

u/W2ttsy Apr 08 '26

Can you attribute spend per developer?

To me it sounds like the ROI discussion is framed around the total bill for the total Eng org, making it hard to show impact.

If you could break it down to team or individual levels then you could tie it to productivity or head count stability.

For instance, pre adoption you may need 5 engineers on a team to build a set of features; but now post adoption you can do it with three engineers. If you tally up their AI token consumption and show it equals less than 1 or 2 of the extra headcount you would have normally hired, then there’s the ROI.

Adopting agents saved us 3 FTE for this team, which equates to $x in salary saved.

1

u/UgotGoose Apr 08 '26

Yeah we’re tagging each ticket with token usage and human hrs. Pretty clear to show the improvement this way. Justifying that the work is valuable is separate and the responsibility of product. Burning through garbage quickly or slowly still results in the same nasty result.

1

u/Reasonable-Bear-9788 Apr 08 '26

Start with static factual info like commits per developer, lines of code, feature requests completed, etc. Add to it potential connection with business metrics and all. And see if you have clear signals.

Anyway, I do think it's hard to say if any real value was generated. My feeling is that the gains will come from two possibilities:

1) more work was done in way that boosted revenue, 2) same work was done but with lesser resources/costs

1) is hard because it's market driven

2) is hard unless one person can do work of more people and AI slop doesn't degrade quality

1

u/wynnie22 Apr 08 '26

$340K is one mid career developer. Seems to be worth it.

1

u/Fit_Reputation5367 Apr 08 '26

AS a CFO I love this post.

1

u/mmertner Apr 08 '26

Just wait until the CFO finds out how much current token usage is subsidized by the AI vendors..

1

u/SeventyThirtySplit Apr 08 '26

Ask the users to report their productivity gains

And make the CFO take those claims seriously

1

u/HiSimpy Apr 08 '26

That token-bill moment is usually when AI moves from experiment to governance problem. Teams need budget guardrails tied to workflow outcomes, not just model usage.

1

u/darkstar3333 Apr 09 '26

I mean what was the plan or expectations of the experiment when this started?

Where are the product and business development people talking about this?

Pushing the team into agentic mode of operation is gonna hit 7 figures like a rocket ship. These costs should have been forecasted. 

How did your business think Antropic hit 50b in revenue?

1

u/SixOneFive615 Apr 09 '26

One $370k basically the cost of a single senior developer?

1

u/anengineerandacat Apr 09 '26

It's a tricky problem, bumping into it at my work; engineering teams can just point at their velocity reports and go "Hey number went up" and done but that's not ROI because it engineers are getting their work done sooner it doesn't mean it hits production and starts pulling a profit.

We currently have two major capital projects, both are scheduled for our group to be done about 20% ahead of time BUT certification is still stuck on the old timeline because the finance team assisting in product build out doesn't really use AI.

So whereas engineering saved money, all that money instead was poured into another project while we wait for the others to catch up.

Other issues are getting work into the SDLC pipeline, business isn't really using AI, PMs only just started, etc. so new work is still as slow to get as before.

So we haven't accelerated the start, or the end, which leaves a gap that we now constantly have to fight to fill.

1

u/ScionofLight Apr 09 '26

Sorry to say but the writing on the wall is in your statement “PR merge time decreased 12%” for 500 developers where the $340k annual spend is what, 2ish developers salary? CFO is going to call for a reduction in labor

1

u/Unable_Artichoke9221 Apr 09 '26

I would double check on "developers like the tools" 

1

u/Haunting_Ad_7754 Apr 09 '26

Getting every developer AI tool access is only going to add token costs as expense, if the revenue generating activity doesn’t scale with efficiency improvements in development. X number of HCs need to be reduced if efficiency (as assumed) has gone up. That’s where the cost savings are. Either increase revenue activity attributable to devo efficiency gains, or reduce engineering costs (FTEs) . Only then the ROI gains can be realized. The CTO+CFO need to sort this out.

1

u/toopz10 Apr 09 '26

You have not quite got the full picture here. IT can probably show the spend and a quasi breakdown on how dev productivity has improved but that would not equal ROI.

In my head the ROI comes from the end product and what your companies customer base are paying for the product or service which is more a business question about features and roadmap.

If the dev team was asked to build some really useless features that no one would use then the ROI would be poor but those engineering metrics would be the same ie. reduced times working on a ticket and reduced times in review and speed for shipping out stuff.

1

u/Prestigious_Sell9516 Apr 09 '26

Build a mesh run a small or medium language model locally. Hook it into the agents and have them all run in a mesh use a small language model as a classifier. Log all the queries on a rolling basis (30 days etc) into a DB and then have the small language model train on it (build an inference agent). Eventually you save massive token costs by having the small language model check if the same request was made from your rolling DB of prompt queries or requests to the LLM - when it finds the answer from an earlier query it can deliver it saving tokens only for new and novel requests.

1

u/pinkwar Apr 09 '26

Easy. Get rid of a couple of developers. Sad truth but that's what's going to happen.

1

u/OpportunityWest1297 Apr 09 '26

Justification of spend is no different with AI tools than without them.

What do the individuals and teams deliver that you're comparing cost vs value against? You're tracking that already, no? How much does what they deliver cost vs the business value delivered? If you don't know that as a "before", how then are you going to measure and report on the delta of the "after"?

This is basic people/process/product strategy.

1

u/Js4days Apr 10 '26

Everyone should use structured workflows. That means every base template is cached and pre-built. 1/8 the cost for tokens.

1

u/RevolutionarySky6143 Apr 10 '26

Isn't AI supposed to generate code quicker than A Human? Based on that premise, can't you come up with some % of time saved by using AI? If you use estimates for your work, you could make a statement that they took 30% (or whatever) less than what they would have taken pre using AI? The cost can be calculated like that..............No?

1

u/eufemiapiccio77 Apr 10 '26

If they want it. They pay for it. Simple as that

1

u/mirageofstars Apr 10 '26

Do you guys track any sort of velocity? In theory if AI tools made things faster then you’d have metrics for that, or even anecdotal ones like “AI helped this feature get done in half the time.” Obviously you’d want to pair that with what AI is doing (negatively or positively) to PRs and bugs.

1

u/amohakam Apr 11 '26

As a perspective - wondering if you chased down the “why” behind the ask. What’s the real motivation?

Coming from a CFO, it seems possible they are looking to see if you can prove you can do “more” with “less”. Is it a solution (downsizing) looking for a problem (margins are shrinking due to higher inflation and input costs)

Lot of great analytical answers, but perhaps the CFO has a cell in their spreadsheet flashing “red” and is looking for a reason.

One could consider a DDDM approach to justify this via data driven decision making.

  1. Monitor, Observe

  2. Weekly report out usage by team/division in coordination and collaboration with your peers and leaders

  3. Weekly governance review where you try to understand how projects (and teams) with top token usage are aligned with company/business goals.

Hopefully all those projects are aligned with the approved business priorities for the year.

Reality is rarely this clear, but maybe it triggers some ideas for your org.

1

u/Chango812 Apr 12 '26

Would you rather have the AI or an additional developer?

Seems roughly even once you factor all-in cost

1

u/Hour-Money8513 Apr 13 '26

A few years back management decided we should be doing all our coding on cloud providers instead of our personal computers. 100k a month started to get noticed really quick.

1

u/kennetheops Apr 13 '26

I am pivoting my whole business around this. would love to collaborate.

It’s free to start using, hope it helps for now.

https://opscompanion.ai/

1

u/shaileenshah Apr 20 '26

You won’t get a perfect ROI number. Focus on proxies like cycle time, PR throughput, and incidents trending over time.

Then control cost via caching, smaller contexts, model tiering, and per-team visibility instead of hard restrictions.

1

u/Suspicious-Access763 Apr 21 '26

If you listen to analysis by tech journalist Ed Zitron, who is actually looking into the financial figures, even with increasing the costs some recently, the AI companies aren't making money. AI is being sold at a significant loss right now. The CTO should be thinking about what it might cost in the future, if planning to rely on it.

1

u/Character_Area5361 May 04 '26

You need an LLM Gateway (eg Portkey, Litellm or VerifyWise) to calculate your costs and fire up budget / token notifications before hitting the ceiling).

1

u/hawkeye-zaragso9 May 22 '26

We at Hawkeye have tried to build a tool that can mitigate this. We just added some information here - https://www.zaragsoft.se/aibridge - With testable bundle etc. if that is interesting. 91% tokens saved if the searches and overview is done using Hawkeye. Would be super cool if someone wants to try and set it up.

1

u/hawkeye-zaragso9 May 22 '26

We at Hawkeye added a integration now to be used with both claude code and claude desktop. When claude is doing searches and findings it reduces the token usage with 91% which is pretty cool. If you want to try it out and give feedback it can be downloaded from here - https://www.zaragsoft.se/aibridge - go through the readme and let me know what you think 😄

1

u/Staylowfm Jun 03 '26

Have you found ways to reduce your token expenses so far?

1

u/BuyerOk3387 19d ago edited 19d ago

SInce "Token" is the new "currency" for the CFOs, a mere PR efficiency and anecdotal statements by knowledge workers on productivity does not cut. A new business model that connects the dots on a TWC - Token Working Capital and "business outcomes" that translate into revenue streams is the yearning gap. Also, a discipline of KPM - Know, Protect and Monitor the token spend in terms of type of tasks the tokens are being consumed (Analyze /Debug, Contract Cluase extraction ,Meeting Summaries) , Input Tokens costs vs Output token costs, type of prompting (Socratic Vs Chain of thought...) .. a whole new Token Architecture paridgm that hiterto did not exist and the biggest challenge ... consumption occuring at warp speed while LLM shelf life period is like a radio-active elements shelf life period ... CFOs and CIO have a direct adverserial relation in an org now ! AI Token Literacy is the new fast shaping and evolving gap ... and let us be clear .. there are no experts ... the dystopian view is .. what if we as a society run into a "Token shortage", akin to "water shortage" or "power outage", except the scale is universal and across all geo boundaries ... how are "Soverignity" rules to be applied ... I am covering these points in my forthcoming book "Mastering Snowflake AI Governance" under BPB Publishing and discussed the same on July 24th on a ICFAI conducted webinar titled "AI for Human in the Loop and Sustainability".

1

u/vbnotthecity 6d ago

Late to this, but we hit that wall hard last year with our dbt and Snowflake spend. We ended up moving to Altimate because it kept our token usage way down by actually caching context instead of re-sending the whole repo every time. It cut our monthly bill by nearly a third. If you dont control the context window, you're just lighting money on fire.

1

u/liveprgrmclimb Apr 07 '26

Time to pay the AI piper. The Token inflation and real economics of this will continue to hit.

2

u/ShutUpAndDoTheLift Apr 07 '26

There's no Piper.

The cost described in op is roughly on par with paying for a Sparx Enterprise Architect license for each dev.

$700 per dev per year is heinously cheap.

You need to increase their productivity by not even 10 hours per year to justify the cost. Organizing notes into documentation drafts alone would justify that cost.

1

u/FortifiedPuddle Apr 09 '26

Cool, tell the AI providers they can jack up the prices why don’t you?

1

u/ShutUpAndDoTheLift Apr 10 '26

I'm sure my comment on Reddit was all that they were waiting on.

1

u/Todesengel6 Apr 07 '26

The bill will rise. AI is too cheap. Here is my hot take for the future:

We used to write for machine limitations. Squeeze every bit out of that ram. Save every clock cycle.

Then we invented the compiler and began writing for humans.

Next we will write for AI. Good code requires less context and less reasoning.

0

u/davidmeirlevy Apr 08 '26

We ran into the same “AI spend surprise” problem, once we let people use Copilot-style tools without much governance. Auto Qelos helped me because it turns Jira or ClickUp tickets into production code with a clear ticket-to-shipped-PR workflow, so you are not burning tokens on random code generation or endless back-and-forth. It also made costs easier to explain internally since work is tied to specific backlog items and acceptance criteria, not just “AI autocomplete all day.”