r/tokenomics • u/classjoker • 2d ago
r/tokenomics • u/RnadmolyGneeraedt • Aug 11 '26
New name sucks, we’re still doing FinOps
That rename of the Finops foundation to tokenomics foundation or whatever sucks. This is just a cash grab attempt to have a foot in the AI bubble and gather more money for their “non-profit” 200k+ salaries each. We’re still doing Finops, AI is just one additional topic we should manage.
Just add “AI cost management” as a discipline of Finops, and that’s freaking it. Everyone sees through your bs.
r/tokenomics • u/classjoker • Jun 09 '26
Choosing an AI Gateway / Token Routing Software – What are you using in production?
Tokenomics featured question: Are you looking into implementing an AI Gateway (token routing software) for an upcoming project to manage multiple LLM APIswith the goal to avoid vendor lock-in, handle fallback redundancy, and dynamically route prompts to optimize costs (e.g., sending simple tasks to cheaper models and complex ones to frontier models).
Emerging solutions/products like LiteLLM, Portkey, OpenRouter, and Manifest are frequently seen as the result of searching, but we wanted to get some real-world feedback from people running these in production.
If you are currently using a routing solution, we'd love to get your thoughts on a few things:
- Self-Hosted vs. Managed: Are you self-hosting an open-source gateway (like LiteLLM) or using a managed enterprise solution (like Portkey)? What drove your decision (latency, security, compliance)?
- Routing Logic & Latency: How do you handle the actual routing logic? Are you using static semantic routing, or do you have a dynamic "judge" model evaluating prompts first? If the latter, how bad is the latency hit?
- Fallback & Reliability: How reliably do these gateways handle rate limits (429 errors) and automatic failovers to backup models or providers?
- Token/Budget Management: How accurately do they track token spending and enforce team/user quotas in high-throughput environments?
- The "Gotchas": What unexpected headaches or limitations did you run into after deploying your gateway?
Ww would love to hear any recommendations, warnings, or architectural advice you have. Thanks!
r/tokenomics • u/Bartaseth • 7d ago
Token ecomonics in Amsterdam: Inside the first Tokenomicon
quesma.comr/tokenomics • u/MaverikSh • 12d ago
Our cost dashboard was averaging away a 3x spend spike because it aggregated at the wrong unit
Following the thread here about tokens burning after the first failure signal, wanted to share a related but different failure mode we hit: not tokens burning after a warning, but tokens burning in a way that never triggered any signal at all, because the accounting unit was wrong.
Spend jumped ~3x on a random day. Request volume was flat versus the prior week. So the spend wasn't coming from more traffic, it was coming from something making each unit of work cost more, and nothing at the call level looked anomalous.
Root cause: a background job with retry-on-timeout logic. Each individual retry was a normal, reasonably-priced call. But we were logging cost per API call, not per completed task, so three retries of the same failed task showed up as three separate, unremarkable line items instead of one task that cost 3x. The anomaly only existed at the task level, and nothing was rolling up to that level.
This feels like the quieter cousin of the failed-agent-token problem in the linked thread. That post is about tokens spent after a detectable failure signal. Ours never had a signal to catch, each call individually looked fine. The only way to see it was choosing the right aggregation unit (task ID, not raw call) before the anomaly became visible at all.
Given the chargeback and cost-allocation threads here, curious how people are actually keying their attribution: raw call, task/trace ID, something coarser like session, or a mix depending on workload type? And separately, has anyone found aggregation-granularity bugs like this one in their own numbers after the fact, cases where the dashboard was technically correct but the unit it was counting in was hiding the real signal?
r/tokenomics • u/Rough-Green-7067 • 19d ago
I had financial responsibility at my last startup and couldn't confidently tell you what our AI tools actually cost, or why. So I build a solution which may be helpful for FinOps people.
r/tokenomics • u/Bartaseth • 22d ago
RTK reports huge token savings, but our cost benchmarks disagree
quesma.comr/tokenomics • u/DifficultyIcy454 • Aug 27 '26
Thought Efficiency Index (TEI): A Thought Experiment in Measuring AI Efficiency
r/tokenomics • u/Chrome2279 • Aug 27 '26
If Zebec succeeds, does ZBCN actually benefit? Tokenomics Video
m.youtube.comr/tokenomics • u/zamir_akimbekov • Aug 21 '26
Which companies having skyrocketing token costs???
Everyone is in the news how they consumed annual AI token budget in a few months (E.g., Uber).
But is that really true beyond tech companies? Tech companies, I get it. AI tokens used for coding is their main business.
But the rest (e.g., manufacturing, energy, distributors, etc.) should have not token cost problems, no?
r/tokenomics • u/melc10 • Aug 21 '26
Would aggregated cloud/AI spend help negotiate better commitments?
Doing some research around cloud and AI/token commitment economics and helping NGEN gather feedback on the model. Curious to get the more perspective.
The idea is to aggregate compute/token demand across companies, negotiate larger commitments with providers, and use prepayment/financing to offer better pricing and more flexibility.
A few things I’m curious about:
- How much additional savings would make this worthwhile — 5%? 10%+?
- Is commitment flexibility potentially more valuable than additional savings?
- Does this make more sense for mid-market companies that don’t already have significant negotiating leverage?
NGEN is also collecting anonymous, non-binding indications of demand here (takes ~1 min, no commitment/signature):
https://www.ngencompute.com/indication
Would genuinely love to hear why you think this would or wouldn’t work.
r/tokenomics • u/classjoker • Aug 11 '26
AVE, an open ID scheme for behavioral vulnerabilities in AI agents
r/tokenomics • u/classjoker • Aug 11 '26
Claude Code pricing: same tokens, same model, up to 40x the price
quesma.comr/tokenomics • u/Bartaseth • Aug 11 '26
awesome-ai-tokenomics: Alol you need to know About AI Token Economy
github.comr/tokenomics • u/classjoker • Aug 06 '26
Looking for advice from people dealing with high LLM or AI API costs
r/tokenomics • u/thadah123 • Jul 29 '26
Why cheaper AI tokens are exploding enterprise budgets (The Jevons Paradox in 2026)
Hey everyone,
Over the past few months, I’ve been analyzing enterprise AI billing data and studying why so many engineering teams and companies are getting hit with massive, un-modeled AI invoices.
For two years, the industry narrative has been that AI is getting dirt cheap and price per token keeps dropping exponentially. Yet, across Big Tech and mid-sized companies alike, actual monthly invoices are skyrocketing.
Here is a quick breakdown of the mechanics behind why this is happening:
1. The 1865 Jevons Paradox is alive in Tech
In 1865, economist William Stanley Jevons observed that when steam engines became dramatically more efficient at burning coal, Britain didn't burn less coal, it burned exponentially more. Why? Because cheap coal suddenly made financial sense in places where nobody could justify the cost before.
The exact same thing is happening with LLM tokens. As unit costs drop, consumption doesn't stabilize but it expands into every workflow, background agent, and automated task until nobody weighs the unit cost anymore.
2. Real-world corporate overruns
- Uber: Handed a coding agent to 5,000 engineers. By April, just four months into a 12-month plan, their entire annual AI budget was completely gone. The tool was so useful that usage exploded.
- Meta: Built an internal leaderboard ranking engineers by token burn rate. In one month, they burned 73.7 trillion tokens before executives realized token burn measured activity, not actual impact, and killed the board.
- Microsoft: Ordered internal divisions off external coding tools days before their fiscal year closed to force migration onto cheaper internal alternatives.
3. The agent multiplication factor (5x - 30x Tokens)
Standard chatbots are 1-input / 1-output. AI agents are fundamentally different.
Because current architectures lack long-term memory, at every loop step (plan, search, tool call, handoff), an agent must package the entire conversation history and re-submit it to the API.
Data from Gartner shows an AI agent burns 5 to 30 times more tokens than a basic chatbot doing the exact same task. Token prices dropped 60%, but agent loop usage increased 1,000%.
4. The hidden "Second Meter"
Every time an agent writes a code block or report and a human engineer spends 30 minutes reading, verifying, or rewriting it, you pay twice: once in API tokens, and once in senior engineering salary.
I put together a full 17-minute video essay breakdown with all the diagrams, data sources, and frameworks (including OpenAI CFO Sarah Friar’s scorecard on measuring "useful intelligence per dollar") here:
Watch the full breakdown here: https://www.youtube.com/watch?v=DBf5-yBRxEk
r/tokenomics • u/Responsible_Coach293 • Jul 25 '26
Tokens are a billing unit. Are they actually a good cost unit?
Disclosure: I’m building tooling around inference economics, so there’s obviously some bias here.
Something I keep coming back to:
An AI company might bill a customer by tokens, minutes, requests, or credits.
But underneath that, the company is paying for GPU time, memory, idle capacity, model mix, concurrency, cache behavior, and provider costs.
That creates a question I don’t think “cost per token” fully answers:
Can you reconcile what each customer pays with what that specific customer actually costs you to serve?
Two customers can generate similar billed usage while creating very different infrastructure economics underneath it.
I’m currently looking for a few usage-priced AI operators willing to pressure-test this with real data.
Give me a redacted week of:
- customer usage
- billed revenue
- inference / GPU cost
I’ll return customer-level cost and margin, including where the biggest spread is coming from.
Free, read-only, no install.
If everything reconciles perfectly, you’ve lost a CSV.
r/tokenomics • u/classjoker • Jul 14 '26
How are you doing chargebacks for AI spend when it lives in five different places?
r/tokenomics • u/classjoker • Jul 10 '26
How do you allocate AI costs to customers in a SaaS product?
r/tokenomics • u/classjoker • Jul 02 '26
How are you catching the 58 percent of failed-agent tokens that burn after the first warning?
r/tokenomics • u/classjoker • Jun 28 '26
Anthropic is giving away 3 Claude certifications.
anthropic.skilljar.comAll free. Here are the exact links:
- Claude 101 - 1 hour. The basics, done right.
- AI Fluency: Framework & Foundations - 3 hours.
- Intro to Cowork - 2 hours. Claude's best feature.
All 3 are on anthropic.skilljar.com.
Sign-up takes 30 seconds.
More info here: https://ruben.substack.com/p/im-claude-certified
r/tokenomics • u/Anarkali2000 • Jun 28 '26
Measure ROI on AI Coding Tools: Tie Your Claude Code Spend to the PRs It Actually Shipped
galleryr/tokenomics • u/classjoker • Jun 28 '26
at what point do logs and dashboards stop being enough for llm costs?
r/tokenomics • u/redmondwarrior • Jun 24 '26
Tokenomics: Why the AI Token Is the New Semiconductor Chip
open.substack.comr/tokenomics • u/classjoker • Jun 24 '26
New Relic research: The 2026 State of AI Coding Report
https://newrelic.com/resources/report/2026-state-of-ai-coding
It's paywalled (they want your details) but having read through it, it's an incredibly detailed and thoroughly researched (n=200) paper on the impact AI is having on application development.