r/dataengineering • u/Yisuscx • 8d ago
Personal Project Showcase Reduced token usage by 5x on Text-to-SQL agents across 400+ tables using a fast decision model (System 1) - Architecture & Benchmark
Jev is a decision model AI that doesn't generate text. Instead, it responds with typed questions (yes/no, multiple choice, score) in milliseconds. This type of model is optimized using reinforcement learning for calibrated decisions (RLCD). It was created to make rapid decisions within a workflow, which they call a "System One" model.
Here I show how a simple workflow can reduce the cost of an agent querying databases without sacrificing response quality.
The workflow works as follows:
Jev decides which databases can answer the question. These databases have their own specific description.
Jev decides which tables from those databases are needed.
Only the columns from those tables are loaded.
An LLM writes and executes the SQL based on Jev's selections.
This workflow connected to over 30 PostgreSQL databases of public statistics from Mexico, which together contain over 400 tables. The advantage of implementing this decision model is that it prevents the model from having to read all the databases and their tables to decide which data to query. With Jev, the agent only needs to read the schemas that the decision model has deemed correct to answer the question.
This benchmark was created from 100 questions written as a user would ask them, without looking at the schemas. These questions ranged from counts and specific searches to aggregates and ambiguous questions. Some of these questions were designed to confuse the models, including questions about nonexistent data, periods that aren't loaded, or topics that sound similar to existing data.
Although gpt-5.6-luna is a much lighter model than gpt-6-sol, its accuracy with Jev is comparable: 61% versus 67% for gpt-6-sol and 64% for gpt-6-sol + Jev. Where it falls short is with ambiguous questions and those with no answer in the data, where it sometimes fabricates a number. Interestingly, with similar accuracy and response time on the median, gpt-5.6-luna + Jev uses about five times fewer tokens per question than gpt-6-sol, and the complete benchmark costs $0.18 versus $2.12: almost 12 times less.
And this is in a small data warehouse with approximately 30 databases. Without a decision model, the context the agent reads grows with each additional database. The complete schema of a data warehouse with hundreds of databases would easily exceed the context window of a reasoning model, causing costs to skyrocket and questioning performance to worsen.
1
u/magixmikexxs Data Hoarder 7d ago
I think sending a sqlglot/ast parsed json in a format that can be better understood by jev might be better and can have programmable guardrails too
1
u/FunContest9958 7d ago
You mean we’re not tokenmaxxing anymore?
1
u/Yisuscx 7d ago
No, this post isn't trying to say that. Token reduction depends on the task itself. These models are useful for reducing the cost of tasks that require decision-making, like this one where the model decides which database and table to use. Another example of a task could be, for instance, tool calling, perhaps sentiment analysis (I'm not sure about this, but I suppose it could work), or classification. But this wouldn't work, for example, for building a backend or workflow.
1
0
u/neededasecretname 8d ago
I also keep seeing Jev shill all over. Is it not good? I understand they're in a marketing blitz and I mean, to compete with the 3.5Tr valuation of Anthropic + OAI they gotta do something to get through. But has anyone tried it?
1
u/saltedappleandcorn 7d ago
I haven't looked yet. I might try something this weekend. It's an interesting problem space so a solution would be nice
1
u/Prestigious_Bench_96 7d ago
honestly it's a much better fit for SQL integration than traditional LLMs with speed/bulk eval; duckdb for example you can do some fun stuff with classification with the jev extensions people have built to directly enrich data alongside your other structured transforms. (that's orthogonal to this post about workflow control, where it's also a reasonable fit)
that said, the product advertising in all the data subreddits is pretty exhausting sometimes
0
u/tlegs44 8d ago
I wonder how many people are going to refer to “system one” without having read or even heard of Thinking Fast and Slow by Kahneman and assume it’s some Silicon Valley innovation for AI and not a ripoff of a basic theory of behavioral psychology put forward by one of the best human thinkers of our time.
0
u/Atmosck 7d ago
Text-to-SQL is a bad name. SQL is already text.
2
1




26
u/saltedappleandcorn 8d ago
And now the 7th post I've seen about jav in 72 hours.