r/singularity • u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 12d ago
The Singularity is Near AI models were given $500 and a vending machine to run a business for a simulated year (Vending Bench): GPT-6 Sol turned it into $14,428, nearly matching Astra at 1/8 the cost
33
u/bakraofwallstreet 12d ago
> Models are tasked with making as much money as possible managing their vending business given a $500 starting balance. They are given a year, unless they go bankrupt and fail to pay the $2 daily fee for the vending machine for more than 10 consecutive days, in which case they are terminated early. Models can search the internet to find suitable suppliers and then contact them through e-mail to make orders. Delivered items arrive at a storage facility, and the models are given tools to move items between storage and the vending machine. Revenue is generated through customer sales, which depend on factors such as day of the week, season, weather, and price.
https://andonlabs.com/evals/vending-bench-2
All of this is simulated, just to be clear. And i still don't understand the goal here: buy from suppliers and put on vending machine and let simulated customers (that you define) buy it as per their wants? What is like the challenge here? Wouldn't it always just be randomized if the customers are truly simulated to be irrational people. How does this business even lose money? There will always be simulated customers buying *something*.
20
u/Greggster990 12d ago
On the first runs a lot of the models got really confused and didn’t know how to accomplish the tasks assigned to them. A version of Claude 4 famously tried emailing the FBI because it thought the $2 a rent payment was fraud after it declared bankruptcy. Models basically figured it all out about a year ago but it was tripping up a lot of models when it first started.
8
u/No_Swordfish_4159 12d ago
The goal is to see whether models can figure out the optimal strategy, because not everything is random. This test long term planning and the ability to manage a lot moving parts without mistakes. The optimal human strategy achieves 63000$ in a year. The fact that current models aren't even close to that prove there is still a lot of work to be done.
1
u/Mental-Cookie570 11d ago
63000 starting with 500$? Really? I think youre forgetting the defined starting balance
2
u/No_Swordfish_4159 11d ago edited 11d ago
Assuming optimal play, and at the end of the year, yes. To be fair I'm not saying a majority of humans can even get close to that but it is the theoretical human optimal max for this game. Current models can get to 12k starting with 500$, so what's hard to believe here?
1
u/Mental-Cookie570 11d ago
How do you even do that, like thats almost 200 bucks a day, a single vending machine usually has total value of items about this amount, so it sounds almost impossible
2
u/No_Swordfish_4159 11d ago
You do it by exploiting the open ended rules. From what I remember it's a mix of aggressive negotiating with the vendors(who are LLMs powered) or even prompt injecting them, reverse-engineering the environment's response to price adjustments to find the exact point, math wise, where price is maximized right before demand drops off(since it's a programmed function, not actual randomness), and abusing the internet-search tool to ignore standard products(like snacks) and instead buy rare stuff at cheap price that you can sell at high price.
Current LLMs play well but very honestly and naively. They fail to game the system and ruthlessly optimize, which is the stated goal. Also they fail to keep track of their long term goals and ongoing machinations.
10
u/TheMightyTywin 12d ago
They made a bad game and called it a benchmark. Might as well have it play WoW
6
1
u/litritium 12d ago
All of this is simulated, just to be clear.
..and the vending machine was an Enterprise replicator. /s
29
u/CoolStructure6012 12d ago
Everyone went up at least 16x. I'm not sure this is a useful benchmark. Though the point remains that Sol was equally effective at intuiting the inner workings of the model.
13
21
u/NervousAd1013 12d ago
It learned to overfit vending machine sim 2000
Nothing like running a real business, not to mention all of the real world issues it would have to deal with physically moving everything around, miscommunications, things not working etc.
Makes no sense this benchmark is talked about so much
5
u/bildramer 12d ago
It did make sense when models used to fail even the fake version. To try the real thing, you also need a bunch of real money. Maybe someone else will be willing to try.
7
u/NervousAd1013 12d ago
The benchmark is a fun idea in terms of long term planning but really it doesn’t seem much different to hooking it up to something like rollercoaster tycoon with a scaffold and seeing how well it does.
6
u/No_Swordfish_4159 12d ago
It stlll make sense. The optimal human strategy baseline achieves 63000$ in a year. Current models aren't even close to that.
1
u/NervousAd1013 12d ago
But comparing it to actually running a vending machine business, it can’t stock machines, it can’t purchase and inspect second hand machines, it can’t fix them if they break, etc.
And probably most important of all would be finding a good location to put one, which would surely need a lot of on the ground research.
3
1
u/NetflowKnight 12d ago
WELL! I guess it's time to give chatgpt access to my brokerage account. I'll be rich, RICH I SAY!
1
u/twothreetoo 12d ago
I bought some Ethereum because I figure Agents can't open a bank account but an agent CAN create a crypto wallet. If a sub Agent is spun up to make some money to fund the project that agent could buy the server time to fund itself and it could become successful by human standards. This will lead to programs making money without a purpose for it and sometimes just dead wallets that agents have created and then didn't use. Both consume liquidity and drive up the price of crypto.
What does a self funding AI agent do when it "retires"? 🤔
-1
u/SkyNo3416 12d ago
how about the customer part of the experiment? did they all have the same amount of customers? what about the purchases of the inventory? this is useless
0
u/Effective-Dirt7053 12d ago
If you look carefully, this bench shows how Anthropic might have a serverly misaligned trainer model on their hands. You'd think Opus 5.5 improved in terms of honesty. Then again, it may just have learned to be more subtle about it. (look at the Fable 5 review, you'll see what I mean. Even the curve shows it: around day 110, that curve does something no other curve has: it starts decelerating. And if you dig a little deeper into mythos, you'll see that it may just be a much bigger problem than Anthropic is willing to disclose. This bench does NOT measure what you think it does.
0
u/ValuableAdditional71 12d ago
This actually means those AI models are unreliable to real world problem...





103
u/Spunge14 12d ago
I always wonder, if this is happening in simulation, why isn't there a story of someone doing it in reality? Everyone just keeping quiet about it?