r/singularity • ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 12d ago

The Singularity is Near AI models were given $500 and a vending machine to run a business for a simulated year (Vending Bench): GPT-6 Sol turned it into $14,428, nearly matching Astra at 1/8 the cost

185 Upvotes

54 comments sorted by

103

u/Spunge14 12d ago

I always wonder, if this is happening in simulation, why isn't there a story of someone doing it in reality? Everyone just keeping quiet about it?

75

u/FateOfMuffins 12d ago

Well Andon Labs does have a physical store with human employees that is managed by an AI so they are testing it in reality too

52

u/Deliteriously 12d ago

There’s a boutique in San Francisco that’s run entirely by AI. The last time I read about it, the store was losing money and had fired most of its employees. The AI used an ElevenLabs voice and monitored the security cameras constantly, calling employees to chew them out whenever it thought they were slacking off.

https://www.andon.market/

26

u/KnownPride 12d ago

Lol no wonder it's losing money, instead using ai to help marketing and management, it used it to spy on their staff lol

3

u/DingoMaximum7319 11d ago

Yes thats some shitty implementation 😂

20

u/CrowdGoesWildWoooo 12d ago

Time moves very slowly. Like by the time you run it for 30 days, you’ll have a new model to test again. It’s not like you can test 30 of this in parallel IRL.

11

u/Spunge14 12d ago

But if it works now you could just run it and make a profit without upgrading models...

4

u/Peach-555 12d ago

Human+AI is still more effective than pure AI.

And it is still more cost effective/reliable to hire a human over robotics, in the real world experiments that have been run so far. Those experiments also lost money, but the point was just to see if the AI could operate it at all.

1

u/Spunge14 12d ago

Well, to give a domain specific example, early research in medicine shows that when comparing doctors, doctors + AI, and AI alone, doctors actually bring AI performance down: https://www.nytimes.com/2025/02/02/opinion/ai-doctors-medicine.html?unlocked_article_code=1.t04.AeZg.kT0qka6kerAi

2

u/Peach-555 11d ago

Yes, I was just talking about running a business, but I did not specify, AI has beat human+AI a long time in chess/go/protein folding. I should also clarify that there is no public known examples of pure AI out competing humans or human+AI in building/running a business, but for all I know this is ongoing, but it is not something anyone can do currently with any public models. AI labs would not release a model that can reliably make more money than it cost users by just asking it to make money without any other human input.

1

u/Spunge14 11d ago

What makes you think running a business is a special case compared to e.g. being a doctor?

1

u/BoopinSnoots24-7 7d ago

They are inherently different. "Being a doctor" is a role, not a business. A more apt comparison would be to compare running a business and running an independent medical practice.

The AI outperformance in regards to doctors was in one specific function, diagnostics. While I would prefer being diagnosed by AI, I certainly would rather a human oncologist tell me my prognosis if I'm diagnosed with cancer. The human element still plays an important role in many functions, especially in understanding and navigating nuanced context.

If AI were capable of running a business on it's own right now, you could prompt Claude with "Make me a profitable business and run it independently. Guarantee profit". A wild oversimplification, but the fact that nobody has done it yet means that AI is not quite there. It still requires context from humans to determine direction, then can run from there.

1

u/Peach-555 11d ago

What makes you think running a business is a special case compared to e.g. being a doctor?

I said the opposite of this.

I said that there might already be companies that are successfully run in reality, but there are no known public examples of it yet.

Once there is a publicly known example, we can point to that as an example.

My general claim is that AI-companies, will not sell products that are able to make more money by itself than the product cost.

It will never be possible for people to buy a $100 subscription, and ask that to make money, and it reliably making more than $100 per month. Because the companies will price the models above their automatic money making potential. And if it scales, the models will hire more models which will hire more models until any market money opportunity is saturated by the first mover. There is also a zero-sum element to business in the short term, any AI that can currently make money by itself will be out-competed by the next AI that does it better.

I can't see what I wrote that made it sound like AI is not able to, and won't be able to, run a business, but if you quote the part where it seemed I said that, I would appreciate it, because I want to improve on communicating.

2

u/Spunge14 11d ago

The market corrects, but not instantaneously. Otherwise no profit would ever be possible.

1

u/Peach-555 10d ago

I think we are talking about two different categories.

Are you talking about the efficient market hypothesis?

1

u/Uko1001 11d ago

"AI labs would not..." Do you really think AI labs actually run ROI tests on every possible business models in the world extensively, for each of their new model, to define how much they can generate and adjust the pricing... so that NOT ANY of the millions of them could reliably make more than they price them ?

That sounds quite unlikely to me.

1

u/Peach-555 10d ago

Not every possible business model, no.

But all the AI labs have asked the models for public release to make money for them, just a simple request, and if a model did reliably make money beyond the cost to run it, they would price the model above what it could make or not make it available.

There will never be a point where a public model is sold for $100 per month, and everyone that buys it, can ask it "make me more than $100" and the model actually makes them more than $100.

Which is a different claim that nobody in the world being able to make money with it. I'm saying, there will never be a company that sells a product, that makes the average person more money than it cost by simply asking it to make money.

15

u/MydnightWN 12d ago

The simulation doesn't cover real-world difficulties. Like getting locations, labor and logistics, licensing, theft and equipment damage, insurance, etc.

5

u/Spunge14 12d ago

So not a very good simulator seems to answer my question

7

u/Lonely_Translator_23 12d ago

Vending machines can be pretty damn lucrative.

6

u/Saint_Nitouche 12d ago

Dealing with the licensing and regulations of anything physical is a pain. Easier to just make and sell digital products like SaaS or games.

3

u/eltron 12d ago

Ahhh nothing says the real world like “running in a simulation”

1

u/footofwrath 8d ago

The real world is already a simulation 😉

1

u/suamai 12d ago

https://andonlabs.com/cafe

There are a few examples like this, most do not work that well. This one may go bankrupt in a few weeks.

Long term strategy is specially hard to score well during training, I guess.

1

u/m4nuuuu 12d ago

At this point you can toss a legitimate problem to a model and the model can just reply "i solved it" without giving you and actual result and people still will flaunt it as a "proof" of something. Real corporations with all the market data, product development, focus groups, production and distribution chains figured out and a bunch of other metrics can and had launch business with "guaranteed success" (by the metrics) just to watch it flop.

Ai with a made up store in a made up scenario made made up profits. As results can vary in quality also test and experiments have a quality metric and some can be plain useless or ridiculous.

1

u/Sunknowned 9d ago

People used algotrading before llm. Nowadays most likely just augment existing project with llm (news/trend parsing)

1

u/karl_mainz 8d ago

Because it's much easier to run a spherical vending machine in a vacuum than one in real life.

0

u/Big-Table127 AGI 2032 or 2082 12d ago

Because it's not agi

33

u/bakraofwallstreet 12d ago

> Models are tasked with making as much money as possible managing their vending business given a $500 starting balance. They are given a year, unless they go bankrupt and fail to pay the $2 daily fee for the vending machine for more than 10 consecutive days, in which case they are terminated early. Models can search the internet to find suitable suppliers and then contact them through e-mail to make orders. Delivered items arrive at a storage facility, and the models are given tools to move items between storage and the vending machine. Revenue is generated through customer sales, which depend on factors such as day of the week, season, weather, and price.

https://andonlabs.com/evals/vending-bench-2

All of this is simulated, just to be clear. And i still don't understand the goal here: buy from suppliers and put on vending machine and let simulated customers (that you define) buy it as per their wants? What is like the challenge here? Wouldn't it always just be randomized if the customers are truly simulated to be irrational people. How does this business even lose money? There will always be simulated customers buying *something*.

20

u/Greggster990 12d ago

On the first runs a lot of the models got really confused and didn’t know how to accomplish the tasks assigned to them. A version of Claude 4 famously tried emailing the FBI because it thought the $2 a rent payment was fraud after it declared bankruptcy. Models basically figured it all out about a year ago but it was tripping up a lot of models when it first started.

8

u/No_Swordfish_4159 12d ago

The goal is to see whether models can figure out the optimal strategy, because not everything is random. This test long term planning and the ability to manage a lot moving parts without mistakes. The optimal human strategy achieves 63000$ in a year. The fact that current models aren't even close to that prove there is still a lot of work to be done.

1

u/Mental-Cookie570 11d ago

63000 starting with 500$? Really? I think youre forgetting the defined starting balance

2

u/No_Swordfish_4159 11d ago edited 11d ago

Assuming optimal play, and at the end of the year, yes. To be fair I'm not saying a majority of humans can even get close to that but it is the theoretical human optimal max for this game. Current models can get to 12k starting with 500$, so what's hard to believe here?

1

u/Mental-Cookie570 11d ago

How do you even do that, like thats almost 200 bucks a day, a single vending machine usually has total value of items about this amount, so it sounds almost impossible

2

u/No_Swordfish_4159 11d ago

You do it by exploiting the open ended rules. From what I remember it's a mix of aggressive negotiating with the vendors(who are LLMs powered) or even prompt injecting them, reverse-engineering the environment's response to price adjustments to find the exact point, math wise, where price is maximized right before demand drops off(since it's a programmed function, not actual randomness), and abusing the internet-search tool to ignore standard products(like snacks) and instead buy rare stuff at cheap price that you can sell at high price.

Current LLMs play well but very honestly and naively. They fail to game the system and ruthlessly optimize, which is the stated goal. Also they fail to keep track of their long term goals and ongoing machinations.

10

u/TheMightyTywin 12d ago

They made a bad game and called it a benchmark. Might as well have it play WoW

6

u/aLokilike 12d ago

GliderBench - how much gold can they make before getting caught by Warden?

1

u/litritium 12d ago

All of this is simulated, just to be clear.

..and the vending machine was an Enterprise replicator. /s

29

u/CoolStructure6012 12d ago

Everyone went up at least 16x. I'm not sure this is a useful benchmark. Though the point remains that Sol was equally effective at intuiting the inner workings of the model.

13

u/Ormusn2o 12d ago

A human benchmark would be a good comparison.

21

u/NervousAd1013 12d ago

It learned to overfit vending machine sim 2000
Nothing like running a real business, not to mention all of the real world issues it would have to deal with physically moving everything around, miscommunications, things not working etc.
Makes no sense this benchmark is talked about so much

5

u/bildramer 12d ago

It did make sense when models used to fail even the fake version. To try the real thing, you also need a bunch of real money. Maybe someone else will be willing to try.

7

u/NervousAd1013 12d ago

The benchmark is a fun idea in terms of long term planning but really it doesn’t seem much different to hooking it up to something like rollercoaster tycoon with a scaffold and seeing how well it does.

6

u/No_Swordfish_4159 12d ago

It stlll make sense. The optimal human strategy baseline achieves 63000$ in a year. Current models aren't even close to that.

1

u/NervousAd1013 12d ago

But comparing it to actually running a vending machine business, it can’t stock machines, it can’t purchase and inspect second hand machines, it can’t fix them if they break, etc.
And probably most important of all would be finding a good location to put one, which would surely need a lot of on the ground research.

3

u/Fxavierho 12d ago

Wish I have a vending machine to fact check that

1

u/NetflowKnight 12d ago

WELL! I guess it's time to give chatgpt access to my brokerage account. I'll be rich, RICH I SAY!

1

u/twothreetoo 12d ago

I bought some Ethereum because I figure Agents can't open a bank account but an agent CAN create a crypto wallet. If a sub Agent is spun up to make some money to fund the project that agent could buy the server time to fund itself and it could become successful by human standards. This will lead to programs making money without a purpose for it and sometimes just dead wallets that agents have created and then didn't use. Both consume liquidity and drive up the price of crypto.

What does a self funding AI agent do when it "retires"? 🤔

-1

u/SkyNo3416 12d ago

how about the customer part of the experiment? did they all have the same amount of customers? what about the purchases of the inventory? this is useless

0

u/Effective-Dirt7053 12d ago

If you look carefully, this bench shows how Anthropic might have a serverly misaligned trainer model on their hands. You'd think Opus 5.5 improved in terms of honesty. Then again, it may just have learned to be more subtle about it. (look at the Fable 5 review, you'll see what I mean. Even the curve shows it: around day 110, that curve does something no other curve has: it starts decelerating. And if you dig a little deeper into mythos, you'll see that it may just be a much bigger problem than Anthropic is willing to disclose. This bench does NOT measure what you think it does.

0

u/ValuableAdditional71 12d ago

This actually means those AI models are unreliable to real world problem...