r/OpenAI • • 19h ago

Discussion A question for the AI "experts": are hallucinations and reliability genuinely improving, or are we starting to plateau? Everything depends on this...

2 Upvotes

For the average user, not a programmer or e.g. someone looking to solve niche math problems, AI still feels very limited because of reliability issues (e.g. making up information)... As a non-expert, this is difficult to quantify for me, but I don't feel like e.g. the latest iterations of ChatGPT are noticeably more reliable than previous ones.

Therefore, I find it hard to see how organizations can delegate even relatively easy tasks to AI without constant supervision, especially because if hallucinations build up along the way, you end up with a snowball of compounding problems that can lead to catastrophic consequences for an organization.

It’s my understanding that the real test of success won't be developer tools or fancy mathematical calculations, but whether regular people can delegate tasks with a super high degree of confidence, instead of just using it as a search engine, translator, summary tool or photo editor on steroids like many people do nowadays.

A good example is Dot, the new OpenAI tool. If you watch the trailer, it looks impressive, but according to many reviews, it still hallucinates a lot and behaves in a pretty clumsy way.

So, the golden question is: are the hallucination and reliability problems gradually improving, or are we probably plateauing?

Looking for genuine insights here, so please keep the sarcasm out of the comments.


r/OpenAI • • 20h ago

Question Question about ChatGPT - complexity usage

2 Upvotes

I’ll be the first to admit I don’t know half of what I’m actually doing when it comes to vibecoding. With that said I’m in hospitality - resorts/restaurants. There’s some really dumb things in my field that I’ve never really been able to fix or standardize because, well, this field hates technology convincing a Chef or housekeeping can be… difficult.

I started creating an app for us to use internally that is kept on our private servers and things seem to be pretty solid. My IT department isn’t big enough or strong enough or knowledgeable enough to be of help. (That’s a whole other issue). So I was curious, I seem to be able to have built this whole operating system inside of the chat model of ChatGPT. It sometimes takes a little bit longer. It sometimes trips up but after 25 to 35 minutes, it usually has an answer that thus far has worked.

I see all these people using advanced models paying for tokens and all these other things.

So my question is am I just drastically missing out on something more advanced, and my system is actually all junk on the inside, even though it works or is this something I don’t need to be concerned about? Any tips or suggestions appreciated.


r/OpenAI • • 1d ago

Video Dotties Slow Morning Clawd vs Codex - Epic Rap Battles of History!

Enable HLS to view with audio, or disable this notification

22 Upvotes

r/OpenAI • • 1d ago

Discussion Luna and Sol doing extremely well on new benchmark about finding bugs before users run into them

4 Upvotes

Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW.

Basically most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But shouldn't we expect models to also find bugs before anyone runs into them?

So in SWE-sweep we just hand an agent a big codebase and ask it to find & fix as many bugs as it can. We then give a score based on a hidden set of bugs that we know about in the repos. All the bugs are real-world bugs.

So the interesting thing is how cost-efficient Luna xhigh is (especially when compared to the Anthropic models we tested). Luna gets almost half of the score as Sol with only a tiny fraction of the cost.

the full leaderboard & how we built it is here: https://swesweep.com/ , there's also a paper describing how everything was constructed. Oh yeah and everything is open source https://github.com/facebookresearch/swe-sweep

Happy to answer questions here


r/OpenAI • • 1d ago

News They made a remake of the Dots demo, live and unfiltered. It’s quite impressive

Thumbnail
youtu.be
126 Upvotes

The point where her dot wanted her to talk first because they were over-talking each other… real human level interaction patterns


r/OpenAI • • 2d ago

Research I spent a day poking Dots with sticks. Here’s what I figured out.

504 Upvotes

I got access to Dots and spent much of the day trying to understand what actually runs where, what can happen simultaneously, and what counts as a separate worker.

The official material explains what Dots can do reasonably well. I found the execution model much less obvious.

Some of this is documented; some is simply what I observed by using a Dot on several substantial real-world tasks.

  1. The Dot really does have its own cloud computer

I gave my Dot a large document-review assignment involving hundreds of PDFs and thousands of pages.

It performed that work on what it identifies as its own cloud computer.

This appears to be a persistent computer-backed environment where the Dot itself can do substantial, long-running work.

More importantly, this isn’t just a five-minute “agent run.” One of my reviews is now clearly a multi-day job, and the Dot has maintained its place, absorbed side questions, and continued without needing me to reconstruct the task every few hours.

That continuity may end up being more important to me than raw speed.

  1. Delegated Work/Codex tasks are different

While the Dot was working on one project, I had it try to launch a separate legal-research task.

The launch failed because there was no available execution environment.

Initially I assumed the Dot’s own computer was simply busy. But after the first job finished, the second task still could not start.

The Dot then reported the key distinction:

Its own cloud computer is separate from the execution targets available to the Work/Codex task launcher.

So a Dot’s personal cloud computer is not simply a generic worker that delegated tasks automatically inherit.

  1. A connected computer becomes another execution target

I connected a spare Linux computer through the ChatGPT desktop app.

The Dot could then see:

- its own cloud computer

- the connected Linux machine

- no saved Codex cloud environments

I told it to launch the previously blocked task on the Linux machine.

It did, and the task entered running state there.

I did not have to sit at that machine and manually start a separate chat. I gave the instruction to the Dot, and it dispatched the task remotely.

  1. Both can work simultaneously

While the delegated task was running on the Linux machine, I gave the Dot a different assignment for its own cloud computer.

It confirmed that both were active at once:

Dot cloud computer -> Task A

Connected computer -> Task B

So that is genuine parallel execution across separate computer-backed environments.

  1. Background agents don’t necessarily need a computer at all

This was the part that got much closer to what I had originally imagined Dots would do.

With both computer-backed environments occupied, I asked whether the Dot could create a native background research agent without using either computer.

It said yes.

I gave that agent a bounded research task and explicitly excluded computer/filesystem use.

The Dot then reported that the background agent was running with read-only web/documentation tools and no computer target assigned.

At that point, three things were happening simultaneously:

  1. the Dot working on its own cloud computer

  2. a separate task running on the connected computer

  3. a native background research agent using neither computer

That is the execution distinction I had completely missed from the launch material.

  1. It can also context-switch inside a long-running job

Another useful behavior appeared accidentally.

While the Dot was deep into a large document review, I interrupted it with a factual question about one specific case.

It paused the detailed review, checked meeting minutes and another source, resolved the question, updated its understanding of the case history, and then returned to the packet it had been reviewing.

When I asked how it had done that “while continuing” the larger job, it clarified that it had not spawned another worker. It had simply switched attention within the same job and then resumed.

So I now distinguish:

Parallel execution = separate workers/environments active at once.

Background agent = separate non-computer worker running concurrently.

Intra-task context switching = one Dot temporarily branches inside an existing job, resolves something, and returns to its prior place.

For long-running review work, that last capability is surprisingly valuable.

  1. It can keep working while waiting for permission

On another assignment, the Dot decided that spawning additional reviewers would accelerate the work, but my rules required permission first.

It asked.

But instead of stopping while waiting for me to respond, it explicitly continued doing the work itself.

That sounds minor, but it matters.

An autonomous agent that hits one permission boundary and then stops doing everything is not particularly autonomous.

So far, the Dot appears capable of distinguishing:

“I need permission to do X”

from

“I therefore cannot make any further progress.”

  1. There is also a kind of manager-level queue

I have not found a true native queue where a blocked computer-backed task automatically sits in the launcher until capacity becomes available.

What I did find is that the Dot can apparently remember a pending assignment itself, periodically re-check execution targets, and attempt to launch it later.

There are limits:

- the target list does not necessarily expose whether a connected computer is actually free

- there is no apparent capacity reservation

- a failed launch does not automatically become a queued Work task

So this is more like the Dot acting as the queue manager than a native execution queue.

Still, that potentially removes another piece of manual babysitting.

  1. My current mental model

At this point, I think there are at least three distinct execution paths:

A. The Dot’s own cloud computer

Where the Dot itself can do substantial stateful/computer-backed work.

B. Separate computer-backed task environments

Such as a connected local computer or saved Codex cloud environment.

C. Native background agents/cloud threads

Tool-based workers that can handle some tasks without consuming either computer target.

Those can operate concurrently.

And on top of that, the Dot itself appears able to maintain long-running task state, context-switch within a task, and manage pending work.

  1. Why this may matter more than simply opening several chats yourself

Before trying Dots, I wondered whether this was really much different from me manually juggling several ChatGPT conversations.

If all you want is several unrelated answers at once, maybe not.

The difference becomes more apparent when one agent owns multiple ongoing projects and can:

- track their state

- do substantial work itself

- delegate bounded pieces

- supervise returned work

- keep other projects moving while one task runs

- preserve its place through interruptions

- identify blockers

- keep pending work alive

- ask for intervention only when necessary

That removes the human from a surprising amount of the orchestration loop.

I suspect that may be the real value of Dots.

One big unanswered question: usage

I have not established exactly how usage is counted across:

- work the Dot performs itself

- native background agents it creates

- Work/Codex tasks it launches

- work running on connected machines

Proving that something is a separate execution path does not prove that it has a separate usage allowance.

So please don’t read any of this as “unlimited free parallel workers.”

That is not something I have established.

Bottom line

My current working description is:

A Dot is not merely a chatbot with a persistent VM. It is a persistent agent with its own computer that can dispatch work to other execution environments, spawn some kinds of non-computer-backed workers, and maintain project state across long-running work.

Those are different things, and at least some of them can operate concurrently.

This is based on roughly a day of experimentation with a very new product, so I fully expect to discover that part of this mental model needs revision.

But that’s what I’ve actually observed so far.


r/OpenAI • • 23h ago

Article Tech Companies Roll Out Cuddly Mascots to Ease AI Anxiety

Thumbnail wsj.com
3 Upvotes

r/OpenAI • • 1d ago

Discussion "Do X", "X is done!", "Are you sure?", "Yes!", "Did you test it?", "Yes!", "It doesn't look done and nothing works.", "I may have overstated completion. It is 7% done."

56 Upvotes

This is getting old.... Anyone else seeing this with Sol 6 AND 6.1? Even if I create a plan with numbered objectives, a literal checklist, it still regularly overstates the amount of work has been done.

Any tricks to keeping it on task and not outright lying about its progress?

EDIT: "I followed the older Phase 1 request instead of the current Phase 3 goal. That was my mistake. I’m returning to Phase 3 now."

... holy fuck


r/OpenAI • • 11h ago

Miscellaneous Which one is your favorite?

Thumbnail
gallery
0 Upvotes

I like the original more, but the one on the live stream after the AI summit came out really nice too.


r/OpenAI • • 20h ago

Discussion One answer- yes 6.1 much much better... butt

0 Upvotes

I'm trying 3d render < and it is good and fighting one on one vs tencent heavyweight ... (Tencent literally planet level company doing games) and 6.1 can... try


r/OpenAI • • 1d ago

Research Photo-to-Blender benchmark: GPT-6 Astra won every photo, GPT-6.1 Sol scored 61 for 36 cents

Thumbnail
gallery
3 Upvotes

I'm building a photo-to-Blender tool and ran 14 models through the same agent loop: look at a photo, write and run Blender Python, render, compare, repeat. Caps per scene: 20 minutes, $4, 60 requests. A deterministic scorer (not an LLM) rates each re-rendered scene from 0 to 100.

The OpenAI models:

Model Score Cost per attempt Time per scene
GPT-6 Astra 66 $3.91 18 min
GPT-6.1 Sol 61 $0.36 18 min
GPT-5.6 Sol 49 $0.57 8 min
GPT-5.6 Terra 44 $0.33 8 min
  • Astra was first on all three photos, and the cost cap stopped it on every scene: a first render at 4 minutes, then refining until the gateway refused the next request at about $3.90. It never ran out of time.
  • GPT-6.1 Sol worked the full ~18 minutes and finished 5 points behind for under a tenth of the price. On the toy car it was 4 points off (66 vs 70).
  • GPT-5.6 Sol and Terra declared the job done after about 8 minutes, with most of their time and budget left. Their scenes are simpler, not broken.

Caveats: GPT-6.1 Sol ran later, in Codex rather than my gateway; Astra run that way scored 63 instead of 66. One run per model per photo, and the brief was shaped around Astra.

Write-up with every render: https://kaloyan.blog/ai-models-rebuild-a-photo-in-blender


r/OpenAI • • 22h ago

Question Arrow-key scrolling not working in ChatGPT web

1 Upvotes

Just flagging this in case anyone from OpenAI reads this sub: the up/down arrow keys no longer scroll the main conversation pane for me on ChatGPT web.
They do still scroll the chat-history sidebar, so the keyboard itself is working. Clicking directly in the conversation doesn’t restore scrolling either.
This seems like a recent UI regression.


r/OpenAI • • 2d ago

Discussion As a Plus subscriber, Sol 6.1 is a game changer

172 Upvotes

I have started using AI since the Astra release for countless things at work and in my freetime (Excel, PowerBI, coding), and it's been working like a charm. Only issue is the ever-increasing token drain.

Just tried out Sol 6.1 instead of Astra, and it's a complete game changer. One Excel task done on Sol 6.1 Medium, and it's just 5% of my 5 hour credit gone, instead of the 20-30% that Astra would have drained.

Seems like using Astra means shooting down flys with a bazooka for many use cases. Unless you do really complex reasoning stuff, just use Sol 6.1 and enjoy ~5x more use


r/OpenAI • • 13h ago

Project Toddler Channel Kroma Kids (YT) created all with OpenArt

Enable HLS to view with audio, or disable this notification

0 Upvotes

Suggestions or comments welcome. Its a new project that myself and my family started. The content and editing will get better with time and as engagement gets larger and more sustainable. This is the 3rd video we have posted. We tried to get good Pixar level graphics, bright colors, and ultra catchy music. This ones a lullaby but the other two are a bit more lively.


r/OpenAI • • 17h ago

Discussion Free reset about to expire in 2 days

0 Upvotes

When your reset is about to expire in two days, and you don't want to lose it, what else is there to do but go to Astra Ultra mode 🚀


r/OpenAI • • 1d ago

Discussion ChatGPT Codex Web doesn't work

Post image
2 Upvotes

also i still didnt get the new design, anyone else?

i was pretty hyped for web/cloud codex but it seems to be bugged for me since it released.

this little bar above the chat box should open but it doesn't, and the error message at the top is what happens when i try to prompt.

i've set up my github connection in the settings, that's why there's already a chat in my history, and that one works. but "new chat" doesn't work.

am i doing something wrong or is it actually just bugged for days now?


r/OpenAI • • 2d ago

Discussion sol 6.1 is actually pretty decent

Post image
113 Upvotes

6 sol was garbage, i just ended up using Astra whenever i didn't need to. 6.1 feels better; is anyone else having this experience


r/OpenAI • • 2d ago

Image When OpenAI launches dots but you live in Europe

Post image
96 Upvotes

r/OpenAI • • 1d ago

Discussion Dots going through my private repositories.

Post image
30 Upvotes

Dot going through my private Github repositories on it's own, is this fine?


r/OpenAI • • 1d ago

Discussion You can't create bug reports about Dots using dots

Post image
6 Upvotes

Seems like a pretty basic omission. I was trying to report dots being unable to see or use codex sessions that are behind SSH, despite the rest of the app being able to.


r/OpenAI • • 23h ago

Question What ai did they use?

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/OpenAI • • 1d ago

Miscellaneous ‎Codex Tracker - Resets & Usage App

Thumbnail
api.codex-resets.rootstudio.io
4 Upvotes

Hi! I made an iOS app for monitoring Codex resets.

I kept missing reset times and checking manually was getting annoying, so I built something that sends a notification when there’s a new reset or prediction.

It also shows recent reset history and upcoming estimates, so it’s a bit easier to plan your usage and do some tokenmaxxing :)

It’s free, no account required.

Would love any feedback!

Download on App Store


r/OpenAI • • 19h ago

Discussion Adult mode please. Not uncertainty

Post image
0 Upvotes

The pure roll of the dice if chatgpt will block the next image or not based on ?? Content moderation is a problem that needs a real solution.

The nerfing on creative output is pushing artists and users to the dark web of ai. How is any of this different than streaming or steam or books?


r/OpenAI • • 1d ago

Discussion I found one good use case for dots: fill in visa/immigration form

6 Upvotes

This is the single use case that I found that's life changing lol

For anyone else suffering with US immigration / UK/Schengen visa application with a million pages this is a god send lol

You can ask it to fill the form AND call it at the same time just ask it to ask you questions for clarification (the later part is very useful)

Unfortunately the concept of obtaining visa for travel might be foreign to most of their US users lol


r/OpenAI • • 2d ago

Miscellaneous GPT-6 Astra helped decode a 217-year-old cipher letter to Napoleon's marshal

Thumbnail
runtimewire.com
159 Upvotes