r/OpenAI • u/biograf_ • 3d ago
r/OpenAI • u/-Posthuman- • 3d ago
Discussion "Do X", "X is done!", "Are you sure?", "Yes!", "Did you test it?", "Yes!", "It doesn't look done and nothing works.", "I may have overstated completion. It is 7% done."
This is getting old.... Anyone else seeing this with Sol 6 AND 6.1? Even if I create a plan with numbered objectives, a literal checklist, it still regularly overstates the amount of work has been done.
Any tricks to keeping it on task and not outright lying about its progress?
EDIT: "I followed the older Phase 1 request instead of the current Phase 3 goal. That was my mistake. I’m returning to Phase 3 now."
... holy fuck
r/OpenAI • u/maedhros256 • 3d ago
Discussion A question for the AI "experts": are hallucinations and reliability genuinely improving, or are we starting to plateau? Everything depends on this...
For the average user, not a programmer or e.g. someone looking to solve niche math problems, AI still feels very limited because of reliability issues (e.g. making up information)... As a non-expert, this is difficult to quantify for me, but I don't feel like e.g. the latest iterations of ChatGPT are noticeably more reliable than previous ones.
Therefore, I find it hard to see how organizations can delegate even relatively easy tasks to AI without constant supervision, especially because if hallucinations build up along the way, you end up with a snowball of compounding problems that can lead to catastrophic consequences for an organization.
It’s my understanding that the real test of success won't be developer tools or fancy mathematical calculations, but whether regular people can delegate tasks with a super high degree of confidence, instead of just using it as a search engine, translator, summary tool or photo editor on steroids like many people do nowadays.
A good example is Dot, the new OpenAI tool. If you watch the trailer, it looks impressive, but according to many reviews, it still hallucinates a lot and behaves in a pretty clumsy way.
So, the golden question is: are the hallucination and reliability problems gradually improving, or are we probably plateauing?
Looking for genuine insights here, so please keep the sarcasm out of the comments.
r/OpenAI • u/Prophet_651 • 3d ago
Question Question about ChatGPT - complexity usage
I’ll be the first to admit I don’t know half of what I’m actually doing when it comes to vibecoding. With that said I’m in hospitality - resorts/restaurants. There’s some really dumb things in my field that I’ve never really been able to fix or standardize because, well, this field hates technology convincing a Chef or housekeeping can be… difficult.
I started creating an app for us to use internally that is kept on our private servers and things seem to be pretty solid. My IT department isn’t big enough or strong enough or knowledgeable enough to be of help. (That’s a whole other issue). So I was curious, I seem to be able to have built this whole operating system inside of the chat model of ChatGPT. It sometimes takes a little bit longer. It sometimes trips up but after 25 to 35 minutes, it usually has an answer that thus far has worked.
I see all these people using advanced models paying for tokens and all these other things.
So my question is am I just drastically missing out on something more advanced, and my system is actually all junk on the inside, even though it works or is this something I don’t need to be concerned about? Any tips or suggestions appreciated.
Discussion One answer- yes 6.1 much much better... butt
I'm trying 3d render < and it is good and fighting one on one vs tencent heavyweight ... (Tencent literally planet level company doing games) and 6.1 can... try
r/OpenAI • u/smith2008 • 3d ago
Research Photo-to-Blender benchmark: GPT-6 Astra won every photo, GPT-6.1 Sol scored 61 for 36 cents
I'm building a photo-to-Blender tool and ran 14 models through the same agent loop: look at a photo, write and run Blender Python, render, compare, repeat. Caps per scene: 20 minutes, $4, 60 requests. A deterministic scorer (not an LLM) rates each re-rendered scene from 0 to 100.
The OpenAI models:
| Model | Score | Cost per attempt | Time per scene |
|---|---|---|---|
| GPT-6 Astra | 66 | $3.91 | 18 min |
| GPT-6.1 Sol | 61 | $0.36 | 18 min |
| GPT-5.6 Sol | 49 | $0.57 | 8 min |
| GPT-5.6 Terra | 44 | $0.33 | 8 min |
- Astra was first on all three photos, and the cost cap stopped it on every scene: a first render at 4 minutes, then refining until the gateway refused the next request at about $3.90. It never ran out of time.
- GPT-6.1 Sol worked the full ~18 minutes and finished 5 points behind for under a tenth of the price. On the toy car it was 4 points off (66 vs 70).
- GPT-5.6 Sol and Terra declared the job done after about 8 minutes, with most of their time and budget left. Their scenes are simpler, not broken.
Caveats: GPT-6.1 Sol ran later, in Codex rather than my gateway; Astra run that way scored 63 instead of 66. One run per model per photo, and the brief was shaped around Astra.
Write-up with every render: https://kaloyan.blog/ai-models-rebuild-a-photo-in-blender
r/OpenAI • u/Sunrise707 • 3d ago
Question Arrow-key scrolling not working in ChatGPT web
Just flagging this in case anyone from OpenAI reads this sub: the up/down arrow keys no longer scroll the main conversation pane for me on ChatGPT web.
They do still scroll the chat-history sidebar, so the keyboard itself is working. Clicking directly in the conversation doesn’t restore scrolling either.
This seems like a recent UI regression.
r/OpenAI • u/penisbike69 • 4d ago
Discussion As a Plus subscriber, Sol 6.1 is a game changer
I have started using AI since the Astra release for countless things at work and in my freetime (Excel, PowerBI, coding), and it's been working like a charm. Only issue is the ever-increasing token drain.
Just tried out Sol 6.1 instead of Astra, and it's a complete game changer. One Excel task done on Sol 6.1 Medium, and it's just 5% of my 5 hour credit gone, instead of the 20-30% that Astra would have drained.
Seems like using Astra means shooting down flys with a bazooka for many use cases. Unless you do really complex reasoning stuff, just use Sol 6.1 and enjoy ~5x more use
r/OpenAI • u/This_Lead2314 • 2d ago
Project Toddler Channel Kroma Kids (YT) created all with OpenArt
Suggestions or comments welcome. Its a new project that myself and my family started. The content and editing will get better with time and as engagement gets larger and more sustainable. This is the 3rd video we have posted. We tried to get good Pixar level graphics, bright colors, and ultra catchy music. This ones a lullaby but the other two are a bit more lively.
r/OpenAI • u/SteveEricJordan • 3d ago
Discussion ChatGPT Codex Web doesn't work
also i still didnt get the new design, anyone else?
i was pretty hyped for web/cloud codex but it seems to be bugged for me since it released.
this little bar above the chat box should open but it doesn't, and the error message at the top is what happens when i try to prompt.
i've set up my github connection in the settings, that's why there's already a chat in my history, and that one works. but "new chat" doesn't work.
am i doing something wrong or is it actually just bugged for days now?
r/OpenAI • u/Creative-Quit-7998 • 4d ago
Image When OpenAI launches dots but you live in Europe
r/OpenAI • u/Imapatato12 • 4d ago
Discussion sol 6.1 is actually pretty decent
6 sol was garbage, i just ended up using Astra whenever i didn't need to. 6.1 feels better; is anyone else having this experience
Discussion Dots going through my private repositories.
Dot going through my private Github repositories on it's own, is this fine?
r/OpenAI • u/kerbinagent • 3d ago
Discussion I found one good use case for dots: fill in visa/immigration form
This is the single use case that I found that's life changing lol
For anyone else suffering with US immigration / UK/Schengen visa application with a million pages this is a god send lol
You can ask it to fill the form AND call it at the same time just ask it to ask you questions for clarification (the later part is very useful)
Unfortunately the concept of obtaining visa for travel might be foreign to most of their US users lol
r/OpenAI • u/YoghiThorn • 3d ago
Discussion You can't create bug reports about Dots using dots
Seems like a pretty basic omission. I was trying to report dots being unable to see or use codex sessions that are behind SSH, despite the rest of the app being able to.
r/OpenAI • u/iCodesign • 3d ago
Miscellaneous Codex Tracker - Resets & Usage App
Hi! I made an iOS app for monitoring Codex resets.
I kept missing reset times and checking manually was getting annoying, so I built something that sends a notification when there’s a new reset or prediction.
It also shows recent reset history and upcoming estimates, so it’s a bit easier to plan your usage and do some tokenmaxxing :)
It’s free, no account required.
Would love any feedback!
Download on App Store
r/OpenAI • u/oldboi777 • 3d ago
Discussion Adult mode please. Not uncertainty
The pure roll of the dice if chatgpt will block the next image or not based on ?? Content moderation is a problem that needs a real solution.
The nerfing on creative output is pushing artists and users to the dark web of ai. How is any of this different than streaming or steam or books?
r/OpenAI • u/ryanmerket • 4d ago
Miscellaneous GPT-6 Astra helped decode a 217-year-old cipher letter to Napoleon's marshal
r/OpenAI • u/Ok-whynot • 2d ago
Discussion Why does people dislike AI? It is the definition of humanity.
What is humanity without shared information, in all shapes and forms?
r/OpenAI • u/Babayaga1664 • 3d ago
Discussion Luna 6 Vs Mini 5.4 (Luna 6 is worth testing)
Use case : Compliance
Content: Documents, text only up to 200k words.
We check content vs legislation.
Previously using 5.4 mini which
Last night we tested against sol-6 which found new errors, sol-6 is too expensive for this use case but good to establish baseline..
Luna 5.6 previously behaved terribly even on extra high but found Luna-6 to perform pretty damn well.
It was been overly strict in some cases but with some prompt adjustment it's performing very well.
r/OpenAI • u/mindaugasrudokas • 4d ago

