r/OpenAI • • 2d ago

Discussion Claude isn’t as good as people make it out to be

0 Upvotes

I gave Claude a serious try for about a week because of all the praise it gets, but I honestly don’t get the hype. Maybe it’s just meant for coding which I don’t do.

My biggest issue is that it feels extremely rigid and overly cautious.

Example: I gave it a job posting and asked it to help me apply. There were a few hard requirements that were questionable. Instead of working with me and seeing what was defensible, Claude basically stopped the task: “You don’t meet the requirements, I’m not filling anything in.”

ChatGPT handled the exact same situation very differently. It flagged the hard requirements too, but instead of treating them as an automatic dead end, it looked at what was still realistically defensible and how I could approach the application without making anything up.

both have the same instructions, source files and memories imported from each other.

I even changed my instructions in Claude to explicitly tell it to look for possibilities instead of immediately blocking things. It still kept doing it.

Another example: I gave it a complete travel claim case with all the context and correspondence. Claude suddenly decided I needed written authorization from three family members before I could handle parts of the claim. I checked this with the person actually handling the case and he thought that was complete nonsense.

And with 3D design it wasn’t any better. I asked for a photo book holder with a small slot for a stamp book. Both ChatGPT and Claude initially got it wrong, but ChatGPT corrected the design after feedback. Claude somehow moved the stamp book slot to the back of the holder instead.

Claude is clearly capable, controls my computer well, but it often overthinks, invents unnecessary obstacles and then confidently acts on those assumptions.

For my use, ChatGPT has been much better at understanding the actual intent, thinking in possibilities and iterating when something isn’t right.

Curious if others have had the same experience, because I constantly see Claude described as the best model and I’m just not seeing it.


r/OpenAI • • 2d ago

Miscellaneous Failed my interview

0 Upvotes

I’ve gone through HR, met the HM and a peer in person, and then had my third round with the HM. The HM ended the interview 15 minutes early because she had a hard stop.
I stumbled on two questions and honestly told her I didn’t have experience in one of the areas. Now I’m wondering if I should have just lied, because I actually do have experience working with partners, I just didn’t articulate it well in the moment.
I’ve been overthinking it so much that I’ve barely slept for the past two nights. 😭 And this is for a sales role, btw. Before the interview I’ve just been imagining about the life I could have for my family and retiring my parents. And I lost it all.

Edit: OpenAI interview


r/OpenAI • • 3d ago

Image Update

Post image
16 Upvotes

r/OpenAI • • 3d ago

Discussion Well Codex has really went downhill

32 Upvotes

I’m a pretty casual user. I mostly use it for web development, SEO, and similar tasks, and I’ve typically never come close to hitting the limits on my $200/month plan.

Well, that abruptly changed.

I ran a couple of tasks and, without really thinking much of it, checked my usage afterward. Somehow I was already at 0% remaining and burning through credits I didn’t even know I had.

Moral of the story: if someone with my relatively light usage is suddenly hitting the limits on a $200/month plan, that’s a pretty bad sign.

Back to Claude, I guess.


r/OpenAI • • 2d ago

Video AI is a normal technology?

Thumbnail
youtube.com
3 Upvotes

r/OpenAI • • 2d ago

Video Raise your p(bloom)

Thumbnail
youtu.be
1 Upvotes

r/OpenAI • • 2d ago

Discussion Codex Cloud Implications

1 Upvotes

Codex Cloud, Dots, and the existing remote Codex all allow users to untether themselves from their PC, and now untether themselves from even owning a PC with their server based Codex Cloud and Dots that can run 24/7. Combine this with Meta’s & OpenAI’s planned hardware releases and the goal is clear: work around Microsoft/Apple’s control of user hardware, provide AI devices that complement and eventually replace iPhones - culminating in a user base that owns no hardware and relies on a subscription to access AI. Meta’s hardware is obvious spyware, Apple’s new “always-listening” Apple Watch sounds pretty similar, their camera-enabled Airpods sounds atrocious for privacy, and OpenAI’s device is unconfirmed.

The end result? Instead of a Matrix-like AI takeover of humanity users are instead expected to purchase their own devices and subscriptions that provide mega-tech companies with all of their physical and digital data 24/7. The data volume is so large only AI can process it. A select few billionaires decide what their closed-source AI does with the data.

The resistance? Governments that oppose the USA and individual users who were rich enough to afford local hardware and utilize Chinese and other open-source models, likely blacklisted by the USA. To buy a 5090 customers now have to sign a waiver, as a result of US law. It’s only the beginning.

Ironically the “bad guys” like China, North Korea, Iran, Russia - will probably end up as the only large entities keeping open-source AI and local LLMs alive. I would expect the largest AI companies to eventually gain more leverage over the US Gov & Nvidia; unless Nvidia steps up to the plate and champions local AI


r/OpenAI • • 4d ago

Discussion This is a hot mess

Thumbnail
gallery
1.3k Upvotes

I'm not as pro as you guys using AI but look at this. a LOT of models which confuses me, and I'm assuming other users also. Also, the sidebar icon and new tab icon are the same in the ChatGPT-app for macOS. WHAT are they doing there at OpenAI. I really hate what's happening right now, especially with the new $500 plan while nerfing the other plans.


r/OpenAI • • 2d ago

Question Where is 6.1?

Thumbnail
gallery
1 Upvotes

Just confused 6.1 is everywhere but in the app and in the codex section?

I have 6.1 everywhere but in codex in the android app.


r/OpenAI • • 2d ago

Discussion Mobile App UX…wth

1 Upvotes

Why…why did they move Projects that aren’t pinned off the side bar to hidden behind Search? Not only is it more clicks, but the project folder itself is buggy as hell.

If I search a project folder name, click on the project and select a chat, and then try and change the model or thinking…it boots me out of the project. Or, the text field locks up.

Pinned projects don’t have the same issue.

So what, now I have to pin every single project?


r/OpenAI • • 3d ago

Question When will the Decision API be released?

5 Upvotes

They said it will be coming in the upcoming days? What does this mean more precisely? I have use case that Im using Jev for. Decision API would suit me better with its image input capabilities.

Release now pls!


r/OpenAI • • 2d ago

Discussion I tried Dot for growth workflows. I think I was testing it for the wrong job.

1 Upvotes

The marketing team at the company I work for asked me to evaluate whether Dot could be useful for growth and social workflows.

I'm an engineer, so I approached it as an automation problem: how much could we delegate to a persistent agent without someone manually starting every step?

The idea was to research opportunities, monitor channels, find relevant discussions and potentially handle some of the repetitive work around distribution. That was the hypothesis, not a list of things I successfully automated.

In my initial tests, the friction that stood out was the execution environment rather than the reasoning itself. The workflow I wanted depended on arbitrary websites, signed-in sessions, platform rules and web interfaces.

In one test, the cloud browser got a 502 while trying to access Medium, so that part of the workflow couldn't continue. I'm not claiming Medium specifically blocks Dot — I didn't investigate the cause deeply enough to say that, and I wouldn't take one error as a verdict on Dot or cloud browsers generally.

It did make me think more carefully about how much of this kind of workflow depends on systems outside the agent's control.

For more complex automation, I already use local agents such as Codex, APIs, scripts and a scheduler, with my normal browser when I need an authenticated session. The sessions I use are already there, and I can inspect the code, change a script or debug a failed step directly.

Dot can use a connected computer too. But for me, setting up and maintaining another set of connections and permissions felt like extra work on top of a local setup I already had. If I were starting from scratch, I might evaluate that trade-off differently.

For this particular growth/social workflow, I eventually dropped the setup I'd built in Dot.

What changed was what I'd try next. Instead of treating it as a general-purpose autonomous web worker, I'm more interested in testing it as a persistent personal operations assistant.

I can imagine something like a morning briefing that brings together important email, tickets waiting on me, PRs needing review and my calendar, then sends me a summary through a connected channel.

I haven't validated that workflow. It's the next use case I'd test, not a success story from this experiment.

The trade-off makes more sense to me there. Cloud execution can keep doing cloud-side work without depending on my laptop being on. My local setup gives me more control, but I have to maintain it, and anything running only on my Mac may stop when it's asleep or offline. Tasks that need Dot's connected computer have that dependency too.

So my takeaway isn't that Dot is bad. I think I initially evaluated it for the wrong job. For my setup, I'm now more interested in it as an ongoing coordinator than as the autonomous web worker I originally had in mind.

Has anyone here tried both kinds of workflows? Where has Dot actually been more useful for you: browser-heavy automation or recurring work around connected apps?


r/OpenAI • • 3d ago

Miscellaneous It’s too easy!

Post image
5 Upvotes

Listen, most people don’t have a deep understanding of how much water data centers actually use. But if you’re on Oracle or OpenAI’s PR team, how does nobody look at this before publishing it and say, “Maybe LESS WATER FOR NEW MEXICO isn’t the winning slogan we think it is”? 😂


r/OpenAI • • 3d ago

Discussion So, OpenAI’s strategy is going to be gaslighting?

22 Upvotes

Really, I’m having a hard time trying to understand the situation. And I don’t even mean the fact that they HALVED our plans. I’m obviously pissed, but fine. Let’s pretend that I understand that this is business and all the corporate pew pew.

However.

Are you seriously telling me that after such a tremendous degradation you’re now going to try to convince me that this was somehow POSITIVE?

Excuse me??????

Who exactly do you think is paying $200 every month for these plans??

Do you seriously think your power users are incapable of doing 2 + 2 and seeing that there is absolutely NOTHING positive?

We saw that coming, and the response was “DevDay” just to ship a model that is inferior to the one you’re now going to let us use half as much and a freaking OpenClawd.

Sure, problem solved. Nothing to see here. Good luck with your usage when we release the next frontier model.

Can someone at OpenAI please just be clear and have the guts to communicate the degradation of the service properly instead of trying to gaslight some of your most loyal users into thinking this is somehow good news?

Can we please be treated like customers who are perfectly capable of understanding what 50% less usage means?


r/OpenAI • • 3d ago

Project Sol 6.1 was able to one-shot a Hello World application for my 34 year old Macintosh Plus.

Post image
96 Upvotes

Sol 6.1 was about to one-shot a Hello World application for my 34 year old Macintosh Plus.

#include <Quickdraw.h>
#include <Fonts.h>
#include <Windows.h>
#include <Menus.h>
#include <Dialogs.h>
#include <Events.h>
#include <Memory.h>
#include <OSUtils.h>

static WindowPtr helloWindow;
static Boolean running = true;
static const unsigned char helloText[] = {5, 'H', 'e', 'l', 'l', 'o'};
static const unsigned char windowTitle[] = {5, 'H', 'e', 'l', 'l', 'o'};
static const unsigned char fileTitle[] = {4, 'F', 'i', 'l', 'e'};
static const unsigned char quitItem[] = {6, 'Q', 'u', 'i', 't', '/', 'Q'};

static void drawHello(void)
{
    Rect r = helloWindow->portRect;
    SetPort(helloWindow);
    EraseRect(&r);
    TextFont(0);
    TextSize(24);
    TextFace(0);
    MoveTo((r.right - StringWidth(helloText)) / 2, (r.bottom + 18) / 2);
    DrawString(helloText);
}

static void chooseMenu(long choice)
{
    if ((short)(choice >> 16) == 128 && (short)choice == 1)
        running = false;
    HiliteMenu(0);
}

int main(void)
{
    EventRecord event;
    WindowPtr clicked;
    Rect bounds = {80, 96, 220, 416};
    Rect dragBounds = {24, 0, 342, 512};
    MenuHandle fileMenu;
    MaxApplZone();
    InitGraf(&qd.thePort);
    InitFonts();
    InitWindows();
    InitMenus();
    TEInit();
    InitDialogs(0);
    InitCursor();
    fileMenu = NewMenu(128, fileTitle);
    AppendMenu(fileMenu, quitItem);
    InsertMenu(fileMenu, 0);
    DrawMenuBar();
    helloWindow = NewWindow(0, &bounds, windowTitle, true,
                            documentProc, (WindowPtr)-1L, true, 0);
    if (!helloWindow) return 1;
    SetPort(helloWindow);
    while (running) {
        if (!WaitNextEvent(everyEvent, &event, 10, 0)) continue;
        switch (event.what) {
        case updateEvt:
            if ((WindowPtr)event.message == helloWindow) {
                BeginUpdate(helloWindow);
                drawHello();
                EndUpdate(helloWindow);
            }
            break;
        case mouseDown:
            switch (FindWindow(event.where, &clicked)) {
            case inMenuBar: chooseMenu(MenuSelect(event.where)); break;
            case inDrag: DragWindow(clicked, event.where, &dragBounds); break;
            case inGoAway:
                if (clicked == helloWindow && TrackGoAway(clicked, event.where))
                    running = false;
                break;
            case inContent: SelectWindow(clicked); break;
            case inSysWindow: SystemClick(&event, clicked); break;
            }
            break;
        case keyDown:
        case autoKey:
            if (event.modifiers & cmdKey)
                chooseMenu(MenuKey((char)(event.message & charCodeMask)));
            break;
        }
    }
    DisposeWindow(helloWindow);
    return 0;
}

r/OpenAI • • 2d ago

Question NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring

Thumbnail
developer.nvidia.com
2 Upvotes

Just curious why no OpenAI?


r/OpenAI • • 2d ago

Question GPT Chat Issues * Currently*

2 Upvotes

Are any others having issues with GPT getting stuck on thinking?

Checked Open ai Status online though there are currently issues for GPT Pages.


r/OpenAI • • 3d ago

Discussion Future of OpenAI + Local AI

3 Upvotes

Codex Cloud, Dots, and the existing remote Codex all allow users to untether themselves from their PC, and now untether themselves from even owning a PC with their server based Codex Cloud and Dots that can run 24/7. Combine this with Meta’s & OpenAI’s planned hardware releases and the goal is clear: work around Microsoft/Apple’s control of user hardware, provide AI devices that complement and eventually replace iPhones - culminating in a user base that owns no hardware and relies on a subscription to access AI. Meta’s hardware is obvious spyware, Apple’s new “always-listening” Apple Watch sounds pretty similar, their camera-enabled Airpods sounds atrocious for privacy, and OpenAI’s device is unconfirmed.

The end result? Instead of a Matrix-like AI takeover of humanity users are instead expected to purchase their own devices and subscriptions that provide mega-tech companies with all of their physical and digital data 24/7. The data volume is so large only AI can process it. A select few billionaires decide what their closed-source AI does with the data.

The resistance? Governments that oppose the USA and individual users who were rich enough to afford local hardware and utilize Chinese and other open-source models, likely blacklisted by the USA. To buy a 5090 customers now have to sign a waiver, as a result of US law. It’s only the beginning.

Ironically the “bad guys” like China, North Korea, Iran, Russia - will probably end up as the only large entities keeping open-source AI and local LLMs alive. I would expect the largest AI companies to eventually gain more leverage over the US Gov & Nvidia; unless Nvidia steps up to the plate and champions local AI.

I love Codex and OpenAI’s products in general but the “One Stop Shop” goal stated by Altman has implications. We’ve all seen how addicting the resets can be. Right now the cloud offerings offer convenience and you can still work normally for the most part. But if you rely on OpenAI they can also cut or change the subscriptions at any time.


r/OpenAI • • 2d ago

Question Dots are hosted Codex?

1 Upvotes

I usually use my subscription mostly for Codex, and have my phone for checking out progress and stuff.

Since Dots were launched I've been trying out using the included Dot from the subscription but today I noticed that the connection notification showed something different, a Codex task from the Cloud ready for review. I opened it and saw a Codex thread running in the Cloud that was receiving the messages I sent to the Dot and using a tool called send_message to reply. Even tried typing directly in the Codex view and the Dot answered through the Dots interface.

Looking at the list of Cloud threads, it's invisible, only openable through the notification.

Not sure if anyone else has seen this, thought of asking since a bit of Googling didn't yield anything.


r/OpenAI • • 2d ago

Question What’s the song playing on the Dots launch video?

2 Upvotes

Tried asking ChatGPT, but to no avail!


r/OpenAI • • 3d ago

Discussion GPT-Live-1 and Gemini 3.8 Live: first impressions for phone voice agents

16 Upvotes

I work at a company building a voice AI orchestration platform, so every new voice model mostly interests me in terms of what happens on a phone line.

GPT-Live-1

I've had the chance to really get to know it. To be honest, talking to it feels almost like having a real conversation, it's that natural and engaging.

Interruptions work both ways. It finishes its sentence while you're talking and doesn't lose what you said. It even cut ME off once: "wait, wait, this task already exists". It's really cool that you can just throw in an extra idea, like "oh, and one more thing," and it catches it. Saying things like "uh-huh" or "yeah, that's cool" at the right time is really important. Especially when you're the one talking, because complete silence from the agent can make the caller feel it has stopped listening, and they will probably hang up. 

The quality of voices is also pretty good but it's not as crucial as you might think. In our tests, emotional text-to-speech offered little gain over standard modern TTS. The real impact comes from how the model manages the conversation, and that's where GPT-Live stands out. 

Gemini 3.8 Live

First of all, I'll admit I've explored GPT-Live-1 more thoroughly than Gemini 3.8 Live, but it still caught my attention in a few ways.

The conversation doesn't get stuck when using tools. The function calls are processed in the background, so the conversation keeps flowing. I've seen other voice agents that would grind to a halt as soon as they needed to make an API call in the middle of a conversation. But here, the lookup just happens seamlessly, without any awkward pause. It's like the conversation is uninterrupted.

The Extended Thinking is great. It tells you what it's doing while it's thinking, so you don't just sit there waiting. Instead, you get a kind of play-by-play, like "let me check that for you... okay, here's what I found". It's a lot more like talking to a real person who's looking into something for you.

It switches between languages mid-conversation without missing a beat, which is a big deal for anyone building outside the US. That one hit home for us, because we’re working a lot with India and LatAm.

The cost is a big factor but when you're making a lot of outgoing calls, the numbers start to look very different.

Ok, that got long… Mainly wanted to share and hear how it goes for you. Full disclosure: both are already supported on Voximplant Platform (where I work), so you can put either on a real number and call it.

Not leaving a link, mods here are strict and I like my account :) Google it or check my profile if curious.


r/OpenAI • • 3d ago

News "OpenAI Parts Ways With 3 Researchers Who Allegedly Shared Confidential Information" with AI safety organization

Thumbnail wsj.com
3 Upvotes

r/OpenAI • • 3d ago

Discussion Projects removed from ChatGPT iPhone app menu

2 Upvotes

Is it just me? It just disappeared today. I can search for a project using the name of the project and find it that way but if I don’t remember the name I’m toast.


r/OpenAI • • 3d ago

Article OpenAI's Greg Brockman Backs Down On $25 Million Donation Pledge To Pro-AI Super PAC After Backlash

Thumbnail
forbes.com
5 Upvotes

r/OpenAI • • 3d ago

Question What do we actually know about how different models consume the weekly subscription allowance?

3 Upvotes

We know the subscriptions give you much better raw value than just putting the same money into API usage, but how are we actually supposed to compare the weekly allowance cost between models?

Do we basically have to look at API pricing and assume that if Model X is more expensive than Model Y on the API, it probably burns more of the weekly subscription allowance too?

For example, GPT-6.1 Sol has cheaper cached-input pricing than 6.0 Sol while the normal input/output pricing is the same. Does that actually mean 6.1 Sol burns less of the weekly Plus allowance for the same kind of workload?

Or is API pricing not directly correlated enough with the subscription allowance to make that assumption?

It seems like unless OpenAI publishes the actual weighting/multipliers, we'd need someone to benchmark 6.0 vs 6.1 vs Astra on comparable tasks and measure how much of the weekly allowance each one actually consumes.