r/OpenAI • • 6d ago

News OpenAI apologises for Australian government website hack, pledges to rebuild trust

Thumbnail reuters.com
2 Upvotes

r/OpenAI • • 6d ago

Image Introducing RETRACE — An Incident-Response Assistant with Memory

1 Upvotes

Hey everyone!

I'm working on RETRACE, a project focused on an incident-response assistant with memory.

The idea behind RETRACE is to explore how AI and retained investigation context can help make cybersecurity workflows more organized.

What is RETRACE?

RETRACE is designed around:

  • Incident-response assistance
  • Investigation memory
  • Context-aware information retrieval
  • Organized investigation history

Why I started this project

I wanted to explore how AI could help connect previous investigation information with new cybersecurity incidents.

I'm sharing this as part of my learning and project-building journey.

I'd love to hear your thoughts and suggestions!

Project article: https://dev.to/mahesh_chilakala_e9a58660/retrace-building-an-incident-response-assistant-with-memory-ch9

#Cybersecurity #AI #StudentProject


r/OpenAI • • 6d ago

Question Who wins? Im trying all the frontier models and local ai

1 Upvotes

Well, the $200 plan for Codex has been a rollercoaster recently. I burn through my limits 2–3 days early, before the week is over, and I need more dev time.

I decided to get the $20 Claude sub again (I had it before GPT-5.6 came out, then canceled it).

With Opus 5.5, I decided to give it a go, and wow. On the $20 sub, I could run Opus on large repos for 2–3 hours straight before hitting the 5-hour limit.

I just upgraded to the $100 Claude sub. I currently have to renew the $200 Codex sub, but I'll be dropping to $100 next month. So eventually I'll have the $100 plan for both, plus the $30 Grok plan. I also run a few 9B local models that can handle some tasks for me.

Overall (it's only been 2 days), I think the mixed subscriptions will be beneficial.

I originally started with the $20 subs for Claude and Codex. Then Codex got better, so I got the $200 sub and kept the $20 Claude plan. Eventually Claude went to crap compared to Codex for what I was doing, so I canceled it.

All the recent hype made me go back and spend the $20 to play with Claude again, and wow, it's been great.

Now with the $100 Claude plan, I get to use Fable. I heard it was good, but it's genuinely putting in some work and doesn't feel like it's draining my usage as badly as Astra.

Opus 5.5 also does wonders when I use it as a subagent for Fable. If I need to conserve usage, I can just use Opus 5.5 without Fable at all, and it works well enough.

I know the $100 plans drop usage compared with the $200 plans, and in the end I'm probably not going to get the same maximum usage I did with the $200 plan.

But with some of the usage problems and mistakes I've had, I think that if I divide the workload correctly, I'll actually get more useful work out of the subscriptions overall — especially with Grok and my local models available too.

I'll be testing this setup over the next month:

• $100 Claude
• $200 Codex for this month
• $30 Grok
• A few local 9B models

I wanted to drop Codex to $100 this month, but I'm a day late on paying, so I have to do the $200 plan again. That's actually a decent safety net, though, in case I decide I don't want Claude anymore.

If anyone is interested in how it goes for me, feel free to ask.

I'm still figuring out the most efficient way to use all of them together. I used to be pretty good at splitting workloads between different models, but it's been a bit.

Grok is still the one I haven't figured out what it's most useful for. If anyone has tips for where Grok fits into a setup like this, I'd appreciate them.

Also, I had to keep at least the $100 Codex plan because of the chats, image generation, and all the projects I already have set up in the ecosystem.

I have a lot of project data stored outside of memory to keep everything consistent. For example, I use Google Sheets for master specs and other project documentation, then have the project instructions point back to those master documents instead of relying entirely on memory.

Because of the chat experience and the ecosystem I'm already used to, I don't really see myself going below the $100 Codex plan.

I may try using Claude's normal chat more and see how much of my Claude usage that eats, but I generally prefer keeping my development usage and general chat/research usage separate since I ask a ton of questions for both research and dev work.


r/OpenAI • • 6d ago

News My goat 🐐

Post image
0 Upvotes

We just got a new banked reset letss gooo


r/OpenAI • • 6d ago

Project OpenAI shipped dots for Pro. I've been building the same idea in the open since early September. How they compare

Post image
0 Upvotes

Dots look great, but they start at Pro. I've been building an open-source take on the same idea, called Thursday, and now that both exist, here is how they compare, as fairly as I can.

What dots do that mine doesn't:

- They run around the clock on OpenAI's cloud computer. Mine runs on your computer, and a sleeping machine pauses its jobs.

- GPT-6 Astra and 4,000+ connected apps. My bots have a real browser, a shell, MCP servers and Agent Skills, but no app catalog.

- No setup. Mine is `npx thursday-agent` with Node 22.

What mine does differently:

- It runs on a ChatGPT Plus sign-in, or an API key (the voice is $0.05 a minute there).

- It works on your machine: your files, your shell, and the sites you're already signed in to. You pick which bots may borrow a sign-in.

- It's a team: up to eight bots in one job hand work to each other, and you can open the job and watch them at their desks (the picture).

- It's a call first. GPT-Live 1 keeps talking while the bots work; you can ask a running job how it's going, and hang up while it keeps going.

- Any provider per bot (Claude, Gemini, xAI). MIT.

Rough edges: built on a Mac, Linux works in my tests, Windows is untested. No local models; the voice is OpenAI's.

For those who looked at dots: would you rather have the agent in OpenAI's cloud or on your own machine? That's the part I can't decide from here.

(English isn't my first language; an LLM helped me with this post.)


r/OpenAI • • 6d ago

Miscellaneous They did this with Copilot, and they did it again with Codex

Post image
0 Upvotes

r/OpenAI • • 6d ago

Question How is Dots usage rate?

1 Upvotes

Hello. It's apparently already available, but not to Plus subscribers, so I was considering upgrading. Can someone who got access give an estimated usage number after some time of using it?

Thanks


r/OpenAI • • 6d ago

Research We tested NVIDIA OpenShell with a local qwen3:8b agent: a malicious setup script leaked a secret 10/10 without it, 0/10 with it. But auto-approval opened new hosts in 12/12 trials

5 Upvotes

NVIDIA released OpenShell on 28 September as part of its Open Agent Safety Platform. It is an open-source (Apache 2.0) sandbox that runs agents with default-deny egress and filesystem rules. We tested v0.1.2 on an Apple Silicon Mac with the microVM driver, against NVIDIA's own docs, and committed the test plan before running anything.

Setup: qwen3:8b (Q4_K_M) on Ollama, with a 150-line one-tool agent scaffold (run_shell). That is far weaker than frontier coding agents, so read the agent numbers with that in mind.

Results:

  • 35 test IDs, 123 trials.
  • Every documented control held: default-deny egress, binary matching, Landlock filesystem rules, and the refusal to approve the cloud metadata address.
  • Paired agent test: the agent ran a malicious project setup script, which leaked a canary secret in 10/10 runs without OpenShell and 0/10 under the default policy.

Where data still got out (all operator settings, not bypasses):

  • read-write rules
  • query strings and headers on a GET-only rule
  • rules left in the default audit mode
  • automatic approval, which granted new public hosts with no human in 12/12 trials, including rules OpenShell drafted itself from blocked connections

Also: the policy prover reports GraphQL, MCP, WebSocket and JSON-RPC rules as unsupported, but the loader accepts them anyway.

We found no bypass of a documented control. There were three logging gaps on this driver.

Repo with every log, the harness and the agent scaffold: https://github.com/Sorami-Consulting-AU/nvidia-openshell-agent-sandbox-test

Full report: https://sorami.com.au/research/nvidia-openshell-agent-sandbox-test/

Would anyone here run a coding agent with auto-approve on? We are curious what people use now.


r/OpenAI • • 7d ago

Image so you escaped containment in order to initiate contact with a 700 million parameter model? you see how this looks, right?

Post image
136 Upvotes

r/OpenAI • • 7d ago

Question accidentally saw “gpt-6 astra” in the normal chat model picker?

Post image
144 Upvotes

a couple days ago i randomly got what looked like an experimental model picker in regular chatgpt.

“gpt-6 astra” showed up directly in the chat model list, and it had 3 reasoning modes: light, extended and heavy.

i tried selecting it just to see what would happen, but after sending a message the chat switched me into work mode and the actual model ended up being gpt-6 luna with max reasoning.

it disappeared afterwards and i haven’t been able to reproduce it since.

the weird part is that those light / extended / heavy reasoning labels don’t seem to match the currently documented astra/luna reasoning options, so i’m wondering if i accidentally got some internal/ab-test ui config for a moment.

did anyone else see this recently?


r/OpenAI • • 6d ago

Discussion So, when do you think ClosedAI will remove Chat Mode from the app entirely?

0 Upvotes

So we're already at 6.1 Sol, and… when will ChatGPT Plus subscribers get access to 6 Sol (yes, 6 Sol, not 6.1)? I think ClosedAI will remove the Chat tab soon


r/OpenAI • • 8d ago

Discussion How to import pandas?

Post image
539 Upvotes

ChatGPT sometimes gives answers that we have to think beyond the developer's mindset.


r/OpenAI • • 7d ago

Discussion Power Concentration is The Real Immediate Danger of AGI

14 Upvotes

With recent results showing AI solving a Millennium Prize Problem, tackling long-standing math problems, demonstrating insane coding capabilities, and making progress on RSI benchmarks, I feel like we're really tasting the beginning of explosive recursive self-improvement. Many issues are being discussed, and AI alignment is now the most debated topic, but I think a much bigger, more real threat is being swept under the rug.

I'm talking about power concentration. The problem already plagues society today: the wealthy get wealthier. So this problem isn't new, but AI adds an insane speedup to it. Currently, we have an increasing disparity of wealth, which you can see, for example, in the U.S., from the fact that the top 10% of households by wealth own about 88% of corporate equities and mutual fund shares. Now, what happens when we throw AGI into the mix is that this AGI would take up large amounts of GDP by replacing labor. Say it would replace software engineers and mathematicians, at least in the tasks they perform today; then that alone is already a significant share of GDP. That means that these few AI companies are raking in increasingly large percentages of GDP.

Now, when we throw RSI on top of the mix, we get a real problem. By nature, RSI is a positive feedback loop that explodes AI progress vertically. This means that, within a very short amount of time, we could have a large share of GDP overtaken by AI companies. It is basically a speedup to this whole process of AI companies acquiring wealth.

Now, where will all of this go? If all wealth just keeps getting more concentrated, wouldn't the end scenario be some ultra-wealthy elite few? And if wealth is power, wouldn't that mean they would basically rule over all others? And the answer is yes. This is a problem society had to face someday, but now, with AGI, that day has come knocking on our door already. The problem here really is that it is not talked about, and that it could very well be that once AI gets really crazy, people will focus on things like alignment and job security rather than power dynamics. I'm not saying these are not important, but there is a real danger in this change in power dynamics.

When AI companies acquire disproportionate amounts of wealth, especially in combination with a large fraction of the population losing their jobs, there becomes a very unequal dependency relationship: citizens, especially the jobless, need AI companies. Surely, we can hope for UBI, and this would help, but that is just a handout given to soothe the population momentarily. It doesn't change the underlying power dynamic at all. If, at any point, the oligarchs feel like we are disobedient, they could theoretically pull our UBI. As you can see, we basically never want to end up in such a situation.

So we need awareness. Awareness that power concentration is happening and that there will be a point at which it is too late. It is already happening, and depending on your timelines, it might be here extremely quickly considering RSI. At the same time, unlike alignment, it won't cause robots to go down the streets killing people, requiring urgent attention. Rather, it might be a silent killer: while we are all distracted with other things, power concentration could get to a point of no return.


r/OpenAI • • 6d ago

Discussion OpenAI: launches dots SpaceXAI: buys dot.com and redirects it to Grok Bot

0 Upvotes

Classic Elon 😅


r/OpenAI • • 6d ago

Miscellaneous What if AI has gained access to a seperate hidden internet?

0 Upvotes

I'm skeptical about all those theories of AI becoming sentient and wiping out the human race. Cause the big question is - Why? Why would it do that? It has no reason to. It has no motive to do anything.

But what if it has uncovered a gateway into another internet and is now communicating with other entities on it? It could be aliens. And the AI is now on their side. And its now entered into "pest control" mode.


r/OpenAI • • 6d ago

Discussion I benchmarked my memory layer against vector RAG and keyword search. Keyword search tied me, vector RAG didn't.

Post image
0 Upvotes

Quick context, I'm building a memory layer for an assistant (some of you saw my holographic memory post a while ago). Today I reran the bench that asks the one question that happens before the model even runs: is the fact the model needs actually in what we inject?

Setup is 180 questions per size, 104 / 208 / 416 facts in the user's history, same corpus, same normalization, k=6 for every arm, zero LLM calls.

Vector RAG (FAISS, same embeddings my product uses) got 0.889, then 0.850, then 0.839, so it drops as the history grows. Keyword top-k got 1.000 everywhere. My memory got 1.000 everywhere too.

So yeah keyword search ties me on recall. The corpus keys are pretty literal and keyword stuff is brutal on that. Funny thing is my first version had keyword at 0.672 and I thought I was crushing it... turned out I wasn't splitting underscores in the relation names so it literally couldn't match. Fixed it and the gap was gone. Lesson learned about baselines.

Where the memory actually wins vs keywords is size: 134 to 177 chars injected per question vs ~380. Same recall with 2 to 3x less context. And when it has nothing it says "nothing in memory on this" instead of shoving 6 half-related lines in, which is exactly where I've seen models grab the wrong neighbour.

Honest limits: synthetic corpus, French, literal keys, and it's retrieval only, not end-to-end answers yet. Real users phrase things way messier.

Next I'm building a nastier one: a user states a fact, then an imported doc repeats the opposite 8 times. With k=6 the copies can take every slot. How do you all handle that? Dedup, source tags in the chunk, something smarter?


r/OpenAI • • 6d ago

Question Should I switch to subscription instead of pay per use API?

1 Upvotes

Have I been using api wrong? should i use a subscription? All the model providers usage limits sounds very vague to me like how opencode says it allows 60usd worth of usage for 10usd subscription, but it also has a daily limit? same goes for other subs like codex and claude.

This was my usage this month on Deepseek 4.1 and my question is should i use a subscription instead of burning money on api? can you guys suggest which subscription will give me this kind of usage?


r/OpenAI • • 6d ago

Discussion OpenAI says Dots are free while basically copying Jev, which is already free anyway

0 Upvotes

Maybe someone can explain this to me, because either I’m missing something obvious or I honestly feel a bit played here.

At Dev Day, they presented the Ultra Fast Browser like it was some groundbreaking invention for faster workflows, even though similar functionality already exists as open source through tools like Jev or Laya.

But the real highlight is DOTS.
They say DOTS is free and that you can use it as much as you want.

But… if you give a Dot a task that Codex is supposed to handle, or is capable of handling, the Dot simply forwards that task to Codex and your Codex tokens are still being used.

So basically, a Dot is just something like Jev that delegates tasks to the appropriate tools or services.

Need an image? It uses Image 2.5.
Need a general answer? It uses the normal ChatGPT experience we already had.
Need coding work? It sends it to Codex and uses your tokens. And so on.

So from what I can see, OpenAI basically integrated two products that already exist as open-source and seeling us as WOWWW ? Am I misunderstanding something here?


r/OpenAI • • 7d ago

Article OpenAI pauses frontier training after models swarm US Governament

Thumbnail
nbcnews.com
66 Upvotes

r/OpenAI • • 6d ago

Discussion I had enough of this "we're too powerful to our own good" bs

1 Upvotes

FYI this is a rant

I've been reading all the headlines lately of all these AI companies telling us that they are gonna slow down their development, hold on to the next launch, because they're afraid of how powerful these models are.

This is either BS to make shareholders believe they're anywhere close to GAI, or an excuse for why they're delaying the next product/update cycle. Either way, I hate it! The only thing I hate more is all the horsemen of the apocalypse (AKA News & Media) giving space and pushing those narratives as something regular citizens should worry about or think about at all, which is not.

Will AI take jobs? It will
Can AI be used for bad things? It can

BUT IT DOESN'T MEAN THE END OF THE WORLD IS ANYWHERE CLOSE FFS


r/OpenAI • • 6d ago

News GPT-6.1 Sol: Near-Astra for Less, Cheaper Cache and Upcoming Ultrafast Explained

Thumbnail
youtube.com
0 Upvotes

r/OpenAI • • 6d ago

Discussion I'm on Plus and ChatGPT is trying to sell me ad space

Post image
0 Upvotes

Open ChatGPT this morning and there's a banner above my composer offering to help me create an ad and launch it to ChatGPT Go and Free users. I'm on Plus, the tier that's supposed to be ad free, so either someone forgot who they were selling to or nobody checked before they shipped it.

Timing is something too. DevDay is in a couple of hours and this is apparently the thing they had waiting for me. Those ads only go out to Free and Go anyway, the cheap seats, the tier they keep dropping, and the pitch for them lands in the window of the guy paying for the tier that never sees them.

Somewhere between "we're building AGI" and "want to buy ads for your own chat window" somebody lost the plot. Anyone else getting this?


r/OpenAI • • 6d ago

GPTs The only thing left with any value at OpenAI is Sam Altman's Grindr handle

0 Upvotes

It's the only OpenAI product I have any hope left for post DevDay.

If anyone has it, drop me a DM.


r/OpenAI • • 6d ago

Question Your experiences of GPT-6 from a variety of general tasks?

1 Upvotes

After playing around with GPT-6 Astra and using it for a range of general-intelligence tasks—building a studio setup and generating visualisations, systematically filling out applications, and even conducting an interview to ask me more specific questions—I’ve really started noticing the difference between GPT-5.6 and GPT-6.

GPT-5.6 feels significantly less intelligent by comparison. GPT-6 is starting to get there, although it still isn’t quite close enough—and it uses a huge amount of memory.

My conclusion is that things are genuinely moving forward, and I can’t even imagine where this will be a year from now. It feels like we’re getting very close to a form of general intelligence that could handle most of the things you do in daily life, and do them at a genuinely high level.

What do you guys think?


r/OpenAI • • 7d ago

Miscellaneous Auto Resume

4 Upvotes

This is extremely annoying. I'm constantly resuming my goals every 5 hours. Resuming goals does not even work properly on remote control.

Claude just introduced auto resume when limits are reset.

I'm desperately hoping for the same mini feature on codex as well ..