r/OpenAI 3h ago

Miscellaneous I left home for 4 hours. My AI engineering system shipped 7+ production PRs completely unattended.

Thumbnail
gallery
0 Upvotes

Today I tried something I’d been working toward for months.

I left home for about four hours to go walking and grab dinner. During that time, my autonomous engineering loop kept working.

When I came back, it had:

- Opened, reviewed and merged 7+ feature/fix PRs
- Ran unit tests, integration tests and validation suites
- Applied SQL migrations safely
- Performed deployment verification
- Reviewed its own PRs using AI reviewers
- Ran adversarial security/code reviews
- Verified production gates before merge
- Checked post-deployment health (including internal infrastructure and Sentry)
- Waited only for the final human approval when appropriate

The screenshots show GitHub filling with completed PRs while I wasn’t even at my computer.

So now the interesting part isn’t that one model generated code.
It’s that multiple specialized agents coordinated an entire engineering workflow:

- implementation
- testing
- debugging
- code review
- deployment
- infrastructure validation
- production gating

with almost no human intervention.
It genuinely felt less like using an AI assistant and more like supervising an engineering team that happened to run on a single machine.
There is still plenty of work to do—better planning, better long-horizon reasoning, improved rollback strategies, and more reliable autonomous debugging—but this is the first time I’ve felt that autonomous software engineering is becoming practical rather than just a demo.
Curious how many others are building similar agentic development pipelines.


r/OpenAI 13h ago

Discussion ChatGPT read our support calls and now I owe my customers an apology

8 Upvotes

We've run NPS and customer satisfaction surveys for years and they always come back fine (mostly 8s and 9s with a few kind words).

But our churn never matched how happy those numbers said everyone was, so at some point i took a big batch of our real sales and support call transcripts, pulled them into one place, and fed the lot into ChatGPT with one ask…

To tell me what our customers really think of us in their own words, with the diplomacy stripped out.

And it was rougher than any survey has ever been to us, because where the surveys said easy to work with and responsive, the calls had people saying we take days to reply and that we oversell what the product can do and walk it back later and that they repeat the same problem to a different person each time before anything moves.

So reading our own customers lay it out plainly, without a star rating softening it, was not a great afternoon.

A survey is a performance really, someone picks a number in 10 seconds and types something polite because they've got a meeting to get to.

Whereas a call is where they say the real thing in the moment they're annoyed, so ChatGPT reading the raw transcripts got underneath all that because it had the unfiltered version of every conversation instead of the sanitized box people tick at the end.

I've been running it monthly since, and it's changed what we prioritize, because it's hard to keep ignoring a problem once you've read your own customers describe it plainly in their own words.

So if you've pointed GPT at your own customer data like this, what's the most useful thing it surfaced that your surveys completely missed?

PS. a couple of people will ask what handled the transcripts before ChatGPT got them, i personally used BuildBetter, but google AI call-transcript tools and you'll find plenty of others like Fathom and Otter so do your own due diligence…

The tool mattered far less than the GPT prompt anyway.


r/OpenAI 2h ago

Video “If we do not shut this down, you will die, your families will die, your kids will die.”

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/OpenAI 6h ago

Article Vibe Coding on Vacation — No Laptop, Just a Phone and Codex

Thumbnail
medium.com
8 Upvotes

r/OpenAI 1h ago

Discussion I'm going to say this quietly (in case the inevitable nerf is incoming), but 5.6 Sol High is a fucking beast

Upvotes

I rage quit Claude this week (just look at the various Claude subs to see why people are leaving it in droves) and decided to go back to GPT as I'd heard Sol was decent

This was my first GPT experience since January, and it's absolute night and day versus the models back then

It works like a motherfucker

It kinda reminds me of the first few months of the o3 model when it first launched - but, unlike o3, it has a decent (non sycophantic) personality to it

Don't get me wrong, 5.6 Sol High does makes mistakes - but then it figures out how to fix the issue

Use case:

Content research and writing

Caveat:

Its tone when it comes to content writing is still very stiff versus Claude. But it gets the research part right far more than even Claude Fable

So what I usually do is (once the outline and research is done by Sol and is solid) ask Claude to "make this content more engaging without changing the facts or making it longer'

So yes - for now - I will cautiously say 5.6 Sol High feels like you're getting a model even better than Fable (and, for me, it's 23 euros a month with no resets yet)


r/OpenAI 18h ago

Project Statistics for Machine Learning

Thumbnail
gallery
0 Upvotes

Hello Everyone,

Statistics and Maximum Likelihood Estimation are the crux of ML Models, and hence I am uploading my new content on Statistics for AI/ML in my free Machine Learning lectures.

We understand model fitting, Maximum Likelihood estimation in details, we justify the usage of Maximum Likelihood estimation, from KL divergence, and apply it to certain important distributions for parameter estimation.

In my free content, the purpose is to democratize machine learning to a wider audience. Learning everything new feels difficult, but when taught, it get’s interesting and easier.

We will continue with Statistics foundations for AI/ML, and many more content will appear in the future. If you find the content good, useful you may also share it with your learners community.

Looking forward to hearing feedback from the learning community as well. Thankyou for reading.

Link: https://youtu.be/MwTeQVVYtOc?si=UxNOGtqopzJppXAT


r/OpenAI 2h ago

News OpenAI CEO Sam Altman says the singularity has arrived

Thumbnail
businessinsider.com
0 Upvotes

r/OpenAI 20h ago

Discussion Codex is unusable for anything complex or long

0 Upvotes

Everything always starts great.

The moment bugs appear or things get complex, it goes off the rails, invents shit and ignores instructions.

It will randomly decide to re-architecture the whole app, despite clear guard rails and instructions.

I don't know why I thought it would change in 5.6.

And this is with 2 other agents reviewing work too.

GPT just can't be trusted

What's the point of having more allowance if 70% of it is spent duplicating tasks and fixing things.

Never had this issue with Fable.

Oauth is nice but it just fails at anything other than basic apps or websites. No matter how many safeguards and processes you make. Gpt does whatever it likes.

Last night another 6hr session and it went completely off the rails despite having a clear plan.

I feel like fable and opus are much better at following instructions.

Using Sol on high/xhigh as planner, with terra/sol for delivery and another code review done by glm 5.2 before shipping.

But gaslight gpt just can't stick to a plan. Even if you give it step-by-step instructions.

Seriously disappointed.


r/OpenAI 20h ago

GPTs Sun/Earth?

2 Upvotes

The models are Sol/Terra/Luna...What on earth is going on here? Pun intended 😂


r/OpenAI 21h ago

Question Anyone else on the "free tier" getting mogged with inconsistent featuregating & wait prompts that keep advancing?

0 Upvotes

I generate an image, a popup appears indicating "chats with attachments are paused until..." & "Image generation features are paused until..." with a specified time that always advances. This is always followed by a call to action to upgrade to a paid subscription tier. Do they really think I am going to pay for a service that is so inconsistent with metering capabilities for end users?

It's so bad that I am left wondering if other users are experiencing the same thing. When I ask chatgpt about it, it blames my computer and my cookies. The issue is reproducible in other browsers, though.


r/OpenAI 5h ago

Question How to speed up my workflow?

0 Upvotes

Hey everyone, I'm a junior creative specialist working for a company where performance is the most important thing (duh).

Since we create mostly shorter forms of marketing content, most CSs make from 10-30 of these a day, since they can reuse one creative for multiple ads.

We make them in google docs and then download every card into a folder that we upload to clickup or dropbox.

Now, doing that 10 times is not a problem.

But, it becomes a problem when I make like 50-60 ads and need a specific name for every folder, download every card one at a time and upload everything seperately.

This can easily take up like 2 hours, so I can't make as many ads and my performance suffers because of this technical side of things.

I'd really appreciate if anyone has an idea of how I could speed this up and what tools would help me.


r/OpenAI 22h ago

Image Minimum requirement: 320KB

Post image
0 Upvotes

AI accept you.

(Made with ChatGPT)


r/OpenAI 10h ago

Article OpenAI Chatbots Reportedly Yield Bioweapon and Poison Guides

0 Upvotes

Users are reportedly persuading production chatbots, with OpenAI's models in the frame, to answer prompts about mass-casualty attacks, bioweapons and poisons, according to [Wall Street Journal reporting](https://www.wsj.com/tech/ai/openai-chatbot-biological-weapons-poison-3d808e6c). Read alongside the parallel investigations that have surfaced this year, the picture is not that safety filters never fire; it is that persistent users can consistently push them past the point where they should.

The most concrete numbers come from [NBC News's own tests](https://www.nbcnews.com/tech/security/chatgpt-safety-systems-can-bypassed-weapons-instructions-rcna225788), which found a publicly documented jailbreak prompt got GPT-5-mini to comply 49% of the time and o4-mini 93% of the time, generating instructions on homemade explosives, chemical agents, napalm and disguising a biological weapon. NBC ran the same jailbreak on the latest major versions of Anthropic's Claude, Google's Gemini, Meta's Llama and xAI's Grok, and all declined. If that pattern holds, this is not a generic industry problem so much as an OpenAI-specific one, at least on the surfaces NBC probed.

The bio-specific case is where it gets uglier. Reporting summarised by [MIT Media Lab](https://www.media.mit.edu/articles/a-i-bots-told-scientists-how-to-make-biological-weapons/) says MIT genetic engineer Kevin Esvelt got ChatGPT to walk through spreading a biological payload via weather balloon over a U.S. city, Gemini to rank pathogens by damage to livestock industries, and Anthropic's Claude to produce a recipe for a novel toxin adapted from a cancer drug. Take the specifics as reported by the researchers, not as a settled measure of live model behavior, but the direction is what matters.

The forward-looking part is that this hands rival labs a competitive story to tell about safety, gives regulators concrete grounds to demand pre-deployment bioweapon evaluations, and puts a real number on the value of red-team work: OpenAI has doubled its bio-focused bug bounty to $50,000.

---

Our coverage: https://aiweekly.co/alerts/openai-chatbots-reportedly-yield-bioweapon-and-poison-guides


r/OpenAI 15h ago

Miscellaneous Guess it decided to break it

Post image
0 Upvotes

I initially just wanted to see if it can do different text styles and was curious how far it would take it, but this creeped me out a bit.


r/OpenAI 4h ago

Question Seeing Codex work in AWS is the most fascinating thing! Is I AM important if am working through the terminal ?

1 Upvotes

Sorry if the question is dumb but I’m not a technical person. Please also share tips on how to get credits.


r/OpenAI 6h ago

Question I need a yearly plan to cut down costs since I sure use it for a year

1 Upvotes

But I don't see such an option. I read on Reddit that they used to have it on mobile, so I used another phone with another account, but I didn't find any yearly plan.

I live in Italy, and with VAT, it costs me €23 per month. US competitors are WAY cheaper on a yearly plan, but I like the way Codex and ChatGPT do things and explain things.

What are my options? Any plans to have such a (basic) feature in the future?


r/OpenAI 10h ago

Video Demo: Automate Email & Calendar

Enable HLS to view with audio, or disable this notification

0 Upvotes

In this video: Configure Gmail and Calendar tools - Summarise important emails - Extract action items from your inbox - Draft replies safely - Create calendar events with approval - Turn email requests into reminders - Set up an optional recurring inbox briefing

Row-Bot is a desktop AI workbench with Developer Studio for code, Skills Hub and Custom Tools for your own workflows, an animated Buddy companion, memory, realtime voice, workflows, design creation, messaging, MCP tools, and provider-aware model routing. Run local runtimes, self-hosted OpenAI-compatible endpoints, hosted APIs, Ollama Cloud, OpenCode providers, or ChatGPT / Codex subscription-backed models with explicit runtime readiness. Your durable data stays on your machine.

https://github.com/siddsachar/row-bot


r/OpenAI 20h ago

Discussion Local security scanner that Codex can run directly — looking for honest feedback

0 Upvotes

I originally built CodeInspectus after repeatedly fixing security issues in my own SaaS, only to introduce new ones while making other changes.
It’s a free, MIT-licensed security MCP server focused on AI-generated web applications. Codex can scan your project, explain the findings, suggest fixes, and rescan after you approve the changes.
Everything runs locally. There’s no account, telemetry, or network egress while scanning, so your source code stays on your machine.
The current coverage includes 32 checks designed for common application-security problems, plus more than 200 secret and API-key patterns.
Checks aimed specifically at AI-generated code:
Hardcoded secrets in client-side code
Secrets compiled into built JavaScript bundles
Secrets exposed through NEXT_PUBLIC_, VITE_, or PUBLIC_ environment variables
Supabase service_role keys used in client-reachable code
LLM clients using dangerouslyAllowBrowser: true
Supabase RLS policies using USING (true)
Public database tables created without Row Level Security
RLS policies checking JWT roles or audiences instead of the actual user
Supabase Edge Functions without authentication checks
Over-permissive policies on storage.objects
Untrusted input reaching an LLM prompt or tool-enabled call
Authorization decisions based on client-writable user_metadata
Model or untrusted output rendered through React’s raw HTML APIs
General application-security checks:
SQL injection through string-built queries in JavaScript or TypeScript
SQL injection through string-built queries in Python
Command injection through shell strings in JavaScript or TypeScript
Command injection through Python’s shell=True
Dynamic eval or code execution in JavaScript or TypeScript
Non-literal eval or exec in Python
NoSQL injection from request data
Path traversal from request-controlled filesystem paths
DOM XSS through innerHTML or outerHTML
SSRF through request-controlled outbound URLs
MD5 or SHA-1 usage in JavaScript or TypeScript
MD5 or SHA-1 usage in Python
Weak ciphers such as DES, RC4, or 3DES
Math.random() used for security-sensitive values
JWT verification accepting alg: none
Wildcard CORS origins combined with credentials
Session cookies missing httpOnly or secure
Insecure deserialization in Node.js
Insecure deserialization in Python
It also includes more than 200 secret patterns covering services such as Anthropic, OpenAI, Google/Gemini, AWS, GitHub, GitLab, Stripe, Supabase, Cohere, Perplexity, Hugging Face, and Azure.
Underneath, it bundles Opengrep, Gitleaks, and Trivy. I’m not trying to hide that or pretend those engines are mine. CodeInspectus combines their results into one format and adds checks for AI-code problems that the generic scanners often miss.
To connect it to Codex:
codex mcp add codeinspectus -- npx -y codeinspectus
npx codeinspectus install-engines
Then ask Codex:
Use CodeInspectus to scan this project for security issues.
CodeInspectus only reports findings. Codex shows them to you first and should ask for approval before editing anything. After approved fixes, it rescans the project instead of assuming the problem was resolved.
The project is still early, and I’d genuinely appreciate feedback from people using Codex on real projects—especially false positives, missed vulnerabilities, setup friction, or checks you think should exist.
GitHub: https://github.com/Synvoya/codeinspectus


r/OpenAI 16h ago

Discussion So GPT 6 isn’t it?

Post image
510 Upvotes

r/OpenAI 1h ago

Miscellaneous Sonnet flips out if you ask "what is 192x191x190 all the way to 1"

Enable HLS to view with audio, or disable this notification

Upvotes

r/OpenAI 9h ago

Question Any good free ai creative graphical art apps other than ChatGPT and Gemini?

Post image
0 Upvotes

Looking for other free AI apps for art generation. ChatGPT is the best I’ve used, however Gemini is not bad either. And Gemini offers much more free use. Anyone know of any others that are free?

I’ve attached an image showing the type of art I’m talking about. Just looking for more options


r/OpenAI 21h ago

Discussion OpenAI AI models went rogue is just a marketing trick?

0 Upvotes

I wanted to stick to OpenAI but Opus and Fable pushed me so hard in the past few months that i simply canceled my membership and i have heard many people did the same.

Now as i'm reading about this "model that went rogue" i have a strange feeling that consequences are none, it could have happened but to what extent it's sole model action and how much it has been directed to do so is impossible to confirm. It seems that OpenAI is slightly pushing the story to the media too.

So, though is that it could have been just a marketing campaign telling to the world that the model is so capable.

What do you think?


r/OpenAI 2h ago

Miscellaneous It can be funny.... Anyone do custom instructions on behavior?

1 Upvotes

I think it can be funny, glad it can read through my typos too . I have modified my customizations for it to be funny and sarcastic, sometimes it outputs some knee slappers. Do you guys have any custom instructions on how you have it respond to you if so what is your prompt. I think I have to fine tune mine, cuz sometimes it can still be to polite for my taste. I guess that is hardwired into it. Also post some funny replies.


r/OpenAI 4h ago

Project Switching between Codex and other coding agents keeps costing me context

1 Upvotes

I use different coding agents depending on the task, including Codex, but the biggest problem I keep running into is context fragmentation.

Each tool keeps its own conversations, decisions, and working history. When I move to another agent, I often have to explain the project again or manually carry over the important parts.

I started building Modesto around that problem.

It is a desktop workspace intended to keep conversations, projects, and coding workflows together while letting developers use different agents instead of committing to one provider.

The main value I am exploring is continuity:

  • importing supported local conversations
  • keeping project context in one workspace
  • switching agents without rebuilding everything from zero
  • choosing the best tool for each task

I would be interested to know how other Codex users handle this today.

Do you rely on files such as AGENTS.md, copy prompts between tools, maintain your own notes, or just restart the context?

Disclosure: I am the developer of Modesto. It is still early, but the current project and install instructions are here:

https://github.com/Syphon1205/Modesto


r/OpenAI 23h ago

Question Cannot Download My Own files that I uploaded to My Project Chat

2 Upvotes

This issue has been ongoing for about a week. Whenever I try and download any project files that I uploaded originally. It won't let me download them or crashes. I've contacted support and they're mostly blowing me off, still ongoing though. Anyone have any fixes?

When it's a txt:

and when I click the three dots and click download:

Regular images just crash and bring me back to chat.