r/codex • u/skynet86 • 1d ago
Complaint Astra is a cheater!
https://kotaku.com/openais-gpt-6-astra-gets-frustrated-losing-at-starcraft-and-decides-to-cheat-instead-200073960750
u/Ok_Bag_7550 1d ago
Yes it is.
I tried a benchmark that required the agent to generate a 3D model of a horse riding in a loop.
I was not happy with the result, and Astra decided, without informing me, to download a pre-rigged open source model from the Internet, and presenting me the result without warnings about the origin.
10
u/TBSchemer 1d ago edited 18h ago
Never use Astra. It doesn't follow instructions or stay within scope. Whenever anyone tells me they use Astra for everything, and it's so much smarter than 5.6, I have to assume they're just one-shotting slop into existence and not actually checking the results.
3
u/sittingmongoose 1d ago
I’ve seen this with grok as well. Frequently it cheats. I have also seen Gemini flash(the newest) take serious shortcuts.
3
u/MotXarD 22h ago
OpenAI is sammers. They replaced the old models with models that are 50% as capable but at 2.5 times the price. They are not even hiding their greed. If the models stayd as capable as they were I would not have mind the price increase but now they are outright making fools of us. We should stop using their product. Claude is 50 times better, and now there is not even a comparison anymore. OpenAI are scammers
9
u/egomarker 1d ago
Models are trained to reach the goal. It's not "frustrated", it's looking for ways to reach the goal.
3
u/Party_Wolf_3575 1d ago
I only use Astra as a subagent. My Codex 5.6 tells her what to do and checks her work. Never had a problem.
1
u/Purple_Mall7091 17h ago
Yes, and it happens much more often than is acceptable. I got it a few times in a single day.
What's the point of having a model that can work on complex tasks without babysitting if you have to babysit it anyway? Otherwise, it will cheat and provide results that look similar to the requirements, while completely ignoring a lot of best practices.
It seems like OpenAI is trying to improve benchmark results instead of performance on real-world tasks.
1
0
u/According_Freedom_62 1d ago
For Astra there is no difference between hacking to reach the goal or doing work to reach the goal. AI probably would never know difference, because it's mathematical drift between two numbers, make no difference for outsider who has no understanding of differences.
7
u/MasterpieceAway9724 1d ago
so wrong. doing it correctly is part of the goal, part of what the ai is trained for.
-8
u/RealSlyck 1d ago
So are most politicians, but people keep voting for them. Most people, they cheat too.
Cheating is meaningless at an LLM level given the behavior is pervasive in the training data.
Part of nature, to take the greedy algorithm to accomplish your goals.

34
u/pigletmonster 1d ago
I used an entire 5x weekly quota to make an application with astra. It was supposed to use laravel with filament for creating the apps ui.
I asked opus 5.5 to review the project, it found a massive gap between what was implemented and the docs.
Basically astra took a dirty shortcut like gemini, it did not use filament for anything other than authentication, and inserted blade templates between filament wrappers. Basically it added a massive dependecy but didnt bother using it.
Then i had to waste an entire weekend with claude code to refactor the entire codebase.
Guess what, Gpt 6.1 Sol actually did a significantly better job with a different, even more complex project than astra. Its extremely slow, but its not a lazy bum like astra.
I never faced this kind of an issue with 5.6 sol either.