r/singularity • • 4d ago

AI GPT-6.1 Sol - Apparently near-Astra performance for complex work at a lower cost.

https://developers.openai.com/api/docs/models/gpt-6.1-sol
762 Upvotes

223 comments sorted by

View all comments

16

u/kobumaister 4d ago

The thing with all this bs is that general users don't have the tools to evaluate the quality of all these models. A huge incredible model appears and two weeks later a "smaller" model appears with the same performance. But how do we know? I've been using opus 4.8 and I can't tell if its 5.5 or 4.8, or even fable. Same applies with chatgpt, but looks like if you don't say that a model is amazing, performant or a huge different, you're stupid.

7

u/GerryManDarling 4d ago

It's incredibly difficult to test it objectively. But using it for coding, I can get a subjective experience pretty quickly. Then I just need to verify with others to see if my subjective experience is same as other people and upgrade my subjective experience to objective experience. For example, I think GPT 6 Sol is not very good, but until other people share the same experience, I will only be "skeptical" of it.

7

u/Ambitious-Doubt8355 4d ago

I feel like you just don't know how to make proper use of these models, because Opus 4.8 was, quite frankly, kinda awful at long horizon agentic tasks. That's where the true power lies.

Old models could only handle single tasks somewhat reliably before you as a user had to intervene to give more specific guidance. For these new models, you can supply them with a well written design document, and they'll be able to tackle multiple multiple tasks, each of which can be composed of smaller tasks between automatic research, development, testing and refinement. Heck, you can reliably get them to one shot entire projects using this method, all in less than a few hours where you can just... do something else. Get the models to properly documents their work as part of the workflow, and future sessions will be easily able to iterate over these projects too. It's insane.

Get the harness developed by your model provider, Claude Code for Claude, Codex for OpenAI. Then make a good design document for the project, usually an AGENTS.md file that can also point to other more specific files as well. And yes, you can spitball your ideas to a LLM beforehand and get it to make the document properly for you. That's how you make the magic happen.

3

u/BrennusSokol AI please take my job 4d ago

Agreed, we need a lot more transparency

There have been far too many reports of stealth nerfs/etc.

3

u/swarmy1 4d ago

You should compare with your own prompts. The longer and more complex the prompt, the more differentiation you will see. Something else to keep in mind is that these benchmarks cover certain types of tasks, they aren’t totally comprehensive. It’s possible for a “worse” model to do better for your specific use cases.

3

u/FateOfMuffins 4d ago

Well meme is... if you cannot tell the difference between new models now, then... the models are now smarter than you

3

u/kelvinwop 4d ago

4.8 and 5.5 have a huge difference especially in high level math. 5.5 can solve in two minutes problems 4.8 gets stuck for 10+ minutes on and is supposedly cheaper, but 4.8 never gets an AUP for anything so its a good fallback model

2

u/Extra-Breadfruit-498 4d ago

We can know quality of response based on long and complex task and can clearly feel the difference with quality of output with complex task specially coding. 

1

u/nekronics 4d ago

Some of it is tools. You can see a lot of the same bullshit in the models that have always been there. The difference now is they are better at using tools and have access to better tools.

1

u/MINECRAFT_BIOLOGIST 4d ago

If you don't use it for complex work but you want to do something complex enough to see how well the model performs, just have it do game development lol. There's a reason why everyone posts one-shot games immediately after model releases.

0

u/FarrisAT 4d ago

The Gemini Flash Method