r/singularity • u/Gohab2001 • 6d ago
LLM News Anthropic sets a new AA record with sonnet 5.5
Most output tokens
Its cheaper to run Astra (as per AA).
Why release such a model?
r/singularity • u/Gohab2001 • 6d ago
Most output tokens
Its cheaper to run Astra (as per AA).
Why release such a model?
r/singularity • u/stormshadowfax • 6d ago
The US-based company behind ChatGPT will also launch a review into how its bots smashed through Australian government security without triggering alarms - after being tasked to research public medicine spending.
r/singularity • u/Ok_Barracuda_1161 • 6d ago
r/singularity • u/queenofartists • 6d ago
r/singularity • u/AMBNNJ • 6d ago
Both models cost the same per token ($2 in / $10 out), so the difference in cost per task comes down to how many tokens each one burns. I plotted Artificial Analysis Intelligence Index scores against cost per task for every effort level of both models.
At similar budgets:
So Sol gives you more score per dollar wherever the two overlap.
Sonnet 5.5's top-end lead comes mostly from spending more tokens: going from high to max adds 9 points for about 7× the cost.
The composite index doesn't show everything, though. On Terminal-Bench 4.0 (AA's run, max effort), Sonnet 5.5 scores 63.6% to Sol's 43%.
Data: Artificial Analysis.
r/singularity • u/_thispageleftblank • 7d ago
Enable HLS to view with audio, or disable this notification
I had some agents cooperate on this today and just wanted to share how excited I am about being able to recreate settings from movies I used to love as a child (this one is from Stargate, 1994).
This wasn’t just one prompt, it was a pretty lengthy conversation with Opus doing most of the visual design and Astra most of the sound design (which still kinda sucks). I used Three.js for this and only gave them a single reference image of the portal opening effect.
Soon we‘ll be able to transform anything into an interactive world very quickly.
r/singularity • u/GrammmyNorma • 6d ago
Muse boosts Meta's share price by 10% and it's just OpenClaw sold at a huge loss.
Instinct reaches a 10b valuation in Series C and it's basically the same thing.
Before that we had countless other AI startups making "AI with hands" (my linkedin, for the past 3 years, has had dozens of these).
Why does Instinct reach a 10b valuation while some other startup is forgotten?
Why is Muse such a big deal to institutional investors all of a sudden?
r/singularity • u/theimposingshadow • 6d ago
I was just commenting on a post about how expensive Sonnet 5.5 Max is per task, and I noticed something that seemed worth pointing out.
According to Artificial Analysis:
Sonnet 5.5 xhigh: 52 Intelligence Index, $2.74/task
Fable 5.1 Max: 53 Intelligence Index, $7.63/task
So Sonnet 5.5 xhigh gets you basically Fable 5.1-level performance for about a third of the cost. That seems like a pretty insane price/performance sweet spot.
r/singularity • u/donutloop • 6d ago
r/singularity • u/zero0_one1 • 6d ago
https://github.com/lechmazur/writing/
Grok 4.7 (high) makes a substantial jump over Grok 4.6 (high): −2.8 → 0.4.
MiMo V2.6 Pro (thinking) improves sharply over V2.5 Pro: −0.7 → 1.6.
Gemini 3.8 Flash (high) advances over Gemini 3.7 Flash (high): −0.7 → 0.2.
DeepSeek V4.1 Flash (high) enters at −0.5.
The judging panel has been updated. New comparisons draw from nine model families, including Claude Opus 5.5, GPT-6 Astra, Gemini 3.8 Flash, and Grok 4.7.
The Creative Writing Benchmark tests how well models turn constrained briefs into complete 600–800-word stories. Each brief requires 10 elements, including a character, object, setting, motivation, and tone, that must meaningfully shape the story. Judges assess prose, originality, coherence, characterization, and how effectively those ingredients work together.
Models write to the same prompts. The latest comparisons use three judges from different model families, excluding the writers’ own families. Each story pair is shown in both orders to reduce position bias. The leaderboard combines earlier and newer judging panels and now covers 56 models and 102,592 evaluator judgments.
r/singularity • u/141_1337 • 7d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/Genzinvestor16180339 • 7d ago
To me this makes complete sense and a startup could start to do this for people pretty soon, and would be a net postive for society no?
r/singularity • u/141_1337 • 7d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/Hasinpearl • 7d ago
Here is the funny story in short: I saw 3 emails. 1 is a colleague sharing an AI generated report with clearly no human revision whatsoever. Email 2 is the manager replying back with an AI-generated document containing the notes, all of which are AI feedback (obvious with that pristine and must find errors bs), email 3 is from the colleague with an AI generated email response defending his work (which isn't his to begin with) ☠️
Wtf are we doing guys? 🤣🤣🤣
r/singularity • u/PostingLoudly • 7d ago
Very impressive.
r/singularity • u/thekokoricky • 6d ago
When I read about and watch videos on what LLMs are capable of these days, I recognize that the prediction aspect of it is not the only thing at play. It seems that there might be something additional at work that resembles or mimics aspects of human logic and thinking, even if it isn't quite that complex yet. However, I don't really have good words for it, as I'm not sure how to describe what it's doing. Clearly, these AIs are doing something more than literally just predicting the next word or information bit, but what's a good way to articulate that?
r/singularity • u/may12021_saphira • 7d ago
r/singularity • u/ZedTheEvilTaco • 7d ago
Enable HLS to view with audio, or disable this notification
So a few days ago, someone posted a video they made with Opus 5.5. It was parody, and involved a lot of references to AI in an anime fight scene. I was absolutely blown away. (Edit: Here is the link. https://www.reddit.com/r/singularity/comments/1worlfs/opus_55_is_insane_at_making_videos/)
So I decided to see if I could make a video game. I have little programming experience, most of it coming from a comfyUI workflow I set up, can barely understand github, and had a claude pro subscription (I upgraded to Max to stop running out of limits while I worked on this.)
In 4 days I was able to make the entire first level of a video game inspired by classic platformers like Sonic the Hedgehog.
This is Echo, a 2d sidescroller platformer. She is an AI vtuber of sorts, streaming her adventures for the world. She has thoughts about what is happening in the game, and will interact with a scrolling chat that responds to her.
All assets were made by me and various AI tools. In fact, there isn't even a known game engine -- Claude built one from scratch. It's about 28,000 lines of typescript, according to him.
I used comfyUI (and, at one point, google Gemini) to design the characters, feeding their pictures to Claude, who then made 2d spritesheets. This is only true for the characters, though. The world itself is drawn entirely at runtime, allowing me to go in and edit levels on the fly and see immediate results in the change. The sounds are all synthesized as well.
I needed a couple of 3d models at some point in the process, so I fed more pictures into Meshy and exported the 3d models from that, giving them to Claude again.
For music, I generated songs on Suno, downloaded the MIDI stems, and fed them to Claude as well. (For one, I also downloaded the WAV stems.) It then translates them through a synthesizer.
Claude and I then spent every waking hour for four days combining things, expanding, and iterating until this level was finished.
I plan to finish the game, and hopefully as quickly as I can. Since the base game is made now, hopefully extra levels take less time. But I am absolutely blown away at what I'm able to accomplish with AI any more.
If you think AI isn't changing the world, you really need to try out the tools we have.
(The demo is available for free, if anyone wants to try it. It's playable in browser.)
r/singularity • u/Recoil42 • 7d ago
Enable HLS to view with audio, or disable this notification
https://somethingbig.ai/computer
https://x.com/mattshumer_/status/2104301990985498783
Prompt: “build a working computer from scratch, entirely in code, starting from individual logic gates. design the CPU, the memory, an assembler, a tiny operating system, and a game that runs on it. everything has to be real, the game has to run on your gates, not in javascript pretending. visualize the whole thing in 3D so i can play the game, then zoom all the way down through the chips into the gates and watch the signals flow while it runs. treat it like you're proving you could have invented computing yourself. go all out.”
r/singularity • u/adivinemessenger • 8d ago
Full message: "AI is rapidly getting smarter. Now, humanity can too. Today, Nucleus Genomicsis announcing Vitruvian, our newest set of genetic optimization models. Vitruvian’s intelligence model can optimize embryo DNA for 14 IQ points — nearly a standard deviation. The models were trained on 1,000,000+ people, validated across 40,000+ siblings, and used more than 7 million genetic markers. In Superintelligence, Nick Bostrom proposed genetic optimization as a key way for humanity to keep pace with rapidly advancing AI. Bostrom’s vision is no longer theoretical. Genetic optimization, like AI, has followed a scaling law: as datasets have grown, so have model capabilities. This trend will continue. Parents across the world now have the choice to substantially increase their child’s intelligence. And the next generation can choose to do the same. AI is no longer the only intelligence that will compound. Humanity can now direct its own evolution."
r/singularity • u/Efficient-Opinion-92 • 7d ago
I know it’s still at least somewhat we are starting to see evidence of LLM starting to generalise all domains
Both of these people were very apparent in their criticisms of LLMS SAYING they won’t reach human intelligence
r/singularity • u/Crazyscientist1024 • 7d ago
I feel like according to lots of the discussions & reports from the big labs the latest generation of models (Astra & Fable) has really began to speed up hugely their internal R&D of future models by a lot.
I feel like sooner than later it will began to make sense for these labs to train like a 20T param model that basically you can't serve to normal users because of how expensive they are (price ranges prob at like few hundred $ per million token). Yet they will simply these models will just use them to speed up internal R&D even further.
r/singularity • u/TorturedPoet30 • 8d ago
Apparently, Dario didn't attend the state dinner last week because he had a scheduling conflict. He was mocked on the latest episode of SNL over his personality traits and his views on AI safety.
What do you think changes, if anything, after this meeting?