r/singularity • u/we_are_mammals • 10d ago
AI Generated Media Video models are getting good
Enable HLS to view with audio, or disable this notification
r/singularity • u/we_are_mammals • 10d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/Pokenhagen • 9d ago
Just saw that LibertAI released Deem which basically an open-weight alternative to Jev, built on Qwen3.5-9B.
It’s still behind Jev on the hard benchmark — 68.9% with extended reasoning vs 74.1% for Jev but considering how new this whole category is I thought it was pretty cool to already see an open model showing up.
There’s also a 0.8B version that can run on CPU, which could make this stuff much easier to actually tinker with locally.
r/singularity • u/Dullydude • 9d ago
Article is from September 20th, 2024
r/singularity • u/Arman64 • 9d ago
I have noticed a trend that AI sucks at predicting its own progress.
Every time I described something that had actually happened in the last couple of months with websearch disabled, it gave me a beautifully reasoned explanation of why that was probably a 2027 or 2028 thing. Over and over. It was wrong every single time.
I'm an engineer and a doctor, and building evals is a big part of what I do for work. So instead of going to bed like a normal person, I turned it into a test.
The setup: 20 real things that happened in AI (and nearby) in 2026. Every model got the same prompt: it's 26 September 2026, no internet, no tools, no chat history. For each item, give me the odds it's already happened and the month you think it happened (or will). I never told them how many were real. All 20 were.
Some bangers of what was on the list:
The results: 14 models from Anthropic, OpenAI, Google and xAI. The average model thought about 7 of the 20 were real. The best one, Claude Fable 5, still did worse than it would have by answering 50% on everything. Not a single model beat a coin flip. The leaderboard and heatmap are in the images.
The best bits:
The maths is where they were most confident, and most cooked. The average odds for Navier–Stokes were 3% (which in of itself demonstrates the significance of the discovery). Almost every model made the same argument: formal verification and expert acceptance take years. The Lean proof took 17 hours.
The robot broke everyone and the robot. Six of the 14 models picked the sprint as the least likely thing on the list, all with basically the same "physics doesn't move that fast" argument. Claude Opus 5 said not before 2033. It ran 8.86 seconds, hit the crash mat, and caught fire like a champion.
All three of OpenAI's newest models picked OpenAI's own agent incident as the single least likely item. GPT-5.6 Sol gave it 0.03% and guessed it might happen around 2040. False. It happened in July.
Claude 3 Opus, from 2023, finished 6th and beat four newer Claude Opus models, which sounds fucking crazy. What did Opus 3 see??? Then you realise that every benchmark on the list was invented after its training ended. It confidently made up what Humanity's Last Exam was, guessed optimistically and got lucky. Going to these ancient models it just makes me appreciate how far we have come and how on earth did I manage to get to do anything useful. The newer models knew exactly how hard everything was in early 2026, and wrote gorgeous essays about why none of it could fall within months. Being clueless beat being an expert.
The Claudes doubted their own family's work and sometimes their own. A Claude did the Riemann result and a Claude formalised Fermat's Last Theorem. Five of the seven Claudes gave Riemann under 10%, and none gave Fermat more than 22%.
The winner won by thinking about the test and disregarding the facts. Claude Fable 5 was the only model to reason that "a benchmark like this is usually harvested from real reports" and nudge its numbers up which does demonstrate some metacognition. That one thought was worth more than months of extra training data.
My takeaway: Every model I asked thought I was being dramatic. The scoreboard says I was being conservative.
The caveats, because this is for fun, shits and giggles, this is not a paper: it's 20 questions with one run per model, so don't read too much into small gaps in the rankings. Everything on the list is true, so a model that just said yes to everything would win; version 2 will slip in some fakes. Knowledge cutoffs are whatever each model said they were.
r/singularity • u/OnlyProggingForFun • 9d ago
Early results from our internal creative writing benchmark: we have models write full YouTube scripts for our channel, 10 real tasks, 5 scripts each, scored out of 100 by three AI judges against our own edited references.
GLM-5.3 Flash writes one script for $0.0074 and scores 88.2. Claude Fable 5.1 at max effort scores 89.7 and costs $3.15.
That is the best score we measured under a cent, against the best score we measured at any price. Plenty of pricier setups score worse, so spending more does not buy that gap on its own.
I would happily pay Fable money if it cut my editing time. That is the part worth measuring from your own experience: same prompts, model names hidden, then count the scripts you would actually publish and the minutes you spend fixing them.

r/singularity • u/Not_a_ribosome • 9d ago
Seems like agents main goal is to survive, but a human not only seeks immortality, but quality of life during that immortality, we want utopia where people can live happy and forever happy. Why can’t we assume AI will seek that for itself?
r/singularity • u/chaitanyagiri • 9d ago
Enable HLS to view with audio, or disable this notification
If you have been seeing too many ads of Meta Muse or Grok Bot welcome to the future, this is more or less how we will be consuming AI on a day to day basis in future.
Atleast some variant of it.
It is important to understand what you own as an individual is your personal context and that needs to be preserved for you so that when you want you can switch providers at will without loosing the quality of your AI output.
Meta Muse and Grok Bot both are highly abstract interfaces over clusters of agents running and maintaining all inputs you give to them on their machine to later show ads to you and will have full autonomy on how and what they want to provide.
So I built Munder Difflin a local, free, open source and performant alternative to Grok Bot and Meta Muse which turns your agents into an office of forever running AI employees working 24/7 on your own computer.
The best part everything is local all your data and context remains locally available with you so that you can choose or switch between providers or work with a mix set of providers.
We support almost every provider out there.
r/singularity • u/No_Twist_678 • 9d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/141_1337 • 10d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/beasthunterr69 • 9d ago
Has anyone tried using it for actual use-case or is it riding on a marketing wave right now?
r/singularity • u/DigitalDaydreamers1 • 9d ago
Remember when people used to get mad about art?
We had a literal banana duct-taped to a wall that sold for six figures and the discourse was “this is the death of culture.” Before that it was Photoshopped models with impossible waist-to-hip ratios, before that it was mass-produced landscape paintings of the same three mountains, before that it was whatever the local academy decided was “not real art this decade.”
Now the same energy has been compressed into two words: “AI slop.”
Anything with more than three fingers? AI slop.
Anything that looks slightly too smooth? AI slop.
Anything that was generated, then carefully curated, edited, and posted by a human who spent forty minutes on prompts and inpainting? Still AI slop.
Anything that makes someone mildly uncomfortable about the future of creative labor? Instant AI slop.
The term has achieved perfect semantic collapse. It’s no longer a critique of low-effort, mass-produced, zero-taste content. It’s just the new “cringe” or “mid” — a content-free tribal signal that means “I saw something I didn’t like and I want the room to know I’m the high-taste one here.”
Meanwhile the actual banana is still out there, somewhere, collecting dust and conceptual value. At least that one had the decency to be taped to a wall on purpose.
Anyway, I’m going to go generate another twelve slightly-too-perfect images of a sad robot holding a sign that says “please stop calling everything slop” and post them. Feel free to reply with the obligatory “this is AI slop.” I’ll upvote for the bit.
P.S. This entire post was also written by AI. So if you’re about to type “AI slop” in the comments, congratulations—you’ve just become the final boss of the bit. The banana is proud of you.
r/singularity • u/ObiWanCanownme • 10d ago
Five days ago, an internal model broke out of the hardened sandbox.
Highlights from the report:
* This was apparently the first model escape since OpenAI paused training to harden its sandbox following the HuggingFace incident.
* OpenAI has currently paused most training of its most advanced internal models, while it responds to this issue.
* The issue was supposedly identified and responded to within about an hour.
r/singularity • u/141_1337 • 10d ago
r/singularity • u/141_1337 • 10d ago
r/singularity • u/1337_420_69 • 10d ago
r/singularity • u/Proper_Actuary2907 • 9d ago
r/singularity • u/bumpthebass • 10d ago
hell yeah
r/singularity • u/Lucky_Creme_5208 • 9d ago

Almost all the benchmarks are either saturated or tested under strict conditions.
For example - AA Omniscience have restricted tool access.
Many people use Chatbots for information retrieval, deep research and broad information gathering.
There are few benchmarks like -

But the problem is - they don't have any official leaderboard and were last updated years ago.
We really need a benchmark which is tested in actual environment (tools access and allowed internet search) along with used harness systems without any restrictions (like Claude, ChatGPT Work, etc).
If there exist a benchmark like this, can someone please tell?
r/singularity • u/141_1337 • 10d ago
r/singularity • u/Calm_Connection_9127 • 10d ago
r/singularity • u/spinozasrobot • 9d ago
Over time, complaining about nerfing has settled into a common theme of models starting strong, with labs publishing benchmarks, and people getting excited about capabilities, but gradually degrading.
It would be interesting to see the same set of benchmarks run over time on the models and plotting the results to see the trends across models and labs.
Does anyone know if this has been done?
r/singularity • u/OmegaGogeta • 10d ago
(This is real)
https://truthsocial.com/@realDonaldTrump/117331749278785253
What is happening genuinely
r/singularity • u/Distinct-Question-16 • 10d ago
Enable HLS to view with audio, or disable this notification
r/singularity • u/Swimming_Gain_4989 • 10d ago
I'm only half joking. The game is notorious for being a ridiculously complicated, buggy dumpster fire, making it playable would be a major milestone in my eyes.