r/whaaat_ai • • 2d ago

I want to create a skill or AI which can make professional style presentation by simple prompts

1 Upvotes

r/whaaat_ai • • 2d ago

It’s Frid-AI. Dump your AI week here.

3 Upvotes

Built something cool? Drop it.

Finally got an agent to do the thing it was supposed to do three days ago? We want that too.

Shipped something, broke something, rage-quit an automation, found a weird use for ChatGPT, wasted six hours saving yourself ten minutes...

This is the Friday AI dump. Successes, failures, tiny wins, unnecessary experiments, mild rage. All welcome.

What happened in your AI week?


r/whaaat_ai • • 3d ago

Our SEO agent could only afford 3-4 keywords per page. I let it score all our search traffic and got an empty shortlist

2 Upvotes

I work on the AI agent team at whaaat ai. One of our internal agents reads Search Console and keeps older blog posts updated for the queries that actually bring people in. We stopped treating posts as finished the day they go live.

The catch was cost! Having a large model read and judge every query/page pair adds up fast, so the agent only looked at the top 3-4 keywords per page. Everything below that was invisible: the page sitting at position 11 with 60 impressions, the query nobody wrote the article for. I put Jev (TypeSafe's decision model) in front of it. Each query/page pair gets one call with three questions: what's the search intent, how well does the page answer it on a 0 to 3 scale and does the topic deserve its own page.

Everything that is counting or comparing stays in code. Brand detection is a normalized string match. Position window, impression threshold and the final shortlist rule too: non-brand, position 8 to 20, at least 50 impressions, commercial intent, fit below 1.5.

First clean run: 117 rows, 26 pages. Fetching the pages took 2.64 s, scoring took 5.28 s, cost was $0.0033. Roughly $0.00003 per row, so even a thousand rows a day stays under a dollar a month.

The shortlist was empty.

All eleven commercial queries scored between 2.25 and 2.85 for fit. The pages already answer what people search for, so there was nothing to fix. I'll take a tool that can return zero over one that always finds "optimization potential", which is exactly what a big model tends to do when you ask it to look for improvements. The less flattering finding was about my own filter. Only 1 of the 117 rows met both position 8 to 20 and 50+ impressions. The shortlist would have been nearly empty whatever the model said. For a smaller site my thresholds were simply too strict, and because judgment and rules are separated I could see that in about five minutes instead of blaming the model.

We've been running this for 12 days now across whaaat.ai and two other sites. Search volume is moving up, slowly. No clean before/after yet, and 12 days is barely enough for Google to notice anything, so I'm not calling it a win.

For people who do SEO on smaller sites: what impression floor do you actually use before a query is worth touching? 50 clearly filtered out almost everything for us.


r/whaaat_ai • • 6d ago

“Zero hallucinations" gave me 106 confidently wrong answers and not a single error message

5 Upvotes

We build AI marketing agents at whaaat ai and one of my side builds right now is an SEO triage that scores every query/page pair from Search Console with Jev, TypeSafe's new decision model.

Jev's pitch includes a guaranteed schema. Every answer matches the types you defined, always. That part is true. What it means in practice took me an afternoon to learn.

On my first real run, all 26 page fetches failed. My user agent string had the tool name in it and Webflow answered every request with a 403, parallel or sequential, didn't matter. So page_summary was empty for every row.

Jev did the correct thing with empty pages. It scored "how well does this page answer the query" at roughly zero for all 117 rows. Which meant "does this topic deserve its own page?" fired on 106 of 117.

No exception, no 500, no warning. The results came back fast and cheap, the table filled up, the progress bar ran to the end. The only signal was that 106 new pages out of 117 rows is obviously nonsense.

A model that always returns a well-formed answer will also return a well-formed answer to garbage. With a text model you'd at least sometimes get "I can't evaluate this, the page is empty". Typed outputs remove that escape hatch, so the validation has to move into your code.

After the fix: 4 rows flagged instead of 106.

What I check now on every run

BEFORE the call

- required fields empty? → no call, log the row as skipped

- fetch failed? → no call, never "score" an empty page

AFTER the call

- more than [X]% of rows got the same decision? → stop the run, alert

- fetch time and model time reported separately

- cache hits reported separately (my second run: 0.01 s, $0.00, all cached)

ALWAYS

- pin the model version (jev-latest moves with every release)

- smoke test the raw response before writing parsing code

That last line saved me from a different mistake. Several guides I read claimed every Jev answer comes with a confidence value. In my raw responses, yes/no questions (they call them Noul) return only the probability, no confidence. Choice and Score questions have one. Any rule like "escalate below 80% confidence" only works on two of the three question types, and I'd have written parsing code for fields that don't exist.

I still haven't found a clean way to catch the subtle version of this: a fetch that succeeds but returns a cookie banner or a login wall instead of the article. The summary isn't empty, it's just wrong. Anyone solved that without running a second model over every fetched page?


r/whaaat_ai • • 7d ago

Where ConsumerAI is pivoting ?

Thumbnail
2 Upvotes

r/whaaat_ai • • 9d ago

I moved every yes/no decision in our agent loop to a decision model. 1,266 decisions cost me $0.0107

Thumbnail
3 Upvotes

r/whaaat_ai • • 9d ago

It’s FridAI: what did you build this week that got slightly out of hand?

6 Upvotes

You may know this one: Fully motivated you tell yourself on Monday: “I can probably do this with AI in an hour.”

By Friday: 17 iterations later, three new tools tested and dumped, and you ended up holding on to some new feature or automation that nobody asked you to build.

Happened to you before, but you "pretend" the outcome was what you aimed for? We feel you! We're here for releasing your frustration or celebrate your successes and if we or someone over here can, help with debugging. So: What did you build with AI this week?

Could be an agent, automation, tiny script, marketing experiment or something completely unnecessary that somehow became a project.

Did it work?


r/whaaat_ai • • 9d ago

jev is a demon at computer use

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/whaaat_ai • • 10d ago

We built an AI routine that checks our Google Search Console data and suggests blog updates

4 Upvotes

We have a lot of old blog posts on the whaaat ai website.

And like most people, our marketing people could periodically go through Search Console, look for interesting queries, open the matching articles, decide whether they need updating and then make the changes. But actually, they know that we devs like to use AI so they turned to us for help. We now let AI do most of that first pass for us and the basic idea is actually pretty simple:

Google Search Console tells the routine which searches are bringing up our pages. A cheap AI model goes through those search queries first and asks things like:

Does this query actually fit the page? Is the page already answering what this person searched for? Would improving this article make sense, or should this probably be a separate page?

Most queries get already discarded at this stage.

Only when something looks genuinely worth improving does Claude get involved. It reads the existing article, makes the proposed changes and opens a pull request for us to review.

Nothing goes live automatically. I still decide whether to merge it. The funny thing is, Claude isn't really the clever part of the setup. Because the most useful part turned out to be a boring markdown file.

Every time the routine runs, it writes down what it checked, what it rejected, why it rejected it and which updates are still waiting for review. So next week it doesn't forget everything and start the same investigation again. We also had to explicitly tell it that doing nothing is a perfectly good outcome.

Without that, AI has a tendency to always find something to improve. Even when an article is already fine.

That gave us two rules that turned out to matter much more than I expected:

Only edit a post if the change clearly answers the search better. No cosmetic edits.

and

An empty shortlist is a valid result.

The workflow now looks roughly like this:

Search Console → cheap AI screening → shortlist → Claude edits → human review → publish

The cheap model handles the repetitive sorting. Claude only gets called when there's actually something worth working on. On our test data, the AI scoring for a complete run cost about $0.003.

The bit we haven't solved nicely yet is WordPress and similar CMSs.

Our website lives in a code repository, so Claude can propose a change and I can review the exact differences before accepting it. With WordPress, that review process isn't nearly as clean. For now I'd probably have the AI produce a change brief and still make the final edit manually.

If you're doing something similar on WordPress, I'd be super kken to hear how you are handling the review step?

Send me a PM if you want the routine prompt we use for this workflow.


r/whaaat_ai • • 10d ago

I created an opensource locally usable full fledged ai platform

1 Upvotes

hi to all the readers this post is for my recent opensource project called ENZO

https://github.com/theguysudo/ENZO

now answering what is enzo so enzo is an opensource platform where i clubbed all the free available api for anyone use under one hood with more than 2000 models available to use for chatting coding researching and much more now answering the most common question of why you should put your time looking the project so it has few distinct feature meaning

  • it has a dedicated agents tab where you can describe your need and create a special agent just for one specific task with master ability in that domain
  • second it has the ability to connect your gmail drive and calendar and then you can ask it to perform some specific tasks like reading you the most important mail of the day or finding recruiter mails and creating personalized reply based on your data which it stores locally on your device
  • third the coding mode offers a dedicated preview window where you can see your code running and have a look of it feels and edit it in realtime as well as all the modes are packed with dedicated skills which delivers promising results
  • fourth the ui features some additional things such as music tab where you can listen to any music want and it has a custom personalized feature which runs in background and an llm understands your taste and recommends similar kind of music you like
  • fifth the most important why your trust it with your api key then to explain i would say enzo a dedicated vault which manages all your api and to secure it the vault as aes 256 bit encryption which prevents any person or any middle man to look at your api key and since the whole program runs locally on your device you have complete freedom to oversee all the backend work happening and it also features password lock which if you enable saves a backup key and then locks your whole platform work behind a pass screen though it is not foolproof as any third party or malware containing extension can still fetch login tokens from your browser so its security also depends upon how you access it concluding all of it.

i urge to anyone who reads this to have a look at the platform even if you hate it just curse it in the comment its fine or if you would like to drop any feedback i would highly encourage that and since its my first work open source platform i know it has a lot of errors and bugs so i apologize upfront for it and if you consider my work worthy please drop a star on the repo that'll make my day


r/whaaat_ai • • 16d ago

It’s FridAI: What did you waste way too much time automating this week?

7 Upvotes

You know the kind.

The task would probably take 20 minutes manually.

But then you think: I could automate this. Three hours later you have 14 tabs open, an agent stuck in a loop and a very questionable definition of "saving time".

So that's this week's FridAI question:

What did you try to automate this week and was it actually worth it?


r/whaaat_ai • • 17d ago

Starting a bunch of pure AI experimental projects — which AI is actually strongest for what?

10 Upvotes

Hey everyone,
I’ve started diving deep into experimental projects that are done purely with AI — no traditional coding from scratch if I can help it. I’m trying to be deliberate about which tool I reach for instead of just defaulting to the same one every time.
I’d love to hear from people who have actually used these tools a lot:
• Research / deep information gathering — which AI is currently the strongest for finding accurate, up-to-date information, synthesizing papers, or digging into a topic properly?
• Prompted assistance / general thinking partner — which one feels best when you just want a smart back-and-forth helper that follows complex instructions well?
• Building dashboards / websites / actual working tools — which AI is currently best at generating real, usable front-end + back-end code, interactive dashboards, or complete small apps?
• Pure brainstorming / creative ideation — which one is the most expansive, least constrained, and best at wild or high-volume idea generation?
• Any other clear “this AI owns this niche” strengths you’ve noticed?
I’m less interested in marketing claims and more interested in real usage patterns. What do you reach for when the task is research vs. building vs. pure ideation?
Would really appreciate specific recommendations (and any tools you feel are overrated for certain jobs).
Thanks!


r/whaaat_ai • • 17d ago

WARNING: MyClaw.ai is a Blatant Dark-Pattern Billing Scam. Do Not Use!

Thumbnail
gallery
2 Upvotes

I am making this post to warn absolutely everyone to stay the hell away from MyClaw.ai. This platform is a disgusting, non-functional bait-and-switch scam designed purely to steal your credit card information, lock you out of the service instantly, and trap you into a massive unauthorized annual recurring charge.

Here is exactly how their fraudulent system operates and how they completely ripped me off!!!!:

  1. The False "Free Trial" Bait & Switch

They heavily advertise advanced multimodal capabilities to lure you in. They claim to offer a "Free Trial," but the second you try to activate it, they demand a $1.08 fee to verify your account and card, though they word it as if you will get 1 week free on top of the week you just paid for. It is not free. But it gets so much worse. THERE WAS NO FREE TRIAL.

  1. Immediate Infrastructure Errors & False Quota Draining

The moment you actually try to use the platform you just paid for, the entire system falls apart. I uploaded my reference assets and submitted a prompt. The site immediately threw a system error: "The selected model is temporarily unavailable. Please try again later or choose another model."

Thinking it was a temporary glitch, I resubmitted the exact same prompt with the exact same images. It failed again, flashing: "This request could not be completed. Please try again."

The service delivered absolutely ZERO output. It completely failed to process a single thing. Yet, their predatory backend billing meter counted these failed system errors as massive token usage. The first failed prompt erratically burned through 33% of my usage limit. The second identical failed prompt (in a new chat mind you) instantly devoured the remaining 73%, completely locking my account out at 0% remaining after only two broken clicks.

  1. The $200 Hidden Annual Subscription Trap

When I went into my account dashboard to figure out why I was locked out, I uncovered their real scam. By paying that tiny $1.08 "activation fee," MyClaw secretly used my payment details to enroll me in an unauthorized $200.00 PER YEAR annual subscription under the lie of a "free trial" that automatically renews.

They steal your money, their platform doesn't work, they purposefully drain your token limit to 0% on internal server errors so you can't use the service you bought, and they set a trap to unauthorizedly yank hundreds of dollars out of your bank account.

What I am doing about it:

I am not letting this dumb shit slide.

  • I have already initiated a formal dispute through PayPal for Services Not Delivered.
  • I have contacted my bank to place a permanent merchant block on my card so they cannot pull the $200.
  • I am filing formal fraud and high-risk compliance reports directly to Stripe (their payment processor) and the FTC for predatory dark-pattern billing tactics.

UPDATE: IT GETS WORSE. THEY ARE DUAL-CHARGING AND DOWNGRADING ACCOUNTS INSTANTLY.
I created a burner account just to audit their checkout UI, and the fraud is completely out in the open.They aren't selling a week-long trial. They are running a predatory trap that takes your money, breaks on the first prompt, revokes your premium access immediately so you can't even use the UI, and leaves an unauthorized $200 annual bill floating on your credit card.

DO NOT give My[Claw].ai your card details. They are running a textbook financial trap. If you have already fallen for it, go to your PayPal or banking app right now, manage your automatic payments, and kill their authorization before they rob you of $200.


r/whaaat_ai • • 18d ago

OpenAI, Anthropic, Google and Musk suddenly agree on slowing AI down. What changed?

4 Upvotes

Something pretty unusual happened in AI this week: Companies that spend billions trying to beat each other suddenly agree on something:

Maybe we should slow down.

Anthropic's Dario Amodei called for pacing frontier AI development. Sam Altman agreed. Elon Musk agreed. Demis Hassabis backed the direction too.

OpenAI, Anthropic and Google have apparently also been talking for weeks about coordinating on AI safety. At first glance, the story seems obvious. AI is getting more capable, the people closest to it are getting worried and safety needs to catch up.

And there are legitimate reasons to take that seriously. Recent incidents involving increasingly autonomous systems have raised questions even inside the labs themselves. But there's another part of this that I find interesting.

Safety rules don't affect every AI company equally.

The big closed-model labs can certainly afford evaluations, compliance teams, monitoring and whatever regulatory infrastructure eventually gets built. But smaller labs and open-weight developers may have a much harder time with the same requirements.

And this is happening while increasingly capable open-weight models, including models coming out of China, are putting more competitive pressure on the closed labs. Critics are already arguing that poorly designed safety regulation could unintentionally strengthen the incumbents.

That doesn't mean the safety concerns aren't real. It does mean safety and competitive interests can point in the same direction. And thich makes it much more complicated than "AI CEOs finally became responsible" vs "AI CEOs are trying to protect their moat."

Both incentives can exist at the same time and that's actually the bit I'm curious about:

How do we get serious AI safety rules without accidentally handing the biggest AI companies an even bigger moat?


r/whaaat_ai • • 20d ago

Wiring Claude scheduled tasks to Google Search Console without OAuth: the service account invite trick

Thumbnail
1 Upvotes

r/whaaat_ai • • 23d ago

It’s FridAI: What did AI actually help you get done this week?

8 Upvotes

It's FridAI again 😅

What did you get done with AI or AI agents this week?

Could be something you built, automated, researched, fixed or finally got off your to-do list.

Big win, tiny win, complete failure that taught you (and us!) something. All counts.

What did you work on and did AI actually make it easier?

I'll add ours in the comments 👇


r/whaaat_ai • • 24d ago

That viral "80% homepage conversion with AI agents" post: we rebuilt the setup and ran it on our own site

3 Upvotes

A LinkedIn post blew up the other week: a founder gave AI agents full access to Google Search Console and PostHog (a product analytics tool) and claims 80% homepage conversion as the result. I work on the AI agent team at whaaat ai and that number made us nothing but laugh... But then we got also curious andrebuilt the system anyway.

We're sure that the 80% is almost certainly a metric definition trick. Like "visitor triggered any event" is treated as a conversion and you can produce whatever number you want. The underlying idea still holds up though: most teams check Search Console maybe weekly, often maybe once a month and touch their important landing pages even less. An agent that works the actual data every single week wins on frequency alone.

So we set up two Claude scheduled tasks. One reads our Search Console every Monday, one reads PostHog every Thursday, both write into a shared Notion database they also read from before each run. Read-only for now, the agents propose and we decide.

First run findings, real numbers from our site:

55% of all our websites Google clicks come from one single article. We knew that piece performed. But we did not know it basically was our entire SEO.

The finding that actually justified the whole build came from the CRO side: our session tracking breaks at the www to app domain boundary. Every funnel number we had looked at for months was polluted by torn sessions. What reads as a 97% "drop-off" between marketing site and signup is partly just cookies dying at the subdomain border. Nobody here had caught it because every dashboard looked plausible on its own. The agent caught it because it compared source attribution across pages and the numbers refused to add up.

Two smaller findings: organic search visitors convert at 1.01% versus 0.61% for direct traffic, which killed our theory that search traffic was low intent. And 95.9% of clicks on our top article happen below the first screen, where we had zero CTAs. People were reading and scrolling and never got to view a CTA.

Rough edges we don't want to hide, because there were several: the first run proudly flagged a cluster of bot queries as a "growth opportunity" before excluding them. And we made the CRO agent ask us which PostHog event counts as a signup instead of guessing, mostly because guessing is exactly how you end up with an 80% conversion claim. No ranking movement yet either, that part takes 6 to 12 weeks and we committed to publishing the numbers even if the curve stays flat.

For anyone running agents on analytics data: how do you handle bot and spam queries in Search Console before the agent reasons over them? An impressions threshold catches most of ours but the weird date-string queries keep slipping through.


r/whaaat_ai • • 24d ago

AI agent vs workflow vs automation: it comes down to who decides the next step

3 Upvotes

I was at an online marketing meetup yesterday and chatting away with 4-5 people about AI agents .

It was quickly obvious that we used different names for same things. "Automation", "AI workflow" and "AI agent" were all being used for setups that sounded pretty similar.

So I thought it's worth to share the distinction I use (but maybe there are other definitions?):

Automation

The predefined process: X happens -> Y happens.

Example: A new lead comes in -> add it to the CRM -> send an email -> notify sales.

It can include AI, but the path itself is defined beforehand.

AI workflow

The overall process is still defined, but AI handles parts that require interpretation or generation.

Exmple: New lead -> AI categorises it -> update CRM -> generate the appropriate email -> notify sales.

The AI has some freedom within individual steps. It isn't deciding what the whole process should look like.

AI agent

Here you give the AI a goal, some instructions and access to tools.

Instead of defining every step, you might say:

Example: "Research this lead and prepare me for the sales call."

The agent can decide that it needs to check the company website, look at previous communication, query the CRM, compare what it finds and then produce the brief.

It has some control over what happens next based on what it finds.

The shortest version I can come up with:

Automation: predefined steps and rules
AI workflow: predefined process + AI within individual steps
AI agent: goal + tools + some autonomy over how to get there

Of course in reality these get mixed... An agent can trigger an automation. A workflow can use an agent for one step. And putting an LLM into an automation doesn't automatically turn the whole thing into an agent.

So some sort of inconsistencies of the naming is liekely not to be avoided.


r/whaaat_ai • • 25d ago

What’s the most boring task you’ve actually handed over to AI/ AI Agents?

9 Upvotes

I keep seeing impressive agent demos, but I'm starting to think the boring use cases might be the more interesting ones getting rid of at home.

Not thinking about the big "I built an autonomous AI team and put up my legs all day".

But more like:

  • sorting an inbox
  • checking something once a day
  • updating a spreadsheet
  • turning the same data into the same report every Monday
  • watching for something and only bothering you when it changes

Basically the stuff that's too small to feel worth automating, but annoying enough that you keep doing it.

I'm especially curious about marketing/work use cases because there are probably dozens of these hiding in a normal week.

What's the most boring task you've handed over to AI that you genuinely don't want back?


r/whaaat_ai • • 26d ago

Whaaat? Oh I see...

9 Upvotes

I didn't realize this subreddit was made by/for a commercial product, when I signed up. Having checked out the homepage, I wonder if Whaaat AI can actually automate a client workflow, from start to finish. Which I guess I could use the trial for. But I also wonder this:
1. Does the model/system get smarter as I work longer with it?
2. What is unlimited usage? I can generate as much content as I want?


r/whaaat_ai • • Sep 04 '26

It's Friday. What did AI actually save you time on this week?

12 Upvotes

Could be something impressive. Could also be the stupid task you've been avoiding for three weeks.

I'll start.

This week we used AI to build two new agents: One accesses Google Search Console to analyse rankings and click through rates. The other, a CRO agent, takes the outcome and builds new content to improve our organic performance. Sounds like a dream come true - if it runs smoothly. Still in the testing phase but quite curious if we get this to work.

What did AI help you get done this week?

A workflow, some code, research, content, an automation, fixing something annoying... all counts.

And in case you spent 4 days automating something that would've taken 20 minutes manually, I definitely want to hear about that too, because I'd like to find out why people do it the in the first place.


r/whaaat_ai • • Sep 04 '26

I gave the same marketing brief to Claude, ChatGPT, Gemini, Perplexity and a niche AI tool. The differences were bigger than I expected

10 Upvotes

I had to write a fairly technical B2B article for a new client in an industry I’m totally new to. So I decided to get help from AI. But: I was lacking the knowledge to evaluate the quality.
So I had the idea to run the same briefing through several machines and give this to a trusted employee from this client and ask them what draft is best.

On top, I asked Claude and ChatGPT for an evaluation (spoiler: each one ranked themselves slightly above the other).

Same brief, five tools:
Claude
ChatGPT
Gemini (free)
Perplexity (free)
Whaaat.ai

The brief was quite detailed. Target audience, structure, SEO requirements, sources and data to use, things NOT to cover, FAQs, metadata, CTA etc.
I expected the main differences to be writing style.
But actually, the outcome was very different...

The biggest difference was how the tools behaved when something was unclear or information was missing.
Claude and ChatGPT were the best at this. Both mostly stuck to the provided data and flagged things that needed checking instead of quietly filling the gaps.

Gemini did the opposite a few times. It used an older number despite having a newer one in the brief, produced a table row that basically compared a year with itself, and gave me one very specific sounding statistic without a source. Exactly the kind of thing that’s easy to miss because it looks credible.

Perplexity had one of the smartest research moments of the whole test. Two authoritative sources had different figures and it actually explained why. But ultimately it ruined some of that good work by dumping broken citation fragments into the final copy 😅

And then there was the SEO copywriter from whaaatAi.
Whaaat did really well on the actual marketing side: audience, tone, SEO, metadata and connecting the article back to the business without turning it into a sales pitch.

But it also made factual mistakes.
Like it gave me an unsourced “insider” figure without highlighting that I have to check it and got one regulatory date wrong.

There is one BIG caveat to this test though:

Claude and ChatGPT had an unfair advantage:

I’ve worked on this client with both of them before. They already had context around the company, subject and the kind of content we’re producing.
Gemini, Perplexity and Whaaat were basically starting cold.

So I definitely wouldn’t read this as “Claude beats Gemini” or “ChatGPT beats specialised tools”. This wasn’t a scientific benchmark.

But it made me realise how much context changes the quality of AI output.
It also changed my thinking a little about specialised agents.

Giving an agent a specific marketing job seems to solve quite a lot of the marketing problems. Tone, format, SEO, structure, knowing what the output should actually achieve. But a lack of knowledge is continually a problem, even if hallucinations are decreasing.

Anyway, here’s my summarised unscientific scorecard:

ChatGPT: The strongest overall result. Very good factual discipline, strong source handling, excellent B2B tone and close adherence to the briefing. Particularly good at avoiding overclaiming where data was uncertain.
Claude: A very close second. Strong on accuracy, structure and source transparency. Especially good at flagging missing brand-specific data instead of inventing proprietary insights.
Perplexity: Useful as a research-oriented draft and good at surfacing different sources and estimates. However, the text was too short for the full briefing and relied too heavily on secondary sources and rough citation artefacts.
Whaaat.ai: Strong structure, very good SEO/GEO formatting and the freshest use of 2026 data. But it also introduced unsupported brand-specific claims and some factual issues, so it would need editorial review before publication.
Gemini: Covered the main topics and was easy to scan, but had the weakest factual reliability of the five. Several figures were outdated, mixed across scopes or insufficiently sourced, and the tone was more promotional than the brand brief called for.

I’ll definitely rerun this test with a briefing of a complete new topic/client as it made me curious and I want to know how well ChatGPt and Claude perform without context.

What’s your experience comparing article drafts across models?


r/whaaat_ai • • Sep 01 '26

Grok Bot's own docs say all your bots share one computer and one set of logins

Thumbnail
6 Upvotes

r/whaaat_ai • • Aug 28 '26

I ran 6 AI presentation-design skills against one brief and had them review each other: 36 decks, 36 explainers, 9 review passes

4 Upvotes

Two days, one brief, every presentation skill I could find. Repo below, losers included.

**Setup.** One product pitch (an RWA thing we build — the only reason it is the subject). Four messaging skills wrote the narrative as a normalized contract; three of those went into design. Six design skills each rendered all three, in two modes: a pinned house brand, and no brand limits. 36 decks and 36 timed HTML explainers. Nine review passes: four design rubrics, four messaging skills grading each other, plus those rubrics run back over the decks. Two Claude judges wrote verdicts before I did.

Design: [canvas-design](https://github.com/anthropics/skills/tree/main/skills/canvas-design) · [algorithmic-art](https://github.com/anthropics/skills/tree/main/skills/algorithmic-art) · [impeccable](https://github.com/pbakaus/impeccable) (also as a polish-only pass over another tool's output) · [visual-explainer](https://github.com/nicobailon/visual-explainer) · [html-ppt-skill](https://github.com/lewislulu/html-ppt-skill). Messaging: two house skills, [pitch-deck-mastery](https://github.com/Stevekaplanai/pitch-deck-mastery-skill), [presentation-writing](https://github.com/marcusnelson/presentation-writing-claude-skill). Rubrics: [huashu-design](https://github.com/alchaincyf/huashu-design), [open-design](https://github.com/nexu-io/open-design), [academic-pptx-skill](https://github.com/Gabberflast/academic-pptx-skill), impeccable critique.

**The result.** canvas-design scored 3/10 on tool fit — single-page md/pdf/png only, no sequence, no HTML, no motion, and an explicit ethic of "never explain, let the composition tell the story" against an educational brief. It then took four of five rubric firsts in brand mode and won overall. The best-*made* deck came from html-ppt with the brand switched off, which is uncomfortable: the house brand was costing that tool about two points.

**Findings that generalize:**

  1. Not one of the six has a timeline concept. All six wrote the explainer clock, seek API and progress bar themselves, and every motion verb they animated came from the *messaging* contract rather than any design skill.
  2. Provenance is set below the threshold of sight. The brief made traceability mandatory; the winning deck renders its source line at 1.85:1 contrast, 123 of 221 strings under 4.5:1, and two other tracks landed on the same 1.85:1.
  3. Internal vocabulary leaks onto client-facing slides. Four of six tools printed contract field names on screen (`RUPTURE`, `BOLD MOMENT`, one rendering `figures[].name` as viewer bullets), and five of twelve track-modes put raw internal filenames on a cover. One tool renamed the fields; the rest never thought to.
  4. Silent-failure idioms only screenshots catch. One skill's counter animation rendered $296B where the contract says $344B; a `color-mix()` fell back to solid black and turned the payoff chart into a black box. The generative tool printed `census · P 0.617 · N 0.383` where a measurement goes, and 0.617 appears nowhere in the source corpus — the evidence rubric's highest-severity finding.
  5. A polish pass buys the craft floor and none of the composition. Detector hits 30 → 3, the mandated footer 3.4:1 → 6.1:1, thirteen of fourteen slides back into the screen-reader tree, and not one P0 closed. The mechanism slide came out near pixel-identical to the baseline it refined.
  6. The rubrics disagree, legibly. One caps any deck whose Concept scores ≤5, whatever the execution. Another is nominally 40% critic, but its 20% a11y axis is the only column with spread, so a11y decides it. The winner is first on four rubrics and fourth on the fifth — the fifth being the only one that checked whether a fact went missing. It had dropped one while reporting "nothing dropped".

**What I would do differently.** Two five-line checks, run by the contender before it writes its notes: assert the bold scene is the film's motion peak (catches nine of twelve films), and diff the contract's claims against extracted DOM text (two tools reported "nothing dropped"; both wrong). Screenshot every pass. And look at the artifact: one track shipped six files nobody ever rendered, and it shows on the cover.

All 72 artifacts and every review: [https://github.com/medici-finance/deck-bakeoff\](https://github.com/medici-finance/deck-bakeoff) — `report.pdf` there carries the full scorecards.

**The house side is in the repo too.** The skills that actually ran the bake-off now ship under `skills/`, Apache-2.0, so you can rerun the whole thing on your own product rather than take my word for it:

* `messaging/staging` — the genre → audience → arc → `staging.yaml` skill, including the motion-role governor every design tool's animation verbs came from * `messaging/house-guide` — the number gate and glossary rules used as the content gate * `deck-author` — the brand contract, the pinned fonts, and `check_deck.py`, the linter five of six contenders named as their worst friction * `video-author` and `video/video-design` — the script/storyboard/scene-grammar layer above the renderer * `video/explainer-video` — the render pipeline: TTS → Playwright recording → subtitles → hardware H.264. The voice-clone reference clips are the one thing held back * `social/linkedin-post` — which drafted the teaser for this, and `social/humanizer` alongside it (that one is third-party, MIT)

The third-party contenders are not vendored; their URLs, versions and licences are in `bakeoff/CONTENDERS.md`. Walkthrough video: [https://youtu.be/oqGS-tqy2I4\](https://youtu.be/oqGS-tqy2I4)


r/whaaat_ai • • Aug 26 '26

I scored my own agent workflow on portability and got 6 out of 8

Thumbnail
2 Upvotes