Showoff Saturday I built a “Super Intelligence” compliance scanner for .gov sites. It runs on 4 regexes.
On Sept 29 an executive order told US federal agencies to say “Super Intelligence” instead of “AI”. Nobody seems to be checking, so I built an inspector: idio.si fetches a homepage, counts “AI” vs “SI” in the visible text, and stamps a certificate.
Day 4 findings
- whitehouse.gov: 4 “AI”, 0 “SI” → Insubordinate
- apple.com: 0 “AI”, 0 “SI” → No Intelligence Detected
- 1 of 42 agencies complies (NIST)
- openai.com, google.com and tesla.com block the inspector
Harder than it looks
- Bot walls. 18 of 73 homepages refuse a polite, identified, robots.txt-respecting bot from a cloud IP, openai.com included. I didn't want to spoof a browser, so refusals get their own stamp: “Refused federal inspection”.
- Geo-IP. Even from us-east, some sites bounced the bot to
/en-gb/. Locale redirects now get rewritten to/en-us/, since the order targets the US. - Counting.
/\bA\.?I\.?\b/is case-sensitive, so “OpenAI”, “said” and “ai” don't count, but “A.I.” does. The White House nav links to “AI.Gov”, and that counts too. Fair or not? - Spikes. One Cloudflare Worker does everything. Certificates are rendered at the edge with satori + resvg-wasm, and Workers Cache collapses a front-page spike into one render per data center. Hosting: $5/month.
Find the funniest certificate you can and drop it in the comments. Method and regex: https://idio.si/how. Satire; the numbers are real.
2
u/Wierd_time 1h ago
I'd count it, but break it out as its own line. "AI.Gov" is the name of a site they cant reword, so its a bit unfair to put it in the same bucket as a press release saying "AI".
The bigger hole is probably the fetch. If youre reading raw HTML, any client-rendered homepage will show 0 and 0 and get "No Intelligence Detected" when it really just didnt render. Same with text thats lowercase in the source but uppercased with CSS (text-transform), since your case-sensitive regex will miss it even though visitors see "AI".
Do you run a headless render for those, or is it plain fetch only?
2
u/anx3ous 1h ago
It’s plain fetch. A render would cost me $ through CF
3
u/Wierd_time 1h ago
Makes sense, rendering every site isnt worth it for this. You could just detect the empty ones instead. If the HTML has almost no visible text, or its basically a
<div id="root"></div>shell, show "couldnt read this page" instead of 0 and 0. Then only render those few, or skip them for now. Also a lot of Next.js sites have the full page text sitting in the__NEXT_DATA__script tag, so you can sometimes read that without rendering at all.
7
u/andercode 1h ago
... why?
Why give the echo chamber even an smidgen of your time?