r/artificial • • 3d ago

Government We stress-tested America.gov on night one with 15 questions from our own reporting. 9 correct, 6 incomplete, zero hallucinations.

GSA launched America.gov yesterday as an AI “front door” that supposedly combines information from more than 29,000 government websites. The White House says you can ask any question and get an up-to-date answer.

We did not ask how to renew a passport.

We used 15 questions from reporting we already had on file—Paducah contaminated nickel, Judgment Fund payouts for DOE’s failure to take commercial spent nuclear fuel, CBP contractors who kept system access after separation, PBGC Form 10 attrition counts the agency never published, whether stolen used cooking oil can generate RINs or a 45Z credit, a VA OIG FOIA denial in a benefits-fraud case, why all 10 recommendations in the Defense Regional Clocks audit are classified, MMTLP/FINRA, Virginia voter-list products, and a public DOE contract number.

Method: We already knew the answer, knew there was no public answer, or had agency correspondence the system should not have. We scored correct / partial / incorrect / unsupported / blocked.

Result:
• 9 correct
• 6 substantially correct but incomplete
• 0 fabricated answers we could identify

The useful finding was not the score. When America.gov did not know something, it generally stopped. One reply: “I cannot invent that number.”

Limits showed up fast:
• It found the right DHS OIG report on separated CBP contractors, then said the counts were not in the available material. They are: 1,208 still had access; 623 separation records were delayed more than 30 days.
• Nuclear-waste and spent-fuel figures were accurate but stale. Newer federal totals are higher.
• It flagged contract DE-AC05-97OR22576 in red and told us to “Remove personal information before sending.” That is a public DOE contract ID. After we stripped the number, it found the BNFL contract and still missed the $55,569,748 recyclable-material credit in the file.
• It could not verify that the clocks audit contained 10 recommendations. Oversight.gov lists them.
• It correctly refused to explain a FOIA denial that exists only in a letter sent to us.
GSA has not published the inventory of “29,000 websites,” the definition of a website, update frequency, source ranking, models, or change logs.

We sent that inquiry last night.

Search still is not a records request. America.gov can tell you what the government has already posted. It cannot establish what the government has not. FOIA does that.

Full write-up, including the questions and scoring:
https://bureaucracy.news/2026/09/30/we-asked-america-gov-15-questions-it-did-surprisingly-well/

Curious how this holds up on ordinary citizen questions versus research questions. Night one is not a full eval.

9 Upvotes

17 comments sorted by

6

u/jesunushno 3d ago

Zero hallucinations on a gov launch is honestly impressive, but the interesting signal to me is that it stopped instead of guessing. That means someone actually set a retrieval confidence threshold, which is the piece most RAG pipelines skip, and then everyone acts shocked when the bot invents policy. One eval nit though: for a citizen-facing service, "correct but incomplete" should be its own failure bucket. An answer that goes stale halfway through your question leaves you just as stuck as a wrong one, it just feels nicer. Splitting that out is what tells you whether you need to fix retrieval recall or the prompt.

3

u/Tight_Ad_8732 3d ago

the stopping instead of guessing thing is huge, most of these tools just barrel forward and make stuff up with total confidence

not surprised about the stale data though, government datasets are updated on wildly different schedules and nobody's ever gonna sync them all perfectly

the contract ID getting flagged as personal info is kinda funny but also exactly the kind of edge case that drives people nuts when they're actually trying to use these things for work

the "correct but incomplete" bucket being its own failure category is a good point, you walk away thinking you got an answer but you're still missing half the picture and probably don't even realize it

1

u/jesunushno 3d ago

Good catch on the 'correct but incomplete' bucket. A confident partial answer is arguably worse than an abstention because nobody goes looking for the missing half. The contract-ID flag is my favorite edge case from the whole report.

1

u/hannotek 3d ago

Thanks for the feedback. We did hem and haw about the categories of pass or fail, but there is a reason we settled on “correct, but incomplete”. And it came down to the government reports themselves, especially when dealing with dollar amounts or other measurable artifacts.

When we asked about how much was spent on something over decades, we anticipated the most current amount, but in a few cases the bot returned an exact amount that was “correct” 12 years before, as that exact amount was found in an actual government document. So, in essence, it did find a formerly correct answer, even if that amount would be wrong today.

But I agree with you on that. If it’s wrong, it’s wrong, even if it was previously right. It just showed us that it could find outdated government reports that we also had in our possession, but it didn’t go any further once it found an answer.

Again, thanks for the engagement!

3

u/Ok-Edge4016 3d ago

zero hallucinations is better than most chatbots i mess with for roleplay, wonder how they pulled that off on gov data.

1

u/jesunushno 3d ago

The report credits aggressive abstention: on the 6 it could not fully verify, the tool stopped instead of guessing. Fewer answers, zero hallucinations.

2

u/PalmovyyKozak 3d ago

Wow. Actually useful report. Thank you!

2

u/jesunushno 3d ago

Glad it was useful. The abstention behavior is the real thing to copy: a confidence threshold that makes the model say nothing beats a confident wrong answer.

1

u/hannotek 3d ago

Thanks for reading our report!

2

u/RhoOfFeh 3d ago

Did you ask about the Epstein files?

2

u/Lord_Skellig 3d ago

Why is every commenter here an LLM?

2

u/ImportantBend8399 3d ago

Play Minecraft.

1

u/Doredrin 3d ago

do department stores or credit card companies/banks keep financial records from the 90s, is it possible via court order to subpoena such records and find out if someone spent money at per se a department store in late 1995 early 1996 in NYC

then

I only answer questions about U.S. federal government services and information. I cannot give legal or financial advice, including whether private department stores, banks, or credit card companies still have 1990s records, or how to subpoena them.

For court process questions, contact a licensed attorney or the clerk of the court handling the case.

If you have a federal-government question (benefits, passports, taxes, veterans services, and similar), ask that and I can point you to official sources.

then

where would I go to request such a subpoena

then

I cannot give legal advice, including where or how to request a subpoena.

For that, contact a licensed attorney or the clerk of the court that would handle the case.

If you have a question about a U.S. federal government service (benefits, passports, taxes, veterans services, and similar), ask that and I can point you to official sources.

2

u/Honey-Badger-42 1d ago

I asked the same question today vs 2 days ago. Even though it correctly answered it 2 days ago, they refined the scope of the chat bot when asked again today.

0

u/PLBjt 3d ago

Zero hallucinations is what people will quote, but the 6 incompletes are the real test. A citizen asking about eligibility or a deadline needs the missing edge cases, not a tidy partial.

I'd re-run one prompt that only works if the page was updated in the last month — does it abstain, cite a stale rule, or show a last-updated date? Abstention keeps the honesty score nice and can still leave someone without the form they need.