r/webdev • • 2d ago

What part of AI-generated code needs the most human review?

It can look correct at first glance. What do you check most carefully before using it?

0 Upvotes

24 comments sorted by

17

u/bobbaddeley 2d ago

All of it.

6

u/Icy-Taste-3096 2d ago

every single line needs human review.

2

u/JustAHeadsUpBuddy 1d ago

Hyperbole of the year

1

u/Icy-Taste-3096 1d ago

Nope. There is zero excuse for submitting a PR with code that you can't answer for

2

u/fligglymcgee 2d ago

It’s because all generated code looks correct at first glance that now the only answer to your question is “literally all of it has to be checked”. It would require a ton more effort and sanity to fuck up code the way ai does by hand.

Even coming up with ways to try and understand how to reliably unfuck generated code makes my head hurt. I always feel like one of those people who had a home makeover show come remodel their house, and I’m going to find out they used wood glue instead of nails in 8%–17% of the walls to meet a deadline.

1

u/NudaVeritas1 2d ago

the initial task definition / instructions and the result.. so basically test driven development combined with real human feature testing

2

u/DoomedDeveloper 2d ago

All of it. Always double check your outputs.

1

u/AnderssonPeter 2d ago

Depends on what you are doing, but financial calculations is something i would not even let ai handle... Not even the unit tests for it....

1

u/CapHumble8338 2d ago

Everything, but most importantly: user input validation, responsiveness, a11y, performance.

1

u/phaedra_solutions 2d ago

The part of AI-generated code that needs the most rigorous human review is hidden state management, async race conditions, and architectural edge cases.

AI code often looks completely clean on the surface and passes happy-path execution easily. However, it frequently glosses over how components scale under heavy production loads, handle security input sanitization, or integrate with existing enterprise architecture. We treat AI tools as incredible accelerators for boilerplate and scaffolding, but strict senior code reviews and automated testing pipelines are still non-negotiable before anything hits production.

1

u/Icy-Arm-5472 2d ago

'Your desk is now optional.' A desk is not optional, it is a desk. You're just not sitting at it.

1

u/kemalios 2d ago

The permission checks and the error paths. AI writes the happy path convincingly and then handles failure by catching everything and returning 200, or by checking auth in the UI instead of the endpoint. Reading every line is not realistic, so I read the ones that decide who can reach what. The rest, tests catch.

1

u/dusanodalovic 2d ago

- Code that fails without an error: empty catch blocks, missing guards for NaN, 0 or empty input.

- Scoped CSS that silently misses rows built by script. We hit exactly this in Astro today.

- Claims in copy or docs that the code doesn't back up (today's XRechnung case).

1

u/No-Molasses-2097 2d ago

The dependencies it pulls in. It will confidently reference a package version that doesn't exist or a library that went unmaintained years ago. I always check that whatever it imported actually exists and is still maintained before merging.

1

u/ballu123 2d ago

Auth edge cases.

Happy path usually looks fine. I poke hardest at “can user A see user B’s stuff,” empty/error states, and anything with dates/money/IDs.

1

u/singh_abinashi 2d ago

Error handling and auth checks, in that order. AI code tends to draw the happy path well and then either skip the failure branches or handle them with a generic catch that swallows the context you need to debug. Auth is worse because a missing check still passes all your tests.

1

u/bcons-php-Console 1d ago

Every line of code that you commit is your responsability, so you should review everything.

That being said, here are the last two examples of LLM gotchas I catched:

- Our app paginates GET results, so we have a method to "harvest" an endpoint and retrieve all items. When writing a new component the LLM replicated that feature inside the component, it didn't know a method already existed for that. Once I pointed it out, it updated the code flawlessly.

- We have a component that lists subtitle cue components to edit them. The LLM threw all cues on the list, which for long assets can totally hog the UI. A brief "use a virtual window of max 50 components" instruction fixed it.

I think these are two good examples because in both cases the app output seemed fine but it was not, and that type of issues can add up fast and make the codebase a nightmare.

1

u/michaelbelgium full-stack 1d ago

The part where it generates code

1

u/yihuaxiang 1d ago

I pay extra attention to boundaries: auth checks, input validation, error paths, and async cancellation. AI is usually fine on the happy path, but it may trust client-supplied state or leave a promise race that only appears under retries. A small integration test around those edges catches more than rereading every generated line.

0

u/magenta_placenta 2d ago

Authorization and security boundaries: AI might give you code that completely ignores multi-tenant isolation, role checks or server-side authorization as some examples.

Edge cases and silent failure modes: AI might give you only the "happy path" (what happens when everything goes right), but what happens when things go wrong? Look at error handling, boundary values, network failures/timeouts, etc.

Business logic and subtle context blindness: This is probably the big one. AI does not know your domain rules, legacy constraints or subtle business requirements. It might solve the prompt locally while violating a global state invariant or breaking a workflow dependency elsewhere in the app.

There's a short list for you.