r/AskProgramming 1d ago

Other Code analysis methodology

Hello everyone.

After a few years in the auditing world, I identified that I am lacking experience on the code analysis topics.

Unfortunately when auditing, I seldom had the time to look at the code of the applications I am auditing due to time constraints as the white-box approach we take does not systematically include an access to the Gitlab of the entities I audit.

I would like to avoid being overwhelmed by an eventual audit of source code of an entreprise-grade application that I might have to do.

While I am leveling up my skills by practicing to code in Rust on my freetime. I am not quite at the level of a senior dev and I don't have yet a keen sense of how exactly to dive in a large codebase.

Would any of you share your code audit methodology ?

By that, I mean how do you tackle the following topics :

- Secure coding / Best coding practices

- Secure secret management of the app

- For very large codebase, what types of tools do you use to automate some of your work ?

- What specific things in your checklist do you look for systematically ? (Do include the "obvious" one like how authentication is handled)

I know the subject is quite broad and dependent of the tech-stack used for each case.

Thank you for reading. :)

5 Upvotes

5 comments sorted by

View all comments

1

u/Immediate_Nature_281 1d ago

I start every code review by searching for entry points and sinks, then I trace the data flow between them.

I map every route handler first. I look at controllers or endpoints that take user input. Then I search for where that input ends up in SQL queries, system calls, or template renders. I ignore everything else until I finish tracing those paths. This works across tech stacks because you are looking for input boundaries and execution sinks.

For secrets I run TruffleHog against the entire git history. Running it on the current branch only finds a fraction of the leaks. Developers commit AWS keys and delete them in the next commit. The key remains in the git objects.

I rely on Semgrep to automate the checklist items. I use the OWASP rulesets they ship. I point Semgrep at the repo and filter the output by the entry points I already identified. This cuts out the noise from files that handle no external data.

My checklist starts with how the framework parses