Audit Code in Flask: A Technical Due Diligence Checklist

Swati Ahuja9 min read

audit code flask · technical due diligence code audit · code audit checklist · legacy code audit tool · flask security audit · software due diligence checklist

On this page

To audit code for technical due diligence, I split the work in two. First comes a 20-minute automated pass over size, structure, code health, hotspots, dead code and security patterns. Then the human hours go into a manual checklist: read every high finding, check licences and dependencies, look at tests and CI, and ask who holds the knowledge.

StepAutomated or manualWhat it answersWhat it showed on pallets/flask
1. Size and shapeAutomatedHow big is this, what are the main parts, where does execution start128 files, 1,063 symbols, 5 modules, 6 entry points
2. Code healthAutomatedWhich files are hard to change safely7.0 out of 10 overall, maintainability 7.7, static performance 9.9, 144 open findings
3. HotspotsAutomatedWhere change and risk overlap83 files ranked by change and fix history; 0 marked hot
4. Dead codeAutomated, then manualWhat can be deleted, and what only looks unused4 findings, 3 lines deletion-ready
5. Security patternsAutomated, then manualRisky calls in shipped code9 pattern matches, 4 high, all 4 deliberate on reading
6. Dependencies and licencesManual, with toolsSupply chain and legal exposureBSD-3-Clause, 6 runtime dependencies
7. Tests and CIManualCan the team change code without fear12-entry test matrix, Actions pinned to commit SHAs
8. People and processManualWho knows what, how releases happenInterviews, not tools

Scroll the table sideways to see every column.

We picked Flask as the worked example because it is small enough to read in an afternoon and carefully maintained, so an alarming finding there is more likely a tool misreading a deliberate choice. That is the lesson most due diligence reports need.

What a code audit for due diligence actually checks

We have all seen the technical due diligence checklist: architecture, code quality, security, team and process, licensing, yada yada. The agency guides that rank for this search are good on the areas and thin on the code itself. They say "assess code quality" and stop, or quote a two-week engagement without showing the first hour.

The first hour is where automation helps. A machine can rank 128 files, or 12,000, by how risky they are to change, but it cannot tell you whether that risk matters to the deal. That is why the audit splits in two, with an automated pass of about 20 minutes (most of it reading) and then a manual checklist where the judgement goes.

The 20-minute automated first pass

The numbers below come from the public pallets/flask page on repowise, indexed on 13 August 2026 at commit 2a8a38b0. Flask's main branch has moved since, so treat these as a snapshot. The index itself took about a minute, and the rest of the 20 minutes is reading five pages and writing down what to look at next, which is about as close to a done state as due diligence gets.

Step 1: size and shape

Start with the architecture page. It shows 128 source files holding 1,063 symbols across 5 modules, with 6 files acting as entry points into the dependency graph. The largest module is tests, with 52 files, and the map draws 79 files and 223 dependencies between them, leaving out files that have no import links.

That tells you how much there is to read, and whether tests are a real part of the codebase. In Flask the test module is the biggest one, which is a good sign before you have read a line.

Step 2: code health

The code health page scores every file from 1 to 10 using deterministic markers: complexity, nesting, large methods, duplication, error handling, and git signals such as prior bug fixes. repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. Flask scores 7.0 out of 10 overall, with maintainability at 7.7 and static performance at 9.9, and there are 144 open findings across the three pillars.

The useful part for an audit is the list of files by score. The three lowest are src/flask/app.py, src/flask/sansio/app.py and src/flask/cli.py, which hold the Flask application class and the command line interface, where a framework concentrates its logic. A low score there is a reason to read carefully and says nothing on its own about whether the code is broken, so I treat the score as a reading order.

It's fair to ask whether any health score predicts bugs, and we asked it too. Ours scored an ROC AUC of 0.74 for predicting next-six-month bug fixes across 21 repositories and 2,770 files (ROC AUC measures how well a score separates files that later needed fixes from the rest; 0.5 is chance). That is about the same as ranking files by size alone (0.737 vs 0.742), so the score's extra value is in naming why a file is risky. The rest of the caveats are in the validation study.

Step 3: hotspots

A hotspot is a file that changes often and has a history of fixes. The same page ranks 83 files this way, and for Flask, 0 files are marked hot. Nothing in the repo is both changing fast and collecting fixes, which is the pattern you want to see in a target.

In a busier product this table is usually the most useful page in the audit, because it shows where the next bug is likely to land. The hotspot guide explains the method if you want to reproduce it with git log yourself.

Step 4: dead code

The dead code tab found 4 findings: 2 unreachable files and 2 unused exports. Only 3 lines, in src/flask/debughelpers.py, come back as high confidence and deletion-ready.

The two "unreachable" files are examples/celery/make_celery.py and docs/conf.py. Nothing imports them, which is what the tool measures, but one is a script you start a Celery worker from and the other is the Sphinx configuration file that the docs build reads by name, so neither is dead. A person spots that in under a minute, and a tool that walks imports cannot, which is why docs/conf.py only gets 40% confidence. Confirming findings like these is on the manual checklist below.

Step 5: security patterns

The security page shows 0 live secrets and 9 code patterns in source, 4 of them high: three eval matches in src/flask/cli.py and one exec in src/flask/config.py. The same page also says no full security scan exists for this snapshot, so the dependency advisory count is not a result you can rely on here. Run pip-audit yourself (step 6).

Then read the four lines, which is the part people skip and the part that matters. I understand the temptation, because "4 high" already looks like a ticket someone else can pick up.

  • cli.py line 129 calls ast.parse(..., mode="eval"). That parses the --app string into a syntax tree to check it is a name or a function call, and nothing is executed, so the tool matched the word and nothing more.
  • cli.py lines 150 and 152 use ast.literal_eval, which only accepts literals like strings and numbers.
  • cli.py line 1023 runs your PYTHONSTARTUP file inside flask shell, the same way the plain Python prompt does, so it runs a file you chose, on your machine.
  • config.py line 209 is Config.from_pyfile, which loads a Python config file, and executing that file is the documented feature.

All four are deliberate, so in the report they become one line, "pattern matches reviewed, no issue", instead of "4 high severity findings" and a lost week.

The manual checklist

  1. Review every high finding by reading the code, as in step 5. It takes fifteen minutes for Flask and longer for a big target. Write one line per finding: real, deliberate, or false match.
  2. Confirm or dismiss dead code. Mark config files, scripts and plugin entry points as live, and anything left is either cleanup or a sign of abandoned features worth asking about.
  3. Check the licence and every runtime dependency. Flask is BSD-3-Clause and declares 6 runtime dependencies: blinker, click, itsdangerous, jinja2, markupsafe and werkzeug. For a commercial target, list every dependency licence and look for copyleft in shipped code. A small dependency list is a good sign on its own.
  4. Run a vulnerability check on dependencies. pip-audit (Apache 2.0) checks a Python environment against the Python Packaging Advisory Database and OSV. Run it against the lock file the product actually ships, since the loose ranges in pyproject.toml can resolve to different versions.
  5. Look at tests and CI. Flask's test workflow runs a 12-entry matrix: Python 3.10 to 3.15, a free-threaded 3.14 build, PyPy, Windows, macOS, and two runs against minimum and development versions of its dependencies. Its GitHub Actions are pinned to full commit SHAs and start with no permissions, and a separate workflow runs zizmor, a security linter for Actions files. Coverage shows as unknown on the repo page because no coverage report was uploaded, so ask for one.
  6. Check how a release happens. Is there a publish workflow, or does someone run a script from a laptop? Flask has a publish.yaml workflow. In a private target this is often the most revealing question.
  7. Map knowledge to people. Who last owned the riskiest files, and are they still around? The hotspot table shows an owner per file, and in an acquisition that list is the retention conversation.
  8. Read the architecture against the plan. If the buyer wants to add multi-tenancy or a new region, find the files that would have to change and check their scores. The dependency map shows how far a change spreads.
  9. Write questions for the owners. Every unexplained finding becomes a question, and the reading guide has a template for this memo.

If the target is a Flask app, add these

Auditing a product built on Flask is a different job from auditing Flask itself. Add these framework checks, each a grep and a few minutes of reading:

  • DEBUG or app.run(debug=True) reachable in production config. The debugger allows code execution from the browser.
  • SECRET_KEY hard-coded in source, or committed in a config file. It signs session cookies.
  • CSRF protection on forms, usually through Flask-WTF.
  • render_template_string called with user input, which is a template injection path.
  • Raw SQL built with string formatting instead of bound parameters.
  • send_file or send_from_directory with a path built from the request.

Bandit (Apache 2.0) walks the syntax tree of Python code and catches several of these, and Semgrep Community Edition (LGPL-2.1) lets you write a pattern for the rest. Both produce the same kind of output as step 5: matches that a person still has to read.

What the automated pass cannot tell you

The automated pass cannot tell you whether a low-scoring file matters to the business, or see code outside the repository: infrastructure, a second service, a vendored fork. Health and hotspots lean on git history, so a target that squashed its history last month will look calmer than it is, and none of it says anything about the team.

Its value is the order it gives you. Twenty minutes in, you know which files to read, which findings to explain, and what to ask on the first call. Due diligence targets are usually private repositories, and repowise runs the same first pass on a private repo once you connect it. The legacy codebase tools roundup covers the other options fairly if you would rather assemble it from separate tools.

The manual pass still has no done state, but at least now it has a reading order.

How we measured

Every Flask number comes from the public repowise pages for pallets/flask, read on 6 October 2026, describing the snapshot of 13 August 2026 at commit 2a8a38b0. I read the flagged lines, licence, dependencies and CI workflows in Flask's source at that commit. The index time comes from the snapshot's timestamps, while the 20 minutes is a time budget that nobody measured. The scores set a reading order and say nothing about how good Flask is. No full security scan ran on this snapshot, so no dependency advisory result is quoted.

FAQ

What is a code audit in technical due diligence?

It is the part of due diligence that looks at the source code itself: structure, how risky it is to change, tests, dependencies and security issues. It sits next to reviews of the team, process, infrastructure and licences. The output is usually a short report with ranked risks and questions for the owners.

How long does a technical due diligence code audit take?

The automated first pass takes about 20 minutes for a small repo once indexing has finished. The manual checklist takes from a day for something Flask-sized to one or two weeks for a large product with several services, which matches the timelines most agencies quote. Most of that time goes into reading findings and talking to the team.

What is a good legacy code audit tool?

Use one tool per question. A code health or hotspot tool ranks files by risk, a dead code tool finds unreachable code, Bandit or Semgrep finds risky patterns, and pip-audit checks dependencies. repowise puts the first three views and git history on one page, but you still need a dependency scanner and a person reading the results.

How do I audit code in a Flask application for security?

Start with the framework checks: debug mode, `SECRET_KEY` handling, CSRF on forms, `render_template_string` with user input, raw SQL and file paths from the request. Run Bandit for Python-level issues and pip-audit for dependency advisories. Then read every high finding, because many pattern matches are deliberate code, as the four in Flask itself show.

Can an automated tool replace a manual code audit?

No, because tools are good at ranking and counting and bad at intent. In this example, every high security match and both unreachable files were deliberate once a person read them. Use automation to decide what to read, and leave the conclusions to the person reading.