BLOG / ENGINEERING

Engineering

Technical deep dives into dependency graphs, semantic search, and architecture

engineering8 min read

35x or 16%? we published both, and they measure different things

The same tool on the same codebase gives 35.6x fewer tokens on one measurement and 15.9% on another. Both are correct, and they measure different things.

2026-08-06Read →
engineering7 min read

a tool that did nothing came out 43% cheaper than the bare agent

A tool that never called its own server measured 43% cheaper than a bare agent. Prompt caching explains it, and it shows why dollar-cost benchmarks mislead.

2026-08-06Read →
engineering12 min read

how to benchmark a codebase tool so the number survives a rerun

JetBrains reran two token-saving claims and both collapsed. Here is the eight-part methodology we used to build a benchmark that survives an independent rerun.

2026-08-06Read →
engineering10 min read

we benchmarked ourselves against four open source tools and came last

We ran a sealed retrieval benchmark against four open source tools, scored last at 0.228, published it, found the bug, fixed it, and reran the sealed half once.

2026-08-06Read →
Does Code Health Predict Bugs? 21 Repos, 9 Languages, ROC AUC 0.74
engineering11 min read

Does Code Health Predict Bugs? 21 Repos, 9 Languages, ROC AUC 0.74

A reproducible defect-prediction study: 21 repos, 9 languages, 2,770 labeled files, ROC AUC 0.74, and 2.3x more defects caught under a fixed review budget.

2026-06-26Read →
Is AI-written code buggier than human code? We blamed 112,000 commits to find out
engineering11 min read

Is AI-written code buggier than human code? We blamed 112,000 commits to find out

We git-blamed 112,382 commits across 28 repos to test if AI-agent code is buggier than human code. Controlled for size, it isn't, and its lines last longer.

2026-06-09Read →
Process metrics beat structural metrics for predicting defects
engineering12 min read

Process metrics beat structural metrics for predicting defects

Complexity and code smells get attention. Across 21 defect-risk markers and 21 repos, the strongest predictors were evolutionary, not structural, size controlled.

2026-06-03Read →
engineering15 min read

How a 21-Marker Code Health Scorer Actually Works

A code health scorer works because one number is only useful if it is built from signals that map to real maintenance cost. A file can look clean and still…

2026-05-20Read →
engineering13 min read

Building a Deterministic PR Review Bot With Zero LLM Calls

A deterministic PR review bot can do useful review work without calling an LLM once. That sounds odd until you break “review” into smaller jobs: parse the…

2026-05-20Read →
Why vector-only retrieval misses co-change clusters in monorepos
engineering9 min read

Why vector-only retrieval misses co-change clusters in monorepos

Co-change analysis git shows which files move together, who owns them, and what breaks next. See why vector-only retrieval misses refactor signals.

2026-05-19Read →
Our git co-change model failed on a monorepo refactor
engineering10 min read

Our git co-change model failed on a monorepo refactor

Why co-change analysis git overfit a monorepo refactor, and the guardrails we added after 500 commits of history pointed at the wrong coupling.

2026-05-17Read →
The dead-code detector we shipped with Pagerank and git history
engineering11 min read

The dead-code detector we shipped with Pagerank and git history

Code hotspot detection as ranking, not a linter: Pagerank plus git history sorted quiet files from risky ones, so you can ignore the right 3.

2026-05-08Read →