Running Benchmarks (MCP Server)
The MCP benchmark servers provide a programmatic interface for running Juliet and
real-world benchmarks. All results are stored in data/benchmarks.db (SQLite,
WAL mode).
Benchmark Infrastructure
bench/
__init__.py Package marker
__main__.py CLI: python -m bench juliet [--full] [--jobs N]
[--keep-csv] [--compile-commands]
config.py Paths, constants, defaults
db.py SQLite schema, WAL mode, CRUD + query API
analyzer.py TP/FP classifier (Juliet ground truth)
runner.py Parallel CWE runner
machine.py Machine metadata (CPU, RAM, hostname)
SQLite Schema
Table |
Purpose |
|---|---|
|
One row per benchmark (version, SHA, mode, status, machine) |
|
One row per CWE per run (file count, violations, duration) |
|
Every individual aurora-lint finding with TP/FP classification |
|
Pre-computed aggregates per CWE (TP/FP rates) |
|
Per-rule per-CWE counts |
|
Real-world benchmark runs (aurora-lint version, machine) |
|
Per-project per-tool violation counts (+ codebase_commit) |
|
Every individual real-world aurora-lint finding (file, line, rule) |
|
Adjudicated TP/FP oracle keyed on (project, commit, file, line, rule) |
Historical data from JULIET_RESULTS.md and REALWORLD_RESULTS.md
(both retired 2026-09-03 once this backfill made them redundant with
Postgres) has been backfilled into the database.
Benchmark Workflow Protocol
Important
Commit BEFORE benchmark, do not bump the version: rebuild (
cargo build --release) and commit before starting. The run_id issqc-{version}-{sha}, and the SHA is what discriminates runs; the version string is a release artifact, bumped only when a release is cut (seeCLAUDE.md).NEVER modify code while a benchmark is running: The benchmark uses
target/release/aurora-lint. Rebuilding while running corrupts results.Wait for completion: Fast-mode ~8-10 min (4-core), ~3-5 min (24-core). Full-suite ~40-50 min. Check status no more than once every 5 minutes.
Compare runs after completion.
Sequence:
implement -> commit -> build release -> run benchmark -> wait -> analyze
Pre-Benchmark Checklist
All code changes committed
cargo build --releasesuccessfulpython -m bench corpus-checkclean (real-world)No other benchmark currently running (it’s your terminal – you’ll know)
Previous results compared if needed (
python -m bench compare)
Juliet Benchmark
python -m bench juliet [--full] [--jobs N] [--keep-csv] [--compile-commands]
python -m bench status [RUN_ID]
python -m bench compare BASE TARGET
python -m bench runs
python -m bench corpus-check [--json] # real-world checkouts still pinned?
Run identifiers accepted by status/compare:
"latest"– most recent run (default)Full run name:
"aurora-lint-0.3.20-abc1234"Commit SHA:
"abc1234"Historical runs:
"aurora-lint-0.3.17-historical"
Notes:
python -m bench julietblocks until the run finishes – background it yourself (nohup ... &, a second terminal,tmux) to keep working while it runsFast mode (default): per-CWE manifests, CWE-matched rules only. ~10x faster
Full mode: all 307 enabled rules against every CWE. Higher noise ratio
Resume: interrupted runs skip already-completed CWEs on re-run
Per-CWE/per-rule detail beyond what
status/compareprint is a directsqlite3 data/benchmarks.dbquery away (cwe_scans,violations,rule_cwe_breakdown) – there’s no separate CLI subcommand for it
Compile-Database Runs
--compile-commands makes a run pass aurora-lint’s --compile-commands flag,
adding the build’s include search paths and
-D macro state to the cross-file context. It is off by default – a
plain run is unchanged.
The run_id is suffixed -cdb (and, for real-world runs, so is the results
directory), so a with/without pair on the same aurora-lint build stays two distinct,
comparable runs. Without that suffix the second run would collide: Juliet’s
resume logic skips a run_id already marked completed, and the real-world
runner reuses the id for its results directory.
Databases are generated per-host (they embed absolute paths, so they are never committed):
# real-world: one compile_commands.json per checkout root
ansible-playbook playbooks/setup-compile-commands.yml -i "localhost," -c local --ask-become-pass
# Juliet: synthesized, no real build system to capture
python3 scripts/generate_juliet_compile_commands.py
A run requested with --compile-commands errors if the database is
absent, rather than silently running without it – a quietly-degraded run is
indistinguishable from a genuine “the compile DB made no difference” result.
Warning
A compile-DB run is a changed-rule delta, not a like-for-like
comparison. The flag can only add macro/header knowledge, so findings
move to (file, line) pairs that were never adjudicated and fall outside
the ground_truth precision/recall denominator in either direction.
Follow the delta-adjudication protocol in CLAUDE.md before publishing
any precision claim from such a run.
Note
Juliet gains nothing from this. Measured on the synthesized database
(54,486 entries): its only flag is -I<testcasesupport>, which the runner
already passes as -d testcasesupport. A with/without pair on
CWE457/s01 produced an identical 6,783 violations. The plumbing exists for
symmetry and for future Juliet build changes; the real payoff is on the
real-world corpora, whose databases carry genuine per-project include trees
and -D state.
Real-World Benchmark
Local and sequential, by design: one person running one benchmark in their
own terminal, against their own SQLite DB. See bench/realworld_runner.py.
python -m bench realworld-run [--tool sqc,cppcheck,clang-tidy] [--codebase C,C] [--compile-commands]
python -m bench realworld [RUN] [--compare BASE] # FP dashboard
python -m bench realworld-runs # list runs
python -m bench realworld-score [RUN] # measured precision/recall
realworld-run defaults to sqc – the aurora-lint binary; the tool id
is sqc on purpose, see Reproducing the Published Numbers – against
every codebase; narrow either flag as needed. It blocks until every requested combo finishes, then ingests
the aurora-lint results and scores them against the oracle – no separate ingest
step, no polling.
Supported tools: sqc (aurora-lint), cppcheck, clang-tidy, infer,
frama-c
Supported codebases: libcrc, sqlite, mosquitto, curl, hostap,
lua, raylib, pureftpd, sel4, mbedtls, valkey, ventoy
(aurora-lint-only for pureftpd, sel4, mbedtls and valkey — no
cppcheck/clang-tidy baseline yet; ventoy is the Win32 oracle,
Ventoy2Disk/Ventoy2Disk/ only, and <windows.h> is unresolved on a
Linux node for every tool alike)
Note
Remote-host execution (SSH) and background/concurrent run tracking
existed in this module’s MCP-server predecessor and were deliberately
dropped when it became a plain synchronous script – neither applies to
one person running one benchmark locally. If you need to run against a
fleet of remote hosts, that’s the kind of thing the maintainer’s
benchmarking_db infrastructure is for, not this repo.
Per-Codebase Rule Configs
Each codebase carries its own aurora-lint rules manifest in conf/realworld/ (the
real-world analog of a project shipping its own aurora-lint-rules.toml). The
runner reuses it for every run of that codebase via the
CODEBASES[<name>]["sqc"]["manifest"] registry entry, so rules that do not
apply are ignored consistently. There is no shared fallback base: a codebase
with no manifest entry is an error, because a manifest replaces the base
outright and so decides which rules can reach the oracle at all. The config is the
categorical filter (disable a whole rule only when it is inapplicable);
per-finding false positives among enabled rules are recorded in the
ground_truth oracle instead, so analyzer misfires stay measured rather than
hidden. See conf/realworld/README.md for the per-codebase audit workflow.
Because a manifest is standalone and a scan iterates only its enabled rules, a
rule with no entry at all never runs on that codebase and nothing reports the
gap — an omission is indistinguishable from an oversight, which is a defect that
has actually shipped. scripts/check_realworld_manifests.py (pre-commit hook
check-realworld-manifests) asserts every manifest decides every rule, and
that every enabled = false carries a comment naming its reason.
libcrc is fully audited (every enabled-rule finding labeled); the four
large codebases grow their labels incrementally.
Per-Codebase Scan Scope
Each codebase’s CODEBASES[<name>]["sqc"]["extra_args"] entry in
bench/realworld_runner.py also carries --exclude globs that
scope the scan to the shipped product, not the whole checked-out repo —
test harnesses, build tooling, vendored/bundled code, and companion tools
(fuzzers, example plugins, separate CLI utilities) are excluded so they don’t
inflate the violation count or dilute the precision/recall denominator.
These globs are derived from each codebase’s ground-truth oracle scope,
documented per project in docs/design/realworld-corpus-scope.md — exactly
which directories were ruled in/out during that codebase’s adjudication
sweep and why. (That rationale used to live in
data/precision_audit/<codebase>/README.md; data/precision_audit/ is
now local working data, gitignored per
docs/adr/0007-responsible-disclosure-gates-publication.md, and holds
only what your own adjudication pass produces.)
Important
-d/--directories in a codebase’s extra_args does not
restrict the scan — it only adds cross-file pre-scan context (see
Advanced CLI Usage). A codebase’s primary scan root is the whole repo
whenever scan_path is None, regardless of any -d entries in
extra_args. To actually narrow scope, use --exclude globs (or set
scan_path to a single subdirectory, as raylib does for
{path}/src).
When adding a new real-world codebase or revisiting an existing one’s
ground-truth audit, check whether its scope notes call for new
--exclude entries here — a mismatch between the oracle’s labeled scope
and the live scan’s actual scope means dashboard numbers include findings
that were never meant to be measured (or, more subtly, that the ground-truth
denominator no longer matches what’s being scanned).
Warning
That mismatch is not hypothetical, and it is measured. Scope is
declared in three places per codebase and nothing keeps them in sync:
the --exclude globs here (what aurora-lint reads), scope_include /
scope_exclude in data/benchmark_repos.json (what the oracle may
adjudicate), and the codebase’s Scope section in
docs/design/realworld-corpus-scope.md (the rationale the other
two claim to derive from).
Audited across all nine codebases, six agree and three do not — always in the same direction, with the scan wider than the scope, so aurora-lint emits findings that can never be labeled: sqlite 992, mosquitto 168, curl 144. That is 1,304 findings, roughly a fifth of that run’s whole unlabeled pool, unadjudicable by construction. They depress label coverage permanently, with work nobody is allowed to do.
The sharpest case is one category of file treated two ways in the same
suite: curl excludes include/** here, so its installed public
headers are never scanned, while mosquitto does not, so its public
headers are scanned and then declared out of scope by the oracle.
So when you touch either list, change both — and prefer making this one
derive from benchmark_repos.json, which is already the declared
single source of truth for the pins and is already read by
setup-benchmark-repos.yml and corpus-check. benchmarking_db
asserts both directions of drift on every run
(_check_scope_within_scan and _check_scan_within_scope); the
per-codebase reasoning and the audit results live in that repo’s
docs/corpus-scope.md, since it owns the scope predicate.
Note
benchmark_repos.json’s globs are path-aware: * stops at
/ and ** crosses it, so src/** and src/*.c are different
things. The --exclude globs here are aurora-lint’s own and follow aurora-lint’s
rules; do not assume the two spellings are interchangeable when copying
a pattern between the files.
Auto-Scoring
When realworld-run finishes, it ingests the aurora-lint results and auto-scores
them against the oracle: it writes a <run-dir>.score.json sidecar and
prints a one-line measured precision/recall. Scoring only joins findings to
existing labels — it never adjudicates new findings. Re-run any time with
python -m bench realworld-score <RUN>.
The scans and the ingest fail independently. Every tool writes its JSON export
as it goes, so a scan that printed ok is on disk under
results/realworld/<version-sha>/ whatever happens next; if the ingest then
fails, realworld-run says so on an INGEST FAILED line and exits
nonzero, and the run is simply absent (or partial) in bench realworld and
bench realworld-score until the ingest is repeated. Two runs can still be
compared straight from their JSON exports, finding for finding, without the
database.
Typical real-world workflow:
python -m bench realworld-run --tool sqc # blocks until every codebase is done
python -m bench realworld latest # view results
python -m bench realworld latest --compare 0.2.6 # compare against a prior run
Real-World Ground-Truth Oracle (measured precision/recall)
Volume deltas and CWE-aware Juliet rates do not predict real-world precision
(the v0.4.22 audit measured ~2–34% precision for the noisiest rules). The
ground_truth table is a growing, manually/AI-adjudicated TP/FP oracle for
the real-world codebases — the real-world analog of Juliet’s
OMITGOOD/OMITBAD. Because each benchmark checkout is pinned to a fixed git
SHA, a label keyed on (project, codebase_commit, file_path, line, rule_id)
stays valid across aurora-lint versions: only the tool changes, never the code. Labels
are appended over time, never tied to a single run.
CLI:
python -m bench corpus-check # checkouts still pinned?
python -m bench ground-truth # label inventory
python -m bench realworld-score [RUN] # measured precision/recall
python -m bench realworld-unlabeled [RUN] --rule R --project P --limit N --seed S
python -m bench realworld-import-labels CSV --run RUN [--source TAG] [--update]
realworld-score joins a run’s findings to labels for each project’s own
``codebase_commit`` and reports, per rule and overall:
precision = labeled-TP / (labeled-TP + labeled-FP), over the labeled subset of the run’s findings (a sampled estimate; “Label coverage” shows how much of the run is labeled);
recall = known-TPs flagged / known-TPs — a known true bug that stops being flagged drops recall, seeding regression detection;
unlabeled_count / unlabeled_fraction (overall and per rule) —
run_findings - labeled_total, i.e. how much of this run’s findings never got adjudicated. Precision/recall are only computed over the labeled slice, so a rule with a high unlabeled fraction can have a precision number that looks stable while its raw finding count swings heavily underneath it. The CLI text view flags any rule above 50% unlabeled;compare_runssurfaces the same fields (target_labeled_total/target_unlabeled_count/target_unlabeled_fraction) per rule delta so a raw-count regression that outpaces adjudication is visible without manually cross-referencingground_truth.
A run whose codebase_commit has no labels is warned about, not scored.
Incremental adjudication loop (need not be one-shot):
realworld-unlabeled RUN --rule X --seed S --limit N— pull findings with no label yet (reproducible sample);adjudicate them (Claude or manual) into a CSV (
rule,idx,project,file,line,verdict,reason);realworld-import-labels CSV --run RUN— append (existing labels are skipped unless--updatere-adjudicates them).
The first 200 labels were seeded from an early adjudication pass
(adjudication_0.4.22.csv).
Delta-Adjudication Gate
Important
Before citing a precision/recall claim (“precision held”, “FP reduced”, a published table row) for a rule whose detection logic just changed, run a delta-adjudication pass on that rule’s new findings first.
ground_truth labels are snapshotted at (project, commit, file, line,
rule). When a rule’s logic changes (any commit touching
src/rules/cert_c/**/*.rs that alters what it flags, not a pure refactor),
its new findings land on (file, line) pairs that were never adjudicated —
they’re silently excluded from the precision/recall denominator regardless
of direction. A flat precision number computed only over the pre-existing
labeled sample, or a raw finding-count comparison via compare_runs, can
both look clean while the real picture underneath is unmeasured. This is not
hypothetical: a 21-rule sweep in this project once nearly got reported as a
clean net-positive on aggregate raw-count deltas alone before someone
actually adjudicated the new findings.
Procedure:
Pull the rule’s new unlabeled findings (repeat per project, or split after):
python -m bench realworld-unlabeled RUN --rule RULE_ID --project P --json
Derive each project’s in-scope file predicate from its section of
docs/design/realworld-corpus-scope.mdbefore batching, not after. One delta-adjudication pass found 2,548 of 4,026 (63%) raw unlabeled findings were out-of-scope noise (test harnesses, vendored deps, language bindings) — mosquitto alone was 73% contamination. Scoping after batches are already generated means redoing completed adjudication work.Batch (~110-150 findings/batch), adjudicate, and import with
realworld-import-labels— the same workflow as building a fresh oracle.Only after
ground_truthreflects the new lines is a precision/recall claim about the changed rule safe to publish.
A worked example from this pattern: 6 projects, 14 batches, 1,478 findings, 0.7% delta precision — a very different number than the aggregate raw-count comparison suggested.
Comparing Across Runs
Juliet
python -m bench compare aurora-lint-0.3.17-historical latest
Positive FP delta = regression. Negative = improvement.
Real-World
python -m bench realworld 0.2.7 --compare 0.2.6
Competitor Benchmarks (Infer / Frama-C)
The bench/competitors.py module runs Facebook Infer and Frama-C EVA on
Juliet test cases and classifies findings as TP/FP using the same ground truth
as the aurora-lint benchmark (OMITBAD/OMITGOOD guards and procedure names).
Results are written to data/competitor_results/<tool>_<timestamp>.json.
Infrastructure
bench/
competitors.py Infer + Frama-C runners, TP/FP classification, comparison
Default CWE sets:
Tool |
CWEs |
|---|---|
Infer |
476, 690, 416, 401, 415, 761, 762, 121, 122, 124, 127 |
Frama-C |
190, 191, 476, 369, 197, 680 |
Running
# Run Infer on default CWEs (~80 min on 24-core)
python3 -m bench.competitors infer --jobs 8
# Run Frama-C on default CWEs (~7-9 hours)
eval $(opam env) && python3 -m bench.competitors framac --jobs 8
# Run a specific subset
python3 -m bench.competitors infer --cwes CWE476,CWE690
# Compare results
python3 -m bench.competitors compare \
data/competitor_results/infer_*.json \
data/competitor_results/framac_*.json
Timing Estimates
Tool |
CWEs |
Files |
Estimated Time |
|---|---|---|---|
Infer |
11 |
17,232 |
~80 min |
Frama-C |
6 |
11,628 |
~7–9 hours |
Infer uses incremental capture (infer capture --continue) per file then a
single infer analyze pass per CWE. Frama-C runs EVA per-function per-file
(-main <func>), which is the main bottleneck.
Classification Logic
Infer: Findings include a procedure field (e.g.
CWE476_..._01_bad). If the procedure contains _bad or Bad it is
classified as TP; if it contains good it is FP. Unresolved findings fall
back to line-level classification using parse_c_file_sections().
Frama-C: Each file is analyzed once per entry point (_bad function and
_good/goodN functions). Alarms found when the entry point is a bad
function are TP; alarms under a good entry point are FP.
Key Frama-C flags:
-machdep gcc_x86_64— enables GCC extensions (required for Juliet headers)-lib-entry— incomplete application analysis (nomain)-warn-signed-overflow -warn-signed-downcast— needed for CWE-190/191-eva-precision 1— reasonable precision/speed tradeoff
Troubleshooting
Issue |
Solution |
|---|---|
“Benchmark already running” |
|
Old results consuming disk |
|
Results show wrong version |
Ensure commit before build; the SHA is the id |
SQLite locked |
|
Historical run not found |
Data predates SQLite migration; not available |
Resolved Issues
DCL02-C Stack Overflow (Fixed 2026-01-07): Unbounded recursive AST traversal in DCL02-C caused stack overflow on large files (SQLite). Converted to iterative with depth limit.
STR31-C ``detect_manual_string_loop`` Runaway (Fixed 2026-02-25): Caused 36–49% of all violations on 3 of 5 real-world projects. Root cause: the final fallback iterated every line in the source file looking for
memcpy+strlen/string, so one match anywhere caused every loop to generate a violation –jimsh0.calone produced 180,297 violations. Fix: deleted the file-wide fallback, restricted matching to the loop condition and body, improvedis_string_memcpy. After the fix,jimsh0.c’s STR31-C count dropped from 180,297 to 10. (Migrated here 2026-09-03 fromREALWORLD_RESULTS.md, retired that day.)Output Buffer Saturation: aurora-lint emits one status line per rule per file (~100 rules × N files). Always suppress or redirect output during scans:
./target/release/aurora-lint directory/ --export results.csv 2>/dev/null