Projects
ReconForge
Laptop-trained 1.7B LoRA for financial reconciliation exceptions — 0.913 severity-weighted recall against a frontier model's 0.872, 1.000 recall on high-severity cases, zero API cost.
Suture
8B VLM that diffs an underwriting binder against the issued policy: 0.959 recall across 13 discrepancy classes where zero-shot GPT-5.6 vision scores 0.373.
LossBench
Scores money-touching agents on expected operational loss instead of accuracy — severity-weighted, calibrated, replayable, on a hash-chained decision ledger.
Habeas
Form I-9 compliance validation fine-tune that cites the violated clause in M-274 / 8 CFR 274a.2, trained on synthetic-only data.
Substrate
Five agent-systems research units, each a reference implementation paired with the benchmark harness that can falsify its own claim.
Specula
FDA food-label compliance review: label photo in, cited PASS/FLAG report out, grounded in 21 CFR 101 and openFDA enforcement data.
Plumb
AIA G702/G703 pay-application audit that catches math errors, retainage drift, and double-counted change orders, returning corrected figures per line.
Perfectman
Social simulation where AI personas notice unevenly, reply late, lurk, and form private alliances — presence over tick loops.