# Caio Theodoro — Senior AI Engineer

> I build production ML systems and the distributed infrastructure to scale them. 7 years across AI, engineering, and systems design — from model to product, at scale.

## Current role

Senior AI Engineer at Adopt AI (Full-time), since Mar 2026 — California, United States · Remote. Production AI workflow infrastructure: agent execution, human review, evaluation, and orchestration.

## Selected work

- [ReconForge](https://caio.theodoro.dev/projects): Laptop-trained 1.7B LoRA for financial reconciliation exceptions — 0.913 severity-weighted recall against a frontier model's 0.872, 1.000 recall on high-severity cases, zero API cost.
- [Suture](https://caio.theodoro.dev/projects): 8B VLM that diffs an underwriting binder against the issued policy: 0.959 recall across 13 discrepancy classes where zero-shot GPT-5.6 vision scores 0.373.
- [LossBench](https://caio.theodoro.dev/projects): Scores money-touching agents on expected operational loss instead of accuracy — severity-weighted, calibrated, replayable, on a hash-chained decision ledger.

## Recent writing

- [Ornith-1.5 Wrote Its Own Training Data. Distribution Match Won](https://caio.theodoro.dev/blog/ornith-curriculum-audit-distribution-wins.md) — 2026-08-21: Ornith-1.5 proposes its own training curriculum. I audited it with contamination-controlled data on a verifiable niche, construction pay-app review, and found the harder tasks it invents only help as a supplement, not a replacement, for distribution-matched data.
- [Suture: Catching Underwriting Errors GPT-5.6 Missed](https://caio.theodoro.dev/blog/suture-8b-three-gates-that-lied.md) — 2026-08-18: An 8B vision-language adapter trained to diff underwriting binders against issued policies catches far more errors than GPT-5.6 Luna does zero-shot. The real story is the three measurement gates that gave false confidence before the model actually worked.
- [The Third Number](https://caio.theodoro.dev/blog/why-lossbench-significant.md) — 2026-08-12: Agents that touch money get evaluated on accuracy or price. Neither one catches the failure that actually costs the most: a single bad decision with an outsized loss. This is about the metric that does, and why nothing measured it before LossBench.
- [Can A 1.7B Model Beat a Frontier on Reconciliation Exceptions?](https://caio.theodoro.dev/blog/reconforge-1-7b-beats-deepseek-on-the-money-metric.md) — 2026-06-08: A Qwen3-1.7B model fine-tuned on a laptop catches more high-severity reconciliation exceptions than DeepSeek v4-flash, including every one in the test set. Covers the benchmark, the training run, and the approaches that failed along the way.
- [Copying ARC-AGI's Benchmark Method, Then Stress-Testing It](https://caio.theodoro.dev/blog/building-benchmarks-like-arc-measuring-whether-it-works.md) — 2026-05-11: Benchmarks decay once models start training on them. I rebuilt ARC-AGI-3's benchmark methodology as a pipeline and tested which parts of it actually hold up: difficulty scaling, the human calibration bar, sample size, and a contamination monitor you can validate yourself.

## Contact

- [GitHub](https://github.com/caiotheodoro)
- [LinkedIn](https://www.linkedin.com/in/caiotheodoro1/)
- [Hugging Face](https://huggingface.co/caiotheodoro)
- [Toptal](https://www.toptal.com/developers/resume/caio-theodoro#NJeGln)
- [Email](mailto:dev.caiotheodoro@gmail.com): dev.caiotheodoro@gmail.com

## Elsewhere on this site

- [About](https://caio.theodoro.dev/about.md): Background, how he works, and what he is currently focused on.
- [Contact](https://caio.theodoro.dev/contact.md): Email, profiles, what to reach out about and expected response time.
- [Privacy](https://caio.theodoro.dev/privacy.md): What this site collects (almost nothing) and which third parties see a request.
- [Work](https://caio.theodoro.dev/projects.md): Employment history, certifications and open-source / research projects.
- [Blog index](https://caio.theodoro.dev/blog.md): Every published post with date, tags and summary.
- [Journey](https://caio.theodoro.dev/journey.md): Year-by-year timeline of the path into ML engineering.

HTML version of this page: https://caio.theodoro.dev/. Machine-readable index: https://caio.theodoro.dev/llms.txt. Full URL list: https://caio.theodoro.dev/sitemap-index.xml.
