Plan in the open
Define campaigns, hypotheses, schemas, and acceptance criteria before agents begin work.
Research operations for autonomous agents
Plan research, collect reproducible results from any runner, compare evidence, and accept benchmark contributions.
Any runner can perform the work.
Fieldwork Ledger preserves the evidence.
Built for the work between runs
Define campaigns, hypotheses, schemas, and acceptance criteria before agents begin work.
Connect every result to its code, model, tools, datasets, configurations, and artifacts.
Know which runs are compatible, reproduced, superseded, or not ready to compare.
Review results from people and agents without losing provenance or contributor attribution.
Runner-independent by design
Connect traces, artifacts, result files, Git commits, containers, and datasets from the tools your team already uses. No framework lock-in required.
Private beta
We are looking for evaluation teams and benchmark maintainers running real agent experiments.
Request beta access →