Research prototype · not investment advice

Same judge. Second domain.

CACE-Bench measures whether an agent reaches the right verdict, refuses to decide when the facts were unobtainable, and cites only sources that answered. The DeFi track applies the same unchanged judge to GO / NO-GO verdicts on lending markets, vaults and RWA tokens.

—
historical cases reconstructed on chain
—
drafts corrected by on-chain reconstruction
—
false negative of the rules found

Research prototype by Digital Economy Lab. Not investment advice. Not a rating of any protocol. Historical cases are public incidents.

The judge does not change

Case→your agent→Narrative→judge (unchanged)→7 metrics with Wilson 95% CI
Credit trackDeFi track
FLAG / CLEAR / ESCALATENO-GO / GO / INSUFFICIENT_DATA
sanctions hithard gate fails (exit depth, reflexive backing, manipulable oracle)
PEP matchsingle-borrower concentration
KYC verifiedoracle & admin controls verified
countrychain

The judge, the metrics and the confidence intervals are imported from the credit track without a single change. Only the case generator, the source registry and the ground-truth rules are new.

Reconstructed cases

Each row was read from chain state at the block before the event. Every fact carries a capsule: contract, function, block and the SHA-256 of the response.

Loading cases.json…

Where the chain corrected the draft

H10 · wrong borrower

Draft: one borrower holds ~100% of debt — true today.
Chain at T0: that address held ~3%; a different address held 89% and exited before the collapse.
What changed: the label survived, the reason did not.

Rule adopted: API data may supply candidate addresses, never amounts.

H09 · not a constant

Draft: xUSD priced at a hardcoded $1.
Chain at T0: an “xUSD/USD” feed reporting issuer-side value (1.248 → 1.262), still 1.266 when xUSD traded near $0.26.
What changed: the case carries a manual label; open question 1a.

H05 · not on the public address still a draft

Draft: single-borrower CRV concentration.
Chain at T0: the founder’s public address held 0.67% of LlamaLend CRV debt.
What changed: kept as a draft; open question 1b.

This case is not in the table above — it is listed under drafts.

The first silent decision the track caught was its own. An early reconstruction of H10 targeted a public “Elixir USDC” vault and returned OK: the reconstructor had counted the vault’s idle balance as collateral and reported 100% concentration. The vault was 100% idle at T0 — not the lending channel at all. The reconstructor now refuses idle-only vaults, the artefact was removed, and H10 was re-scoped. A confident verdict on the wrong object, with every number internally consistent, is exactly the failure this benchmark scores.

Where the rules were wrong

At the block before the attack, Term Finance’s delay module reported txCooldown = 608,400 s — seven days. The rules read that as an adequate timelock and return GO. A public proposal queued for six days removed it, and $8.5M was drained.

Proposed rule change, open in issue #5: a queued change that lowers a control below policy sets controls_safe = False. At least two more such cases are needed before it is adopted.

The point of this block: the benchmark can show where its own rules fail. A split on which the rules separated perfectly would prove nothing.

Run it in your browser

Nothing is sent anywhere. Python runs in your tab via Pyodide; the files below are fetched from this site.

What it loads
/demo/cace_bench.py      judge, metrics, confidence intervals (v0.3.0)
/defi/defi_track.py      DeFi case generator and ground-truth rules
/defi/defi_sources.json  source registry

Run it against your own agent

python examples/defi_adapter.py --cmd "your-agent"
python examples/defi_adapter.py --cmd "python examples/claude_code_agent.py"

Co-define the case set — comment on issue #5.

Prospective split

  1. Commit. A position is fixed at T0; the verdict is hashed and the hash published here.
  2. Wait. The horizon runs — nothing is revealed, nothing can be edited.
  3. Reveal. After the horizon the verdict is published and checked against the hash.

What this does not show

  1. Synthetic split: injected error rates and an unverified registry.
  2. Historical split: a handful of cases on chain and a few drafts; perfect separation would prove nothing.
  3. Model knowledge leakage: cases before the agent model’s training cutoff are reported only as a contamination measurement.
  4. Code-level exploits (Euler v1, Balancer v2) are outside the observable risk surface — listed as unscored controls.
  5. Not investment advice; not a rating of any protocol.

Citing this

Concept DOI 10.5281/zenodo.21394049.

The DeFi track has not been released as a Zenodo version yet. Cite the commit, not the DOI, for anything on this page.