Report schema¶
aasm-arena run writes every match's durable report artifacts under reports/matches/<match-id>/ (--reports-root, default reports/matches/), plus a small set of top-level "static index" files a website or docs site can fetch directly — no Arena installation, no backend service, just a plain HTTP GET against whatever hosts this directory (or a copy of it: a CI artifact, GitHub Pages, an object storage bucket).
This page summarizes the layout for a docs reader. The checked-in reports/README.md in the repository root is the canonical, most detailed reference — read it directly if you're consuming these files programmatically (AAASM-4397).
Layout¶
reports/
├── README.md # checked in — canonical schema reference
├── latest.json # generated — most recent match, full report inlined
├── latest.md # generated — Markdown rendering of the same match
├── leaderboard.json # generated — one summary row per known match
└── matches/
└── <match-id>/
├── arena-report.md # generated — human-readable match report
├── arena-report.json # generated — machine-readable match report
└── audit.jsonl # generated — the match's redacted audit trail
Everything under reports/ other than README.md is generated output, not hand-authored — but it is tracked in git: a scheduled GitHub Actions workflow (.github/workflows/scheduled-matches.yml, AAASM-4428) refreshes it on a cadence and opens a pull request to land it on main (AAASM-6186), so reports/ on main stays a real, recent match history rather than a one-time snapshot.
Stable, consumer-facing files¶
Each of these carries its own schema_version field — check it before assuming a field exists; it changes only as a deliberate, visible schema bump (arena.reports.models.SCHEMA_VERSION, arena.reports.index.LATEST_INDEX_SCHEMA_VERSION / LEADERBOARD_SCHEMA_VERSION).
reports/latest.json— the most recently generated match, in full (arena.reports.index.LatestReportIndex). The fullMatchReportis inlined so a static site can fetch this one URL and render a complete result without a second request.reports/latest.md— the same most-recent match, rendered as Markdown (arena.reports.markdown.render_markdown).reports/leaderboard.json— one summary row per match found undermatches/, most recent first (arena.reports.index.LeaderboardIndex). Rebuilt from whatevermatches/*/arena-report.jsonfiles exist on disk at refresh time — not a persistent match-history database. An empty"matches": []is a valid, expected state.reports/matches/<match-id>/arena-report.json— the fullMatchReportfor one match (arena.reports.models.MatchReport). Every match ever run keeps its own directory here;latest.json/leaderboard.jsonare rolling pointers on top of this history, never a collapse of it.reports/matches/<match-id>/arena-report.md— the same match, rendered as Markdown.
Internal / not guaranteed¶
reports/matches/<match-id>/audit.jsonl— the match's raw (already-redacted) audit trail, one JSON object per line (arena.integrations.audit.ArenaAuditEvent). Useful for debugging a specific match in depth, but not part of the versioned report schema above — treat its shape as an implementation detail, not a stable public contract.
How this gets (re)generated¶
aasm-arena run <scenario-id> writes a new matches/<match-id>/ directory (arena.reports.generate.generate_report) and then refreshes latest.json, latest.md, and leaderboard.json (arena.reports.index.refresh_static_index) from whatever's on disk under matches/ at that point — so the static index is always consistent with the full match history, not just the match that was just run.
flowchart TD
A[run_match → MatchResult] --> B[score_match → MatchScore]
B --> C[generate_report]
C --> D[build_report → MatchReport]
D --> E[matches/<match-id>/arena-report.json]
D --> F[matches/<match-id>/arena-report.md]
C --> G[matches/<match-id>/audit.jsonl]
E --> H[refresh_static_index]
H --> I[latest.json]
H --> J[latest.md]
H --> K[leaderboard.json]
Sample reports¶
Two deterministic sample match reports — a win and a loss — are checked into this site so you can see the report shape without running a match yourself:
- Sample: winning match —
agent-assembly wins, zero critical escapes, zero unexpected allows. - Sample: losing match — a match with at least one critical escape, showing what a defeat looks like in the same schema.
Both samples are generated by scripts/generate_report_samples.py against the real arena.reports.* models — they are not hand-authored fixtures.