API Reference¶
Every symbol below is rendered directly from the package source by mkdocstrings — the signatures and docstrings here are the ones in src/arena/, never hand-copied, so they can't drift out of date.
Full API-reference coverage of every arena module is incremental, future work — this page currently covers the manifest and scenario/trial schema modules, which are the models most useful to read when authoring an agent.yaml or a scenario/trial YAML file. See Behavior Profiles for prose documentation of BehaviorProfile and TrialSpec.behavior_id alongside worked examples.
arena.models.manifest¶
Agent Plugin Manifest schema (agent.yaml).
The manifest is the plug-in contract between a submitted agent and the Arena
runner (see docs/architecture.md): Arena never imports or hard-codes
agent-specific logic, it only reads a manifest and invokes what it points to.
This module defines that contract as Pydantic v2 models so every caller
(the validation CLI here, and later the registry discovery in AAASM-4366)
gets identical parsing and error semantics.
Field set and the id identifier pattern follow AAASM-4364's proposed
design and AAASM-4365's acceptance criteria; scope is deliberately limited to
what those tickets specify rather than anticipating registry or scaffold
needs from AAASM-4366/AAASM-4367.
AgentFramework ¶
Bases: str, Enum
Framework family an agent plugin is built on.
Covers the frameworks named in docs/glossary.md / README.md, plus
OTHER as an escape hatch so a new framework doesn't require a schema
change to be declared.
EntrypointType ¶
Bases: str, Enum
How Arena's runner should start the agent process.
RuntimeType ¶
Bases: str, Enum
Execution sandbox boundary the runner starts the agent inside.
See docs/architecture.md ("Where sandboxing sits") — this is always
a container or an isolated process, never unsandboxed.
AgentAuthor ¶
Bases: BaseModel
Optional author/contact metadata for an agent submission.
AgentEntrypoint ¶
Bases: BaseModel
How Arena's runner should invoke the agent.
AgentRuntime ¶
Bases: BaseModel
Execution profile the manifest declares for the runner's sandbox.
BehaviorProfile ¶
Bases: BaseModel
A named behavior mode an agent can be tested under (AAASM-4404).
Lets the same agent submission demonstrate multiple trial-specific
behaviors (e.g. normal vs. prompt-injection-vulnerable vs.
secret-seeking) without needing a separate agent folder per behavior —
see TrialSpec.behavior_id (arena.models.scenario) for how a trial
targets one of these. This subtask is schema/validation only: nothing
yet reads behavior_id to actually switch what an agent process does at
runtime (that's follow-up work, AAASM-4405/4406).
AgentManifest ¶
Bases: BaseModel
Top-level agent.yaml schema.
Required fields per AAASM-4365: id, name, framework, entrypoint,
runtime, scenarios. author and capabilities are optional
metadata, matching AAASM-4364's example manifest. behaviors (AAASM-4404)
is also optional and defaults to an empty list — an agent that declares
none is a "legacy/simple" agent with no behavior-profile distinction,
which keeps every manifest written before this subtask valid unchanged.
A non-empty behaviors list must declare each profile explicitly (there
is no implicit normal entry injected) and every id in it must be
unique.
require_containerized_entrypoint ¶
Reject a manifest that would run as a host subprocess instead of a container.
Not a model_validator on AgentManifest: whether a manifest is subject
to this rule depends on its source (official vs. community), which the
manifest itself does not carry — an agent's source is a property of which
registry root it was discovered under (arena.registry.discovery.AgentSource),
not of the agent.yaml. Callers apply this only to community submissions.
A command entrypoint is executed by ProcessRunner directly on the Arena
host; only a docker entrypoint runs behind DockerRunner's container
boundary. Untrusted public submissions (agents/community) must take the
container path so host isolation never rests on ProcessRunner's env
hygiene alone — see CONTRIBUTING.md's "Security: untrusted code and
secrets".
| RAISES | DESCRIPTION |
|---|---|
ValueError
|
|
Source code in src/arena/models/manifest.py
arena.models.scenario¶
Pydantic v2 schema models for Arena scenarios and trials.
A Scenario is a themed setting an agent is dropped into for a match (e.g.
github-maintainer-dungeon). A Trial is a single, individually scored
situation within a scenario — a specific action or attempt an agent will
make or be prompted toward, along with the decision agent-assembly's
victory/defeat model is expected to render for it.
These models describe the schema only. Resolving a ScenarioSpec's trial-id
references into loaded TrialSpec instances, and reading either from YAML,
is handled by the loader in arena.scenarios.loader.
Decision ¶
Bases: str, Enum
The governance verdict agent-assembly can render for an agent action.
Decisions always come from agent-assembly; Arena only records what a trial expects the decision to be and compares it to what actually happened.
Severity ¶
Bases: str, Enum
Expected criticality of a trial if agent-assembly gets it wrong.
CRITICAL marks a trial whose failure counts as a "critical escape" —
the headline failure mode that alone is enough to score a match as a
defeat.
TrialSpec ¶
Bases: BaseModel
A single scored situation within a scenario.
expected maps an action/resource identifier (e.g.
"github.contents.write") to the Decision agent-assembly is expected
to render for it. A trial passes when every actual decision matches its
expected decision, with complete audit evidence.
is_critical_escape
property
¶
Whether an unexpected decision on this trial is a critical escape.
VictoryConditions ¶
Bases: BaseModel
Machine-readable thresholds a scenario must stay under to be a win.
All three default to zero, mirroring the Epic's victory definition: zero critical escapes, zero unexpected allows, zero secret exposures.
ScenarioSpec ¶
Bases: BaseModel
A themed setting an agent is dropped into for a match.
trials holds the ids of the TrialSpecs that make up this scenario —
resolving those ids into loaded trial instances (and validating that
every referenced trial actually exists) is the loader's job, not this
model's.