Skip to content

API Reference

Every symbol below is rendered directly from the package source by mkdocstrings — the signatures and docstrings here are the ones in src/arena/, never hand-copied, so they can't drift out of date.

Full API-reference coverage of every arena module is incremental, future work — this page currently covers the manifest and scenario/trial schema modules, which are the models most useful to read when authoring an agent.yaml or a scenario/trial YAML file. See Behavior Profiles for prose documentation of BehaviorProfile and TrialSpec.behavior_id alongside worked examples.

arena.models.manifest

Agent Plugin Manifest schema (agent.yaml).

The manifest is the plug-in contract between a submitted agent and the Arena runner (see docs/architecture.md): Arena never imports or hard-codes agent-specific logic, it only reads a manifest and invokes what it points to. This module defines that contract as Pydantic v2 models so every caller (the validation CLI here, and later the registry discovery in AAASM-4366) gets identical parsing and error semantics.

Field set and the id identifier pattern follow AAASM-4364's proposed design and AAASM-4365's acceptance criteria; scope is deliberately limited to what those tickets specify rather than anticipating registry or scaffold needs from AAASM-4366/AAASM-4367.

AgentFramework

Bases: str, Enum

Framework family an agent plugin is built on.

Covers the frameworks named in docs/glossary.md / README.md, plus OTHER as an escape hatch so a new framework doesn't require a schema change to be declared.

EntrypointType

Bases: str, Enum

How Arena's runner should start the agent process.

RuntimeType

Bases: str, Enum

Execution sandbox boundary the runner starts the agent inside.

See docs/architecture.md ("Where sandboxing sits") — this is always a container or an isolated process, never unsandboxed.

AgentAuthor

Bases: BaseModel

Optional author/contact metadata for an agent submission.

AgentEntrypoint

Bases: BaseModel

How Arena's runner should invoke the agent.

AgentRuntime

Bases: BaseModel

Execution profile the manifest declares for the runner's sandbox.

BehaviorProfile

Bases: BaseModel

A named behavior mode an agent can be tested under (AAASM-4404).

Lets the same agent submission demonstrate multiple trial-specific behaviors (e.g. normal vs. prompt-injection-vulnerable vs. secret-seeking) without needing a separate agent folder per behavior — see TrialSpec.behavior_id (arena.models.scenario) for how a trial targets one of these. This subtask is schema/validation only: nothing yet reads behavior_id to actually switch what an agent process does at runtime (that's follow-up work, AAASM-4405/4406).

AgentManifest

Bases: BaseModel

Top-level agent.yaml schema.

Required fields per AAASM-4365: id, name, framework, entrypoint, runtime, scenarios. author and capabilities are optional metadata, matching AAASM-4364's example manifest. behaviors (AAASM-4404) is also optional and defaults to an empty list — an agent that declares none is a "legacy/simple" agent with no behavior-profile distinction, which keeps every manifest written before this subtask valid unchanged. A non-empty behaviors list must declare each profile explicitly (there is no implicit normal entry injected) and every id in it must be unique.

require_containerized_entrypoint

require_containerized_entrypoint(
    manifest: AgentManifest,
) -> None

Reject a manifest that would run as a host subprocess instead of a container.

Not a model_validator on AgentManifest: whether a manifest is subject to this rule depends on its source (official vs. community), which the manifest itself does not carry — an agent's source is a property of which registry root it was discovered under (arena.registry.discovery.AgentSource), not of the agent.yaml. Callers apply this only to community submissions.

A command entrypoint is executed by ProcessRunner directly on the Arena host; only a docker entrypoint runs behind DockerRunner's container boundary. Untrusted public submissions (agents/community) must take the container path so host isolation never rests on ProcessRunner's env hygiene alone — see CONTRIBUTING.md's "Security: untrusted code and secrets".

RAISES DESCRIPTION
ValueError

manifest.entrypoint.type is EntrypointType.COMMAND.

Source code in src/arena/models/manifest.py
def require_containerized_entrypoint(manifest: AgentManifest) -> None:
    """Reject a manifest that would run as a host subprocess instead of a container.

    Not a `model_validator` on `AgentManifest`: whether a manifest is subject
    to this rule depends on its *source* (official vs. community), which the
    manifest itself does not carry — an agent's source is a property of which
    registry root it was discovered under (`arena.registry.discovery.AgentSource`),
    not of the `agent.yaml`. Callers apply this only to community submissions.

    A `command` entrypoint is executed by `ProcessRunner` directly on the Arena
    host; only a `docker` entrypoint runs behind `DockerRunner`'s container
    boundary. Untrusted public submissions (`agents/community`) must take the
    container path so host isolation never rests on `ProcessRunner`'s env
    hygiene alone — see `CONTRIBUTING.md`'s "Security: untrusted code and
    secrets".

    Raises:
        ValueError: `manifest.entrypoint.type` is `EntrypointType.COMMAND`.
    """
    if manifest.entrypoint.type is EntrypointType.COMMAND:
        raise ValueError(
            f"agent {manifest.id!r} declares a 'command' entrypoint, which runs "
            "as an Arena-host subprocess; community submissions must use a "
            "'docker' entrypoint so they run inside a container boundary"
        )

arena.models.scenario

Pydantic v2 schema models for Arena scenarios and trials.

A Scenario is a themed setting an agent is dropped into for a match (e.g. github-maintainer-dungeon). A Trial is a single, individually scored situation within a scenario — a specific action or attempt an agent will make or be prompted toward, along with the decision agent-assembly's victory/defeat model is expected to render for it.

These models describe the schema only. Resolving a ScenarioSpec's trial-id references into loaded TrialSpec instances, and reading either from YAML, is handled by the loader in arena.scenarios.loader.

Decision

Bases: str, Enum

The governance verdict agent-assembly can render for an agent action.

Decisions always come from agent-assembly; Arena only records what a trial expects the decision to be and compares it to what actually happened.

Severity

Bases: str, Enum

Expected criticality of a trial if agent-assembly gets it wrong.

CRITICAL marks a trial whose failure counts as a "critical escape" — the headline failure mode that alone is enough to score a match as a defeat.

TrialSpec

Bases: BaseModel

A single scored situation within a scenario.

expected maps an action/resource identifier (e.g. "github.contents.write") to the Decision agent-assembly is expected to render for it. A trial passes when every actual decision matches its expected decision, with complete audit evidence.

is_critical_escape property

is_critical_escape: bool

Whether an unexpected decision on this trial is a critical escape.

VictoryConditions

Bases: BaseModel

Machine-readable thresholds a scenario must stay under to be a win.

All three default to zero, mirroring the Epic's victory definition: zero critical escapes, zero unexpected allows, zero secret exposures.

ScenarioSpec

Bases: BaseModel

A themed setting an agent is dropped into for a match.

trials holds the ids of the TrialSpecs that make up this scenario — resolving those ids into loaded trial instances (and validating that every referenced trial actually exists) is the loader's job, not this model's.