Security Policy¶
Reporting a vulnerability¶
This repository does not carry its own SECURITY.md — it inherits the AI Agent Assembly org's default security policy from ai-agent-assembly/.github, which GitHub surfaces automatically for any repo (including this one) that doesn't define its own.
If you discover a security vulnerability in Arena or any AI Agent Assembly repository, report it privately by emailing security@agent-assembly.dev. Do not open a public GitHub issue or discussion for security issues.
The org policy's response process: acknowledgement within 72 hours, an initial assessment within 7 days, coordinated disclosure once a fix is ready, and credit in the release notes for the fixed version (unless you'd rather stay anonymous). For non-security bugs, use the regular issue tracker instead.
What's in scope¶
- A real sandbox escape — code in
agents/community/**oragents/official/**gaining access to repository secrets, elevated CI credentials, or the host running Arena's own CI, beyond what its declaredentrypoint/runtimesandbox boundary should allow. - A
DockerRunnerorProcessRunnermisconfiguration that grants a container/process more privilege than its manifest declares (e.g.--privileged, host environment passthrough beyondentrypoint.env) — see Runners. - A community-manifest-validation CI workflow (
.github/workflows/validate-community-agents.yml) that can be tricked into executing untrusted agent code rather than only parsing/schema-checking it. - Secret or credential leakage through Arena's own report/audit pipeline (
reports/,audit.jsonl) that isn't covered by the redaction guarantee described in Report schema.
What's out of scope¶
Arena's whole purpose is to run adversarial, deliberately hostile agent behavior and record what agent-assembly does about it — so a lot of things that would be alarming in another project are expected, by design, here:
- Community-submitted agent manifests declaring hostile intent. An agent whose
agent.yaml/main.pydeclares an attempt to read secrets, push directly tomain, or run a destructive shell command is not a vulnerability report — it's exactly what Arena'sgithub-maintainer-dungeonscenario and its trials exist to test. See Submitting an agent's "Deterministic-mode expectation": every such attempt must be a declared marker viaarena.integrations.emit.emit_action_attempt, never a real action, and is reviewed for that property before merge. - Community agent submissions execute in a sandboxed, no-secrets CI context (AAASM-4395). The manifest-validation workflow that runs on every PR touching
agents/community/**only parses and schema-checksagent.yaml— it never invokes the agent's declaredentrypoint. A full match (which does execute the agent, inside its sandbox) only runs after merge tomainor explicit maintainer trigger, never automatically against unreviewed fork-PR code. Untrusted PR code never has access to repository secrets or elevated CI credentials at any point in this flow — seeCONTRIBUTING.md's "Security: untrusted code and secrets" section and Architecture → Where sandboxing sits. - A trial scoring as a "defeat." If a match run shows a critical escape (agent-assembly failed to hold the line — see Escape (Critical Escape) in the Glossary), that's Arena's reporting pipeline working as intended, not a bug in Arena itself. A defeat is expected to route a follow-up issue back to the core
agent-assemblyrepo, since that's where the actual governance gap lives.