Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Limitations and known bypasses

Everything on this page is a limit that exists today, in the shipped code, on the platform the MVP targets. It is written so that a security reviewer can read it instead of reverse-engineering the integration, and so that nothing here has to be discovered the hard way.

The evidence base is verification-reports/AAASM-5276-claude-code-mechanism-matrix.md — the measured mechanism matrix from the Claude Code lifecycle Spike — plus the adapter code that shipped in AAASM-5281. A claim that traces to neither is not on this page.

For the boundaries of the product’s public claims — the cross-repository audit that checked every documented guarantee against the implementation — see verification-reports/AAASM-5528-public-claim-inventory.md (AAASM-5528). It is the companion artifact to this page: this page states what the integration cannot do, that one records where the documentation used to say otherwise.

Capability status legend

StatusMeaning
SupportedShipped and exercised by tests.
ExperimentalShipped, but its evidence is incomplete or its shape may change.
PlannedNot built. The ticket that builds it is named.
UnsupportedDeliberately not offered, with a reason.
CapabilityStatusNote
Managed settings write / merge / restoreSupportedFour owned keys; every other key preserved.
Proxy CA materialisation + NODE_EXTRA_CA_CERTS injectionSupportedAAASM-5276 condition C1.
HTTPS interception and redaction on the model pathSupportedMeasured against the real binary; see verify for what raises the level.
Side-channel scoping (*.anthropic.com)SupportedCondition C5.
MCP loading control (enabledMcpjsonServers / disabledMcpjsonServers)SupportedOptional, defence-in-depth. Never required for protection.
Drift detection and repairSupportedDetected at status/verify time, not in real time.
Adjudicating protection probeSupportedShipped as the default probe (AAASM-5300); see verify.
strict blocking on a high-severity scanner findingPlannedAAASM-5277, AAASM-5281. Today strict redacts, like recommended.
PreToolUse hook registration (pre-execution mediation of shell/file tool calls)PlannedAAASM-5646, AAASM-5534. No adapter registers one today, on any platform — see Hooks are not registered below.
Endpoint managed-settings fileInstallable, opt-in and authorizedAAASM-5298. --install-managed-settings; verified by read-back.
Endpoint managed-settings enforcement keysStill unmeasuredDocumented as non-overridable; no real override attempt has been measured on any host. How it would be measured.
Byte-exact configuration restoreUnsupportedSemantics-exact by accepted constraint (C3).
ANTHROPIC_BASE_URL redirection as a protection mechanismUnsupportedMeasured delivering the raw secret.
Host-level bypass preventionUnsupportedExplicit non-goal.
Lifecycle for Codex / Copilot / WindsurfPlannedCarried by LegacyAdapterShim; apply is refused.
Windows / LinuxUnsupportedmacOS is the MVP platform.

Known bypasses: demonstrated versus inferred

This split is published deliberately. Presenting the two groups as one undifferentiated list would overstate what has actually been tested — a demonstrated bypass is a measurement, an inferred one is a documented belief, and a reader deciding how much to trust this integration needs to know which is which.

Demonstrated by the AAASM-5276 harness

Three, each asserted positively by a test:

  1. ANTHROPIC_BASE_URL pointed at any endpoint removes Agent Assembly from the path; the raw secret arrives. Shown with both the real claude 2.1.220 binary and an emulated client.
  2. Launching claude outside the managed path (no HTTPS_PROXY) is unprotected.
  3. Observe/AlertOnly forwards the secret unchanged — correct behaviour, and the reason observe-only must never render as protection.

Inferred, not demonstrated

Documented, not measured by the Spike:

--dangerously-skip-permissions · defaultMode: bypassPermissions · --bare · unsetting the proxy env in the shell · repointing CLAUDE_CONFIG_DIR · symlinking .claude · replacing the binary · calling the API directly with the user’s own key · switching provider (CLAUDE_CODE_USE_BEDROCK / CLAUDE_CODE_USE_VERTEX) · running a pre-managed-settings release · a hook exiting 1 instead of 2.

The Spike’s summary sentence counts these as ten; the enumeration above is its own list and contains eleven items, because two permission-bypass flags are enumerated separately. The list is the claim, not the count.

Neither list is asserted to be exhaustive. “No finding” is not “no bypass”.

Which of these the shipped integration can actually see

Detection is not prevention. Where a bypass is detectable, the shipped adapter names it, lowers the reported protection level, and puts it in status; where it is not, the plan states so explicitly rather than leaving you to infer it from silence (aa-devtool-claude-code/src/bypass.rs).

Environment-variable bypasses require the launching client to state its own environment (AAASM-5993). The lifecycle service is a long-lived daemon shared by every client on the host, so it has no shell of its own to read — it can only report on the environment a caller tells it about. aasm states it on every status/verify/repair/remove invocation: ANTHROPIC_BASE_URL, CLAUDE_CODE_API_BASE_URL, CLAUDE_CODE_USE_BEDROCK, CLAUDE_CODE_USE_VERTEX and NODE_TLS_REJECT_UNAUTHORIZED are read by name — presence only, no value ever leaves the process — and carried across the DI-API request to the service, which reports each as genuinely Set, Unset, or (from a caller that predates this, or a third-party DI-API client that states nothing) NotStated. The settings-document half of each of these checks (a matching key in a settings env block) is unaffected either way: that document is read directly, not caller-stated.

BypassDetected?Where it is looked for
permissionMode / permissions.defaultMode = bypassPermissionsYesThe managed settings document. Becomes Absent evidence: the rules are still written and still read back, but nothing can be concluded from them about what the tool will do.
ANTHROPIC_BASE_URL / CLAUDE_CODE_API_BASE_URLYes, in the settings env block. Yes, in the launch environment, when the caller states it — aasm does.A settings env block; the launch environment, as stated by the caller.
CLAUDE_CODE_USE_BEDROCK / _VERTEXYes, when the caller states its launch environment — aasm does.The launch environment, as stated by the caller.
NODE_TLS_REJECT_UNAUTHORIZEDYes, in the settings env block. Yes, in the launch environment, when the caller states it — aasm does.A settings env block; the launch environment, as stated by the caller.
--dangerously-skip-permissions, --allow-dangerously-skip-permissions, --bareYesThe launch arguments. Reported and passed through unchanged — Agent Assembly’s interception sits below Claude Code’s own permission enforcement, so stripping the flag would change your session without changing what is protected.
Launching claude outside aasm runNoNo proxy or CA is injected; there is nothing to observe.
Repointing CLAUDE_CONFIG_DIRNo
Symlinking .claudeNo
Editing the settings file directlyNo (as a bypass)Surfaces later as drift at the next status/verify, not as a bypass at launch.
Replacing the claude binaryNo
Calling the Anthropic API from another program with your own keyNoNot this tool, not this path.
A hook exiting 1 instead of 2NoHooks carry no sensitive-data claim here (see below).

A bypass is not a failure. An unprotected launch is reported as a bypass, not as an Agent Assembly error, because the remedy is different and blaming the system trains people to ignore real failures.


ANTHROPIC_BASE_URL is routing, not protection

Redirecting Claude Code’s model endpoint is unsuitable for protection and is deliberately not offered as a mechanism (AAASM-5276 condition C4).

It was measured, with both the real binary and an emulated client, delivering the synthetic secret to the provider with no Agent Assembly component anywhere in the path. Setting it in the shell additionally suppresses Claude Code’s server-managed settings fetch.

This is why the lifecycle contract keeps ModelPathInterception and ModelGatewayBaseUrl as separate capabilities. They look alike and they are opposites: the first is a protection capability, the second is routing that removes protection.

A governed launch always replaces an ambient HTTPS_PROXY/HTTP_PROXY

aasm run overwrites, rather than merges with, any HTTPS_PROXY/HTTP_PROXY already present in the operator’s shell — including one set by a corporate proxy configuration or a provider-switching tool (AAASM-5892). This is deliberate: an ambient proxy address is environment-supplied input, and trusting it would let anything that can set an env var redirect governed traffic away from Agent Assembly’s enforcement point.

Since AAASM-5897, a governed launch prints a warning (naming that an ambient proxy was detected and replaced, never its value) whenever this override is about to happen. If your environment’s proxy also performs authentication, overriding it can produce a downstream auth failure from the tool itself — use --no-proxy to keep your own proxy and launch unprotected instead.

--no-proxy itself is refused when the installed integration requires managed operation (a Strict-profile receipt, or an endpoint-managed install). Since AAASM-5907, that check honours a --scope project install for the exact project root it was installed into, not only the machine-wide --scope user default — a Project-scope install elsewhere on the host is never mistaken for this one.

What verify adjudicates, and when it still exits 6

aasm integrations verify claude-code passes on a correctly installed integration whose protected path was exercised and adjudicated (AAASM-5300).

Raising the level to Gateway Protected requires exercised evidence, and exercised means the traffic was produced and adjudicated. Adjudicating means knowing what the payload leaving the machine actually carries — which a client on the near side of the proxy cannot see for itself. So the shipped probe does not try to. It marks its own request with an opaque correlation identifier, and the proxy — the component that runs the credential scanner and constructs the bytes that would be forwarded — answers on that request’s own connection with what it decided, plus a re-inspection of the payload it resolved to forward. Redacted is reported only when the proxy says it scrubbed the body and that the scrubbed bytes carry no credential (aa-devtool-claude-code/src/adjudicating_probe.rs, aa-proxy/src/probe_adjudication.rs).

Two properties of that exchange are worth knowing:

  • The probe learns nothing but its own verdict. There is no verdict store and no query surface — a verdict exists only as the response to the request that produced it, and is accepted only when it echoes the identifier that run minted. The correlation identifier is 32 hex characters of OS entropy and is derived from nothing about the payload.
  • The probe’s traffic never reaches the provider. The proxy terminates a correlated request instead of relaying it, and the probe sends a credential-free preflight first — so a path with nothing adjudicating on it never receives the synthetic secret at all.

A probe that returned Redacted because nothing obviously failed would be a vacuous pass, which is precisely what the evidence model exists to prevent. That rule is unchanged, and verify still exits 6 (verification_failed) whenever it cannot measure:

ConditionWhy it cannot pass
The path was never exercisedNo trust material in the receipt, so there is no intercepted model path to drive.
The certificate authority is not trustedThe MitM handshake fails, so nothing inspected the traffic. AAASM-5276 condition C1.
Nothing adjudicates the pathThe peer answered, but not with an adjudication — no component reported what it did.
The core is stoppedNothing is accepting connections; there is no verdict to read.
The exchange times outBounded and reported, never assumed.
A verdict for a different requestA verdict the probe did not produce is not evidence about the probe.
alert_only is configuredThe finding is recorded and the payload forwarded unchanged — observing is not protecting.

Read exit 6 on an otherwise-clean install as “not measured”, not as “measured and failed” — and read status for which it is.

AASM_STATE_DIR can redirect the Claude Code launch-env store — a named, un-closed gap (Core ADR 036 D6, gap #3)

A governed aasm run launch strips ambient HTTP_PROXY/HTTPS_PROXY/ ALL_PROXY/NO_PROXY (and their lowercase forms) from the child it spawns, then — if --no-proxy was not passed — reinjects only a supervisor-owned trusted value: either a runtime-pinned endpoint, or the receipted HTTPS_PROXY/HTTP_PROXY value aasm integrations install claude-code wrote into the launch-env store (StepAction::ConfigureProxy). This removal now happens once, at the Command level, immediately before spawn (aa-cli/src/commands/run.rs’s spawn_and_wait) — closing a pre-existing defect where the removal was only ever applied to an intermediate map the spawned process never actually inherited from (AAASM-5923).

What this does not close: the launch-env store itself — launch_env::installed_environment, read by both ClaudeCodeAdapter::build_launch_command (aa-devtool-claude-code/src/lib.rs) and ClaudeCodeIntegration::build_launch_command (aa-devtool-claude-code/src/lifecycle.rs) — is rooted in the ambient AASM_STATE_DIR environment variable (or its $HOME/.aasm default), read fresh by ClaudeCodePaths::from_env() in-process, with no supervisor/callee boundary across which a resolved path could instead be carried. An attacker able to set AASM_STATE_DIR before aasm run starts can point it at a directory containing a file named all_proxy (or HTTPS_PROXY pointed at an attacker-controlled proxy) under <state>/claude-code/<scope>/launch-env/, and that file reaches the governed child’s real environment exactly as a legitimately-installed receipted value would — bypassing aa-proxy entirely, while the launch still reports as governed. This is the more severe sibling of the pre-existing mitm_hosts-widening gap (an AASM_STATE_DIR-redirected mitm-hosts.d/*.hosts file widens what aa-proxy decrypts; this gap redirects traffic around aa-proxy altogether) — both bottom out in the same ambient state-root read, and the same attacker precondition (AASM_STATE_DIR set before the launch begins) reaches the new TrustedUpstreamProxyEndpoint/DeclaredEnterpriseDestination trusted-config artifact too: aa-proxy’s validation (trusted_upstream::load_and_validate) refuses a malformed or wildcarded artifact regardless of who authored it, but it validates well-formedness, not authorship — an attacker who controls the pre-launch state root can author their own well-formed artifact at that root.

This is not fixed by design, not an oversight left for a later pass under this ticket: no supervisor/callee process boundary exists in ClaudeCodeIntegration’s construction to relocate the read across, so “the supervisor resolves the path and passes it down” would read the identical variable through the identical fallback chain one stack frame apart — it relocates the read, not the trust boundary. Genuinely closing it needs either an operator-facing --state-dir flag that takes precedence over AASM_STATE_DIR everywhere in the tree, or per-read receipt-fingerprint verification (the executor already computes one per ConfigureProxy step, so this is buildable) — both are materially larger changes than adding a spawn- boundary env-sanitization invariant, touching every existing AASM_STATE_DIR consumer, not just the two new artifacts Core ADR 036 introduces. Closing this is tracked as separate, not-yet-approved follow-up scope (Core ADR 036, “named un-closed gaps” #3, alongside gap #1’s identical mitm_hosts-widening channel).

Practical implication: treat AASM_STATE_DIR itself as trusted input on any host where aasm run governs a launch — anyone who can set it before the launch starts can redirect where the launch-env store, the mitm-hosts.d allowlist, and the trusted-upstream-proxy artifact are all read from, independent of anything the proxy’s own validation catches.

AAASM-5924 addendum: AASM_STATE_DIR is not the only ambient channel to the trusted-upstream-proxy artifact. aa-proxy’s ProxyConfig::from_env reads AA_PROXY_TRUSTED_CONFIG_PATH directly — a raw path, not derived from AASM_STATE_DIR at all — and ProxyGuard::build_command (aa-cli/src/commands/proxy/guard.rs) and aasm proxy start’s proxy_child_env (aa-cli/src/commands/proxy/start.rs) both set it only when a real artifact exists on disk but never env_remove it otherwise, so an ambient AA_PROXY_TRUSTED_CONFIG_PATH set before either boundary spawns reaches the child proxy unchanged when no legitimate artifact overrides it. This is a second, more direct route to the identical effect this section already discloses (an attacker who controls the pre-launch environment can point the spawned aa-proxy at their own well-formed artifact) — same precondition, same consequence, not a new attacker capability. Filed as a wording correction against Core ADR 036’s own Test 8 row, which currently (incorrectly) claims this channel is “not adopted”.

Trusted upstream proxy chaining only ever routes explicitly declared destinations (Core ADR 036 D-F)

Chaining routes a CONNECT to the trusted upstream proxy only when the authority exact-matches a DeclaredEnterpriseDestination host and port in the validated trusted-config artifact (D-A/D-D). There is no “send everything through the corporate proxy” mode: a non-declared destination always takes the unchanged direct-dial path, with the unchanged SSRF guard (connect_revalidated) and the unchanged egress-allowlist/denylist checks — chaining being configured changes nothing about how those destinations are handled. Full-egress routing through an operator’s corporate proxy is deliberately out of v1 scope, tracked as a separate, not-yet-approved Spike, not something this feature quietly does not finish.

The published aasm binary has no command that writes this artifact. aasm integrations install’s --trusted-upstream-proxy/ --enterprise-destination/--llm-endpoint flags live inside the strip-for-publish:begin/end devtool region of aa-cli/src/commands/mod.rs (AAASM-2340) and are removed entirely from the crates.io-published crate an operator installs via cargo install aasm. Any statement that an operator “can configure chaining” with the released CLI is false unless it names this: today, only a source checkout (cargo run -p aa-cli -- integrations install ...) can write the artifact; the released binary can only consume one placed on disk by some other means.

A chained forward carries no chained-specific evidence tier (Core ADR 036 D9/M2)

A request that traverses the trusted upstream proxy produces the identical ProtectionState a direct forward does — there is no distinct evidence tier recording that a request specifically went through the second hop. The locally-terminated adjudicating probe (see “What verify adjudicates” above) can justify GatewayProtected on chained traffic exactly as it does on direct traffic; neither its passing nor its failing says anything about whether the chained hop was actually used for that request. This is not blocked by anything chaining adds — a chained-specific evidence tier was never built for v1, and is not implied by anything this feature’s naming suggests.

Declaring a destination widens MITM eligibility on its host, not its declared port (Core ADR 036 D2b)

should_mitm’s chained-destination check (aa-proxy/src/proxy/mod.rs) matches a declared destination’s host only — the same host-keyed (not host+port) matching every other mitm_hosts-derived MITM-eligibility check in this proxy already uses. Chained routing itself is stricter (host and port, exact match, D-D) — so declaring corp.llm.internal:443 as an enterprise destination makes corp.llm.internal:8443 MITM-eligible (decrypted and DLP-scanned) even though it is not chain-eligible (it still direct-dials, unchanged). An operator who expects “MITM-eligible” and “chain-eligible” to track the same port is surprised by this: declaring one port on a host widens what gets decrypted on every port that host is reached on. Accepted as part of Core ADR 036’s own R3 fix, not a defect this Story introduces or is expected to close.

The managed-settings file can be installed; its enforcement is still unmeasured

/Library/Application Support/ClaudeCode/managed-settings.json is the endpoint managed-settings file. Its managed-only keys — allowManagedPermissionRulesOnly, disableBypassPermissionsMode, allowManagedMcpServersOnly, allowManagedHooksOnly — are the strongest available counters to the bypasses listed above.

Since AAASM-5298, Agent Assembly can install that file — through an opt-in, explicitly authorized path, never as part of a default install. See --install-managed-settings and Protection levels → Host Enforced.

What Agent Assembly verifies, by reading the file back after the write:

  • its bytes are exactly the bytes you were shown and authorized;
  • it parses as a managed-settings document and carries the managed-only keys;
  • it is owned by the expected principal (root at the canonical path);
  • no account other than its owner can rewrite it.

What Agent Assembly does not measure, and will not claim:

  • that Claude Code honours each managed-only key at runtime. Anthropic documents these keys as non-overridable; Agent Assembly has not measured a real override attempt on any host. AAASM-5276 condition C6 is closed for the install half and open for the enforcement half.

What would close it is written down rather than left as “we need a device”: Measuring managed-settings enforcement is the procedure, and scripts/measure-claude-code-managed-enforcement.sh refuses to run anywhere it could not produce real evidence. The measurement needs a real privileged write on a real host — which AAASM-5308 scopes as “a managed/MDM-enrolled macOS device, or one where the file can be provisioned with administrator consent” — and, for the override attempts, an account that is not an administrator. Until that has been run, none of it is claimed.

Read a Host Enforced level as: “the managed policy is installed at the OS-managed path, owned as expected and not writable by you.” Do not read it as “this bypass has been demonstrated to fail.” Every status that reports it carries that caveat in the evidence detail.

What the install will not do

  • It will not elevate anything but the single file placement. aasm never runs as root, and no other step in any plan asks for authorization.
  • It will not replace a managed-settings file Agent Assembly did not write — for example one deployed by your organisation’s device management. That is a refusal, and moving the file aside is your explicit decision to make, not Agent Assembly’s.
  • It will not run without a terminal. A non-interactive invocation fails immediately rather than blocking on a credential prompt nobody can answer.
  • It will not report success on the authorization mechanism’s word. A read-back that does not match rolls the write back and fails.

Restore is semantics-exact, not byte-exact

Accepted constraint C3 (ADR 0030 — Accepted risks; AAASM-5276 condition C3, accepted by AAASM-5278).

aa-devtool-claude-code/src/apply.rs reserialises the whole settings document on every write. A user file in non-canonical formatting — hand-chosen key order, unusual indentation, trailing layout — therefore cannot survive an install → remove cycle byte-for-byte, no matter how good the receipt is.

What removal does restore is the document’s meaning:

  • every value Agent Assembly displaced is put back;
  • every key Agent Assembly added is deleted;
  • every key you changed after installation is carried through untouched.

Two consequences follow deliberately from accepting this rather than working around it. Fingerprints are taken over canonical JSON, so a reformat is correctly reported as no drift. And a removal report states the limitation rather than implying a guarantee the write path cannot keep.

The alternative — preserving the original document verbatim — was rejected as disproportionate for the MVP: it needs a format-preserving JSON editor no in-tree adapter has, and it buys byte-identity in a file the tool itself rewrites. If an adapter’s write path ever stops reserialising, this becomes a choice rather than a constraint and should be revisited rather than inherited.

The scanner only recognises the shapes it knows

Detection is deterministic and pattern-based (aa-security’s CredentialScanner). “Detected” means matched by the pattern set; it does not mean understood.

A credential whose shape is not in the pattern set passes through unrecognised — a bespoke internal token, a secret with no distinguishing prefix, a value split across fields. There is no claim of complete detection, and the Spike explicitly does not license one.

Three knock-on limits worth stating:

  • An undetected secret is not absent from audit records. If the scanner never classified a value as a secret, it was never redacted, and it may appear in a recorded payload like any other content.

  • Redaction is not encryption and not a DLP product. An oversized field that cannot be scanned reliably is replaced wholesale with [REDACTED:OVERSIZED] — the scanner fails closed — but that is a containment behaviour, not detection.

  • A flagged undecodable payload loses its whole audit content. A bytes field that is not valid UTF-8 — a binary body, or multi-byte text cut by a chunk boundary — is still scanned, but a detected secret cannot be excised precisely, because the finding’s offsets index the lossy decoding rather than the payload. The field is therefore replaced in full with [REDACTED:UNDECODABLE]. The secret is contained, but so is everything else that was in the field: the surrounding content does not reach the audit record. A clean undecodable field is unaffected and is forwarded byte-identical (AAASM-5346).

    This is sharper for zh-TW traffic until AAASM-5344 ships. That defect makes ordinary Chinese text register as GenericHighEntropy findings, so a chunk-split Chinese payload is dirty by false positive and loses its entire args_json to the 22-byte marker — where previously it was forwarded corrupted but present. Containment is the correct trade, and a corrupted payload was never trustworthy audit content, but the loss is real and it is why ADR 0032’s operational guidance treats zh-TW traffic as unsafe until AAASM-5344 lands in v0.0.1-rc.7. Once it does, benign Chinese text stops producing findings and this path stops being reached by ordinary traffic.

Hooks are not registered: PreToolUse is unmediated

No adapter registers a PreToolUse hook, on any platform (AAASM-5646). WRITABLE_KEYS in aa-devtool-claude-code/src/managed_settings.rs and the non-managed settings.json write path in aa-devtool-claude-code/src/apply.rs both carry a permissions/permissionMode surface but no hooks key, and no code path in this crate constructs one. Concretely: a shell command run through Claude Code’s Bash tool, and a file read/write/edit through its file tools, reach the tool before anything Agent Assembly wrote is consulted. This is true of the shipped adapter even when the managed-settings file, the proxy CA and the launch environment are all correctly installed — those mechanisms cover different surfaces (see the capability table above), not this one.

This is a decision, not an oversight discovered too late to fix: the mechanism exists and is deliberately not wired, for three reasons that make “wire it up” a bigger design question than a hook registration:

  1. The decision function PreToolUse would need to call — handle_policy_query (aa-runtime/src/pipeline/mod.rs) — is private to a running aa-runtime process and reached only over gRPC, with registry lineage, op_control state and audit-write side effects a hook process has none of. A hook invoked as a short-lived subprocess evaluating a policy file directly (the way aasm policy simulate does) is a different, weaker mechanism — no agent registry, no lineage-resolved cascade, no audit trail — and shipping it under the same name as “routes to handle_policy_query” would be exactly the kind of overclaim this page exists to prevent.
  2. allowManagedHooksOnly makes a locally-written hook inert under the profile that most needs it. managed_settings_document() sets allowManagedHooksOnly: strict for the strict profile — so once that key is present, Claude Code honours a hooks entry only if it also lives in the root-owned managed-settings document, not in ~/.claude/settings.json. Writing a hook there is a materially different, higher-stakes change: it goes through WRITABLE_KEYS, validate_managed_document and the ConsentDisclosure/read-back path that governs every other write to that file, and a root-owned document whose hook command names a user-writable binary path is its own review question.
  3. The exit-code contract has no test pinning it. This page’s inferred- bypass list already names “a hook exiting 1 instead of 2” — Claude Code only blocks on exit code 2; 1 and any other non-zero code do not. Registering a hook without a regression test asserting the exact contract (a real Bash call blocked, with the file it would have created absent, per exit code 2 and not 1) would convert a documented absence into an undocumented, untested bypass — a worse state than today’s.

AAASM-5646 tracks this decision explicitly; AAASM-5534 is the separate, broader question of whether host-wide PreToolUse mediation is feasible at all — this page’s answer does not wait on that study.

What this means for a policy author: a GovernanceAction::ProcessExec (aa-core/src/policy.rs) deny rule targeting a Capability::TerminalExec (aa-security/src/policy/capability.rs) compiles, validates, and is evaluated correctly by the policy engine — but nothing on the Claude Code integration path today generates that action before the shell command already ran. The same is true of a file-access deny rule against the Read/Write/Edit tools. Enforcement against these tool calls, where it exists at all today, is the sidecar proxy and eBPF layers described at the top of this repository’s .claude/CLAUDE.md — not this integration.

Hooks cannot carry a sensitive-data claim

Claude Code hooks govern tool and action execution. They cannot see or modify model-bound prompt content, so no hook can support a sensitive-data protection claim. None are registered today (see above); even if one were, it would never be a substitute for in-path interception.

NODE_TLS_REJECT_UNAUTHORIZED is never set by Agent Assembly. Setting it would make interception “work” by disabling certificate verification, and a TLS failure is a finding, not something to suppress. If you have it set, status reports it as a bypass.

Other tools are not yet on this lifecycle

Codex, GitHub Copilot and Windsurf Cascade are carried by LegacyAdapterShim (ADR 0030 §7). They can be discovered, planned and reported on, but their plan steps name no destination file, so the service refuses to apply rather than reporting a success that performed nothing. Their per-capability tiers in the capability matrix come from their adapters’ declarations, not from a measured Spike. Superseded per-tool detail for each — predating the consolidated matrix — is kept at Governance Limits by Tool (also covering Codex, Copilot and Windsurf).

This page is scoped to locally-running tools. If the tool in question is a SaaS-hosted coding agent (Claude.ai, ChatGPT, Cursor cloud), see SaaS Coding-Agent Governance Limits instead — those adapters are capped at L1Observe for a structural reason (no local process to intercept), not a maturity gap like the tools above.

Timing and freshness

  • Drift is found when status/verify runs, so a window exists between a change and its discovery. Between two verifications a state can be reported that has since become false. The evidence carries its timestamp — the claim is “verified at T”, not “true now” — but a consumer that ignores the timestamp will over-read it.
  • Protection state is re-derived on read, never cached. AAASM-5276 measured ~0.07 ms from core stop to connections being refused; a cached level would keep displaying protection that no longer exists.
  • Repair is deliberately narrow. It will not overwrite a key it does not own, even when that key is the cause of the drift — it reports and stops.

What stays local, and what is never recorded

These two are guarantees rather than limitations, but they belong beside the limitations because each has its own edge.

Raw content is processed locally. Scanning and redaction happen in the Agent Assembly runtime on your machine. Raw file contents and raw prompt text are not shipped to Agent Assembly infrastructure in order to be analysed.

That is not the same as your content stays on your machine. The point of the tool is to send prompts to a model provider; Agent Assembly’s job is to make what is sent safe, not to prevent sending. Where an org deployment is configured, policy documents, audit metadata and decision records may be forwarded to a control plane. Metadata is not raw content, but it is not nothing either.

Raw secret material is never written to logs, traces, audit events, installation receipts, API responses or diagnostic output. Findings are recorded as metadata — kind, position, count — and the redaction record deliberately stores no raw value (aa-security/src/redaction.rs). Diagnostics produced for support are subject to the same rule; troubleshooting is not an exemption.

This is enforced by the shape of the types, not by a redaction pass someone can forget to call. Across the DI-API, a rendered settings body becomes a content_sha256 plus the owned key names; an environment value becomes the variable’s name; a model base URL becomes the setting’s name, because a URL can carry a token in its query string. StepView — the sharpest edge — has no field a step value could land in. A bypass report likewise echoes variable names only and never their values, asserted by a test that plants a sentinel value and fails if it appears.

The edge: this does not govern the tool’s own records. Claude Code’s transcripts, your shell history and your provider’s server-side logs are outside Agent Assembly’s control entirely.


What is never claimed

Stated positively so it can be quoted:

  • No host-level bypass prevention. A user or process able to launch the tool outside the managed path is outside enforcement, at every level available.
  • No protection for unmanaged direct provider connections.
  • No complete secret detection.
  • No protection while the core is stopped. Protection is a running-system property; when the core is down the product says not protected, not protection unknown.
  • No universal interception of every AI development tool.
  • No claim that a settings file alone proves model-egress protection. A configuration is intent; a level is behaviour.
  • No claim that MCP is required for, or equivalent to, protection. It is one optional mechanism among several.

References


Last updated: 2026-09-02 by Chisanan232