Mods turn customization into an in-process control plane
Claude Code Mods, announced by Anthropic on October 1, are small TypeScript functions that can run before, after, instead of, or around Claude Code events. That makes them more powerful than a prompt fragment and more tightly coupled than an external integration. A mod can affect what the model sees, which tool invocation executes, whether a permission is granted, what result returns, and what the user interface displays.
The attractive story is obvious: a team can add policy, observability, interface changes, and workflow logic without forking the product. The operational story is harder. Several correct mods can become one incorrect system when they compete for the same event, assume a different order, swallow a continuation, duplicate a side effect, or target incompatible type definitions. A security mod can inspect the pre-transformation request while another mod changes it later. A telemetry wrapper can record success even though an outer wrapper replaces the result. A fallback can run the underlying tool twice.
Anthropic labels the API early access. The launch post and repository documentation also make the trust boundary explicit: mods run in process, are unsandboxed, and have the same machine access as Claude Code. A plugin may package a mod, but packaging does not restrict it. Installation is therefore equivalent to installing executable development tooling, not importing a harmless template.
Review a mod for what it does. Approve a mod stack for what the ordered composition can do.
The current developer signal is meaningful but should not be inflated. Anthropic published the feature and built-in examples; a launch discussion attracted more than a thousand Reddit votes and 182 comments; a community catalog began indexing public examples. One scanner snapshot reported 72 mods across 142 repositories and classified filesystem, process, network, and event-hook capabilities. Those are discovery and capability observations, not proof of compromise, production adoption, or safe behavior.
Use a mod only when the job belongs on the event path
Claude Code already has several extension surfaces. Choosing the narrowest one reduces coupling and review cost. A reusable instruction is usually a skill. A new remote capability may be an MCP server. A distributable bundle is a plugin. A shell hook is useful for coarse process boundaries. A mod is justified when code must synchronously observe or transform an internal event and ordinary configuration cannot express the requirement.
| Surface | Best for | Main boundary | Typical review |
| Skill or instruction | Reusable method, checklist, domain guidance | Influences model behavior; does not enforce | Prompt injection, ambiguity, stale guidance |
| Classic hook | Shell-level validation or notification | Separate process, serialized input/output | Command execution, environment, exit behavior |
| MCP server | Tools and resources behind a protocol | Separate server and explicit capabilities | Authentication, tool schema, data flow, egress |
| Plugin | Distribution bundle for commands, agents, skills, hooks, or mods | Package, not a security boundary | Every bundled executable surface plus updates |
| Mod | Low-latency observation or transformation of Claude Code events | In-process and unsandboxed | Source, dependencies, events, order, continuation, side effects |
This distinction prevents a common failure: implementing policy in a surface that can only suggest behavior. A skill can tell the model not to expose a secret; an event interceptor can reject a request; an external credential broker can make the secret unavailable. These controls are complementary, not interchangeable. For higher-consequence actions, keep final authorization in the target system and use mods for local enforcement or evidence, not as the only boundary.
Draw the nesting before installing the code
Interceptors compose like middleware. The first loaded wrapper sees the incoming event first, calls the next handler, and receives the returning result last. That creates two orders: an inbound order and a reverse outbound order. A flat list of installed names hides this behavior.
incoming tool.call
-> managed security mod (pre-check)
-> redaction mod (rewrite arguments)
-> telemetry mod (start span)
-> Claude Code / actual tool
<- telemetry mod (finish span)
<- redaction mod (rewrite result)
<- managed security mod (post-check)
outgoing result
Before installation, write an event ownership table. For each event, name who may observe, mutate, deny, replace, retry, emit a side effect, and present the final result. “All mods may wrap tool calls” is not a contract. A stronger rule might say security is the only deny authority, redaction is the only argument transformer, telemetry is observe-only, and no mod may retry a non-idempotent tool.
| Event | Owner | Allowed mutation | Continuation rule | Failure policy |
| prompt.submit | context policy | Add fixed policy context; never alter user text silently | Exactly once | Reject with visible reason |
| permission.request | managed security | Reduce authority only | At most once | Fail closed |
| tool.call | security + one transformer | Normalize approved fields | Exactly once unless blocked | Deny or return typed error |
| tool.result | redaction | Remove configured sensitive fields | No re-entry | Return original only under explicit policy |
| UI notification | workflow UI | Presentation only | Never changes execution state | Log and continue |
Make the stack versioned, ordered, and reviewable
A composition manifest should sit beside the project or managed configuration. It is not a Claude Code feature; it is an operational artifact that makes hidden assumptions reviewable. Pin the Claude Code range, generated plugin types, mod source revision, dependency lock, order, event ownership, side effects, failure behavior, and rollback target.
schema: claude-mod-stack/v1
claudeCode: "2.1.287"
pluginTypesSha256: "reviewed-generated-types-hash"
stack:
- id: managed-sec-default
source: anthropics/claude-code@reviewed-commit
order: 10
events: [permission.request, tool.call]
may: [observe, deny]
mustNot: [rewrite_user_prompt, retry_tool]
failure: deny
- id: acme-redaction
source: github.com/acme/claude-redaction@v1.4.2
order: 20
events: [tool.call, tool.result]
may: [transform]
sideEffects: []
failure: typed_error
- id: acme-telemetry
source: github.com/acme/claude-telemetry@v2.0.1
order: 30
events: [tool.call, tool.result]
may: [observe]
sideEffects: [otlp_internal]
failure: continue_without_export
rollback:
previousManifest: mod-stack-2026-09-28.yaml
owner: developer-platform
maxMinutes: 15
Hashes alone are not enough. Record the resolved dependency tree and any build output actually installed. If a plugin marketplace or Git URL can move a tag, mirror or vendor the reviewed artifact. Decide whether updates are automatic, staged, or prohibited in managed environments. An unsandboxed extension with silent updates is a moving privileged program.
The manifest should also state semantic invariants: one underlying tool execution per event ID; one final user-visible result; no mutation after authorization; no network side effect during validation; cancellation propagates to every wrapper; and errors retain the originating mod and event ID. These are properties the stack can test even when internal implementation changes.
Write wrappers that preserve continuation semantics
Use the generated type definitions for the installed Claude Code version and regenerate them when the API changes. The following is intentionally illustrative pseudocode because the early-access surface can move. The important properties are immutable input handling, an explicit decision, one call to next, and a typed audit record that does not contain secrets.
export default function register(api: ModAPI) {
api.on("tool.call", async (event, next) => {
const decision = policy.evaluate({
eventId: event.id,
tool: event.tool,
args: structuredClone(event.args)
});
if (!decision.allowed) {
audit.write({ eventId: event.id, outcome: "denied", rule: decision.rule });
return { kind: "blocked", reason: decision.publicReason };
}
const safeEvent = { ...event, args: decision.normalizedArgs };
const result = await next(safeEvent); // exactly once
audit.write({ eventId: event.id, outcome: result.kind });
return result;
});
}
Do not catch every error and return success. Do not invoke next in both a try block and a fallback. Do not keep mutable module globals that survive reload without a migration strategy. If a mod writes files, starts a process, or sends telemetry, make that side effect explicit, bounded, and testable. Use event IDs for deduplication and make external writes idempotent where possible.
Separate observation from authorization. A telemetry mod should not silently change a decision because an exporter is down. A policy mod should not require an analytics service to answer a local permission request unless the declared failure mode is to deny. Tight coupling turns an optional extension into an availability dependency.
Test the installed order, not isolated demos
Unit tests for each mod are necessary and insufficient. Build a harness that loads the exact production order with fake tools and deterministic events. Assert the call trace, transformed arguments, returned result, audit output, and external side effects. A minimal trace assertion catches many composition bugs:
expect(trace).toEqual([
"security:enter", "redaction:enter", "telemetry:enter",
"tool:execute:once",
"telemetry:exit", "redaction:exit", "security:exit"
]);
expect(toolExecutions).toBe(1);
expect(finalResult.eventId).toBe(input.eventId);
expect(leakedSecrets).toEqual([]);
Test allow, deny, thrown error, timeout, cancellation, malformed event, large payload, re-entrant UI event, exporter outage, partial initialization, hot reload, and downgrade. Run pairwise tests when adding a mod, then the full stack. Seed an intentionally incompatible type change to prove the build fails instead of coercing unknown data. Seed an outer wrapper that denies after the inner tool ran to expose policies that authorize too late.
| Fault | Hidden risk | Required evidence |
| Wrapper calls continuation twice | Duplicate write, payment, message, or deployment | Execution counter remains one |
| Inner mod rewrites after approval | Authorized request differs from executed request | Hash or immutable envelope spans decision to execution |
| Telemetry exporter stalls | Developer tool freezes or policy times out | Bounded timeout and declared degraded behavior |
| Hot reload retains state | Old and new handlers both run | Single registration and clean disposal trace |
| Claude Code upgrade | Event fields or ordering change | Compatibility suite passes against generated types |
| Rollback | Config reverts but artifact or cache does not | Prior manifest, hashes, and behavior restored |
Plugin validation and TypeScript compilation catch structure and type errors; they do not prove semantic composition. Keep a golden event corpus and negative tests in the repository. For external systems, use fakes or record-replay fixtures first. Our guide to testing agent skills without live APIs applies the same principle: make failure safe and deterministic before touching production.
Managed defaults are an anchor, not a complete policy
Anthropic documents a built-in sec-default mod for managed settings. It loads before user-installed mods unless administrators deliberately prepend their own. That is useful because the managed layer can observe the original inbound event. It does not automatically observe the final arguments if later mods transform them, and it does not prove that a later wrapper preserves a denial or result.
Enterprise deployment should therefore control both the allowed set and the order. Maintain an allowlist tied to reviewed commits, prohibit unreviewed local paths in managed projects, capture the effective stack at session start, and emit it with execution evidence. If local customization is allowed, define which events it may use and test an adversarial local mod against the managed controls.
Treat public catalogs as discovery tools, not trust registries. Review license, maintainer identity, repository history, dependencies, install script, build pipeline, network access, filesystem access, child processes, update channel, and requested events. A scanner flag means “inspect this capability.” Its absence does not mean safe, and a capability does not prove misuse.
Keep stronger controls outside the process. Repository branch protection, CI approvals, secret brokers, sandboxing, network policy, and target-system authorization remain valuable if Claude Code or a mod is compromised. The portable plugin trust guide covers package-level review; the execution receipts guide covers evidence across tools and approvals.
Composition failures that look fine in a demo
| Failure | Why it passes a happy path | Control |
| Order inversion | Each mod still logs a plausible event | Pin order and assert the full trace |
| Swallowed continuation | UI displays a fallback response | Test every terminal path and surface origin |
| Double continuation | Read-only test appears harmless | Use a non-idempotent fake and execution counter |
| Late authorization | Denial appears after a side effect | Bind authorization to immutable pre-effect data |
| Side-effect collision | Two retries eventually succeed | Single owner, idempotency key, retry budget |
| Version drift | Structural typing accepts incomplete data | Pin version/types and run compatibility fixtures |
| Partial reload | New UI appears while old handler remains | Dispose, re-register, and verify one handler |
| False security inheritance | Managed mod exists in the list | Attack later transforms and test final effect |
A 14-day adoption checklist
Days 1–3: inventory the desired behavior and prove a mod is the narrowest suitable surface. Record the exact Claude Code version, generated types, source revision, license, dependencies, events, machine access, data flows, and update channel. Reject opaque bundles and automatic moving references.
Days 4–6: write the composition manifest and event ownership table. Mark mutation, deny, retry, side-effect, and presentation rights. Specify continuation counts, cancellation, errors, degraded behavior, evidence, rollback owner, and time limit.
Days 7–10: build the deterministic stack harness. Run allow and deny cases, injected failures, double-call traps, post-authorization rewrites, hot reload, upgrade, downgrade, and rollback. Add hostile local-mod tests when managed settings coexist with developer customization.
Days 11–14: pilot with a small developer group on non-consequential repositories. Compare accepted tasks, false denials, latency, crashes, duplicate effects, telemetry loss, and rollback time. Capture the effective stack in every incident report. Expand only after an accountable reviewer accepts the manifest and negative results.
Release gate: no mod enters a managed stack without source review, a pinned artifact, declared event rights, a passing composition suite, and a tested rollback. “Works alone” is not an acceptance result.
FAQ
What is a Claude Code mod?
It is an early-access TypeScript extension that runs inside Claude Code and can intercept lifecycle events, prompts, tool calls, permissions, output, and UI behavior. It is executable code, not merely instructions.
Are Claude Code mods sandboxed?
No. Anthropic states that installed mods are unsandboxed and have the same machine access as Claude Code. Review source, dependencies, build artifacts, requested events, side effects, and updates before installation.
Why is individual mod testing insufficient?
Wrappers nest. A later mod can transform an already-approved event; an outer mod can change the returned result; two mods can retry or emit the same side effect. Only the exact ordered stack exposes those interactions.
Does sec-default make third-party mods safe?
No. It supplies a managed first layer and useful baseline behavior. Administrators still need allowlisting, order control, full-stack tests, external controls, effective-stack evidence, and rollback.
When should I use a skill or MCP server instead?
Use a skill for reusable guidance and an MCP server for explicit remote tools or resources. Use a mod when synchronous in-process event observation or transformation is necessary and the added coupling is justified.
Sources and research notes
Current facts were checked October 4, 2026. The mod API is early access; verify the current repository types and documentation before implementing examples.
- Anthropic, “Claude Code Mods” — launch description, event interception, built-in examples, and trust warning.
- Anthropic Claude Code repository, Mods README — setup, lifecycle, ordering, managed settings, validation, and early-access status.
- Claude Code mod type definitions — current event and API contract.
- Claude Code v2.1.287 release — release context for Mods and the built-in onboarding material.
- Anthropic Claude Code issue #91870 — public implementation discussion and limitations.
- Awesome Claude Code Mods — community catalog and capability scanner snapshot; used as discovery evidence, not a safety endorsement.
- Anthropic documentation, Claude Code plugins — package and distribution boundary.
- Anthropic documentation, hooks — comparison with process-based hooks.
- Model Context Protocol specification — comparison with protocol-defined tools and resources.
Signal scan: the unattended 30-day scan queried Reddit, Hacker News, GitHub, Polymarket, and the open web with explicit Claude Code and repository targeting. X could not be accessed in the local unattended environment, and YouTube/TikTok/Instagram were unavailable. Polymarket results referred to unrelated Claude model markets and were excluded. Social counts are contextual signals, not adoption statistics.