Developer tooling | October 4, 2026

Claude Code Mods need a composition contract

A mod can rewrite prompts, tool calls, permissions, output, and interface behavior from inside Claude Code. The important unit of review is therefore not one clever extension. It is the whole ordered stack as one program.

Event-path codeOrdered compositionStack testsSources checked Oct 4
Layered event pipeline showing several Claude Code mods around one tool call

Mods turn customization into an in-process control plane

Claude Code Mods, announced by Anthropic on October 1, are small TypeScript functions that can run before, after, instead of, or around Claude Code events. That makes them more powerful than a prompt fragment and more tightly coupled than an external integration. A mod can affect what the model sees, which tool invocation executes, whether a permission is granted, what result returns, and what the user interface displays.

The attractive story is obvious: a team can add policy, observability, interface changes, and workflow logic without forking the product. The operational story is harder. Several correct mods can become one incorrect system when they compete for the same event, assume a different order, swallow a continuation, duplicate a side effect, or target incompatible type definitions. A security mod can inspect the pre-transformation request while another mod changes it later. A telemetry wrapper can record success even though an outer wrapper replaces the result. A fallback can run the underlying tool twice.

Anthropic labels the API early access. The launch post and repository documentation also make the trust boundary explicit: mods run in process, are unsandboxed, and have the same machine access as Claude Code. A plugin may package a mod, but packaging does not restrict it. Installation is therefore equivalent to installing executable development tooling, not importing a harmless template.

Review a mod for what it does. Approve a mod stack for what the ordered composition can do.

The current developer signal is meaningful but should not be inflated. Anthropic published the feature and built-in examples; a launch discussion attracted more than a thousand Reddit votes and 182 comments; a community catalog began indexing public examples. One scanner snapshot reported 72 mods across 142 repositories and classified filesystem, process, network, and event-hook capabilities. Those are discovery and capability observations, not proof of compromise, production adoption, or safe behavior.

Use a mod only when the job belongs on the event path

Claude Code already has several extension surfaces. Choosing the narrowest one reduces coupling and review cost. A reusable instruction is usually a skill. A new remote capability may be an MCP server. A distributable bundle is a plugin. A shell hook is useful for coarse process boundaries. A mod is justified when code must synchronously observe or transform an internal event and ordinary configuration cannot express the requirement.

SurfaceBest forMain boundaryTypical review
Skill or instructionReusable method, checklist, domain guidanceInfluences model behavior; does not enforcePrompt injection, ambiguity, stale guidance
Classic hookShell-level validation or notificationSeparate process, serialized input/outputCommand execution, environment, exit behavior
MCP serverTools and resources behind a protocolSeparate server and explicit capabilitiesAuthentication, tool schema, data flow, egress
PluginDistribution bundle for commands, agents, skills, hooks, or modsPackage, not a security boundaryEvery bundled executable surface plus updates
ModLow-latency observation or transformation of Claude Code eventsIn-process and unsandboxedSource, dependencies, events, order, continuation, side effects

This distinction prevents a common failure: implementing policy in a surface that can only suggest behavior. A skill can tell the model not to expose a secret; an event interceptor can reject a request; an external credential broker can make the secret unavailable. These controls are complementary, not interchangeable. For higher-consequence actions, keep final authorization in the target system and use mods for local enforcement or evidence, not as the only boundary.

Draw the nesting before installing the code

Interceptors compose like middleware. The first loaded wrapper sees the incoming event first, calls the next handler, and receives the returning result last. That creates two orders: an inbound order and a reverse outbound order. A flat list of installed names hides this behavior.

incoming tool.call
  -> managed security mod (pre-check)
    -> redaction mod (rewrite arguments)
      -> telemetry mod (start span)
        -> Claude Code / actual tool
      <- telemetry mod (finish span)
    <- redaction mod (rewrite result)
  <- managed security mod (post-check)
outgoing result

Before installation, write an event ownership table. For each event, name who may observe, mutate, deny, replace, retry, emit a side effect, and present the final result. “All mods may wrap tool calls” is not a contract. A stronger rule might say security is the only deny authority, redaction is the only argument transformer, telemetry is observe-only, and no mod may retry a non-idempotent tool.

EventOwnerAllowed mutationContinuation ruleFailure policy
prompt.submitcontext policyAdd fixed policy context; never alter user text silentlyExactly onceReject with visible reason
permission.requestmanaged securityReduce authority onlyAt most onceFail closed
tool.callsecurity + one transformerNormalize approved fieldsExactly once unless blockedDeny or return typed error
tool.resultredactionRemove configured sensitive fieldsNo re-entryReturn original only under explicit policy
UI notificationworkflow UIPresentation onlyNever changes execution stateLog and continue

Make the stack versioned, ordered, and reviewable

A composition manifest should sit beside the project or managed configuration. It is not a Claude Code feature; it is an operational artifact that makes hidden assumptions reviewable. Pin the Claude Code range, generated plugin types, mod source revision, dependency lock, order, event ownership, side effects, failure behavior, and rollback target.

schema: claude-mod-stack/v1
claudeCode: "2.1.287"
pluginTypesSha256: "reviewed-generated-types-hash"
stack:
  - id: managed-sec-default
    source: anthropics/claude-code@reviewed-commit
    order: 10
    events: [permission.request, tool.call]
    may: [observe, deny]
    mustNot: [rewrite_user_prompt, retry_tool]
    failure: deny
  - id: acme-redaction
    source: github.com/acme/claude-redaction@v1.4.2
    order: 20
    events: [tool.call, tool.result]
    may: [transform]
    sideEffects: []
    failure: typed_error
  - id: acme-telemetry
    source: github.com/acme/claude-telemetry@v2.0.1
    order: 30
    events: [tool.call, tool.result]
    may: [observe]
    sideEffects: [otlp_internal]
    failure: continue_without_export
rollback:
  previousManifest: mod-stack-2026-09-28.yaml
  owner: developer-platform
  maxMinutes: 15

Hashes alone are not enough. Record the resolved dependency tree and any build output actually installed. If a plugin marketplace or Git URL can move a tag, mirror or vendor the reviewed artifact. Decide whether updates are automatic, staged, or prohibited in managed environments. An unsandboxed extension with silent updates is a moving privileged program.

The manifest should also state semantic invariants: one underlying tool execution per event ID; one final user-visible result; no mutation after authorization; no network side effect during validation; cancellation propagates to every wrapper; and errors retain the originating mod and event ID. These are properties the stack can test even when internal implementation changes.

Write wrappers that preserve continuation semantics

Use the generated type definitions for the installed Claude Code version and regenerate them when the API changes. The following is intentionally illustrative pseudocode because the early-access surface can move. The important properties are immutable input handling, an explicit decision, one call to next, and a typed audit record that does not contain secrets.

export default function register(api: ModAPI) {
  api.on("tool.call", async (event, next) => {
    const decision = policy.evaluate({
      eventId: event.id,
      tool: event.tool,
      args: structuredClone(event.args)
    });

    if (!decision.allowed) {
      audit.write({ eventId: event.id, outcome: "denied", rule: decision.rule });
      return { kind: "blocked", reason: decision.publicReason };
    }

    const safeEvent = { ...event, args: decision.normalizedArgs };
    const result = await next(safeEvent); // exactly once
    audit.write({ eventId: event.id, outcome: result.kind });
    return result;
  });
}

Do not catch every error and return success. Do not invoke next in both a try block and a fallback. Do not keep mutable module globals that survive reload without a migration strategy. If a mod writes files, starts a process, or sends telemetry, make that side effect explicit, bounded, and testable. Use event IDs for deduplication and make external writes idempotent where possible.

Separate observation from authorization. A telemetry mod should not silently change a decision because an exporter is down. A policy mod should not require an analytics service to answer a local permission request unless the declared failure mode is to deny. Tight coupling turns an optional extension into an availability dependency.

Test the installed order, not isolated demos

Unit tests for each mod are necessary and insufficient. Build a harness that loads the exact production order with fake tools and deterministic events. Assert the call trace, transformed arguments, returned result, audit output, and external side effects. A minimal trace assertion catches many composition bugs:

expect(trace).toEqual([
  "security:enter", "redaction:enter", "telemetry:enter",
  "tool:execute:once",
  "telemetry:exit", "redaction:exit", "security:exit"
]);
expect(toolExecutions).toBe(1);
expect(finalResult.eventId).toBe(input.eventId);
expect(leakedSecrets).toEqual([]);

Test allow, deny, thrown error, timeout, cancellation, malformed event, large payload, re-entrant UI event, exporter outage, partial initialization, hot reload, and downgrade. Run pairwise tests when adding a mod, then the full stack. Seed an intentionally incompatible type change to prove the build fails instead of coercing unknown data. Seed an outer wrapper that denies after the inner tool ran to expose policies that authorize too late.

FaultHidden riskRequired evidence
Wrapper calls continuation twiceDuplicate write, payment, message, or deploymentExecution counter remains one
Inner mod rewrites after approvalAuthorized request differs from executed requestHash or immutable envelope spans decision to execution
Telemetry exporter stallsDeveloper tool freezes or policy times outBounded timeout and declared degraded behavior
Hot reload retains stateOld and new handlers both runSingle registration and clean disposal trace
Claude Code upgradeEvent fields or ordering changeCompatibility suite passes against generated types
RollbackConfig reverts but artifact or cache does notPrior manifest, hashes, and behavior restored

Plugin validation and TypeScript compilation catch structure and type errors; they do not prove semantic composition. Keep a golden event corpus and negative tests in the repository. For external systems, use fakes or record-replay fixtures first. Our guide to testing agent skills without live APIs applies the same principle: make failure safe and deterministic before touching production.

Managed defaults are an anchor, not a complete policy

Anthropic documents a built-in sec-default mod for managed settings. It loads before user-installed mods unless administrators deliberately prepend their own. That is useful because the managed layer can observe the original inbound event. It does not automatically observe the final arguments if later mods transform them, and it does not prove that a later wrapper preserves a denial or result.

Enterprise deployment should therefore control both the allowed set and the order. Maintain an allowlist tied to reviewed commits, prohibit unreviewed local paths in managed projects, capture the effective stack at session start, and emit it with execution evidence. If local customization is allowed, define which events it may use and test an adversarial local mod against the managed controls.

Treat public catalogs as discovery tools, not trust registries. Review license, maintainer identity, repository history, dependencies, install script, build pipeline, network access, filesystem access, child processes, update channel, and requested events. A scanner flag means “inspect this capability.” Its absence does not mean safe, and a capability does not prove misuse.

Keep stronger controls outside the process. Repository branch protection, CI approvals, secret brokers, sandboxing, network policy, and target-system authorization remain valuable if Claude Code or a mod is compromised. The portable plugin trust guide covers package-level review; the execution receipts guide covers evidence across tools and approvals.

Composition failures that look fine in a demo

FailureWhy it passes a happy pathControl
Order inversionEach mod still logs a plausible eventPin order and assert the full trace
Swallowed continuationUI displays a fallback responseTest every terminal path and surface origin
Double continuationRead-only test appears harmlessUse a non-idempotent fake and execution counter
Late authorizationDenial appears after a side effectBind authorization to immutable pre-effect data
Side-effect collisionTwo retries eventually succeedSingle owner, idempotency key, retry budget
Version driftStructural typing accepts incomplete dataPin version/types and run compatibility fixtures
Partial reloadNew UI appears while old handler remainsDispose, re-register, and verify one handler
False security inheritanceManaged mod exists in the listAttack later transforms and test final effect

A 14-day adoption checklist

Days 1–3: inventory the desired behavior and prove a mod is the narrowest suitable surface. Record the exact Claude Code version, generated types, source revision, license, dependencies, events, machine access, data flows, and update channel. Reject opaque bundles and automatic moving references.

Days 4–6: write the composition manifest and event ownership table. Mark mutation, deny, retry, side-effect, and presentation rights. Specify continuation counts, cancellation, errors, degraded behavior, evidence, rollback owner, and time limit.

Days 7–10: build the deterministic stack harness. Run allow and deny cases, injected failures, double-call traps, post-authorization rewrites, hot reload, upgrade, downgrade, and rollback. Add hostile local-mod tests when managed settings coexist with developer customization.

Days 11–14: pilot with a small developer group on non-consequential repositories. Compare accepted tasks, false denials, latency, crashes, duplicate effects, telemetry loss, and rollback time. Capture the effective stack in every incident report. Expand only after an accountable reviewer accepts the manifest and negative results.

Release gate: no mod enters a managed stack without source review, a pinned artifact, declared event rights, a passing composition suite, and a tested rollback. “Works alone” is not an acceptance result.

FAQ

What is a Claude Code mod?

It is an early-access TypeScript extension that runs inside Claude Code and can intercept lifecycle events, prompts, tool calls, permissions, output, and UI behavior. It is executable code, not merely instructions.

Are Claude Code mods sandboxed?

No. Anthropic states that installed mods are unsandboxed and have the same machine access as Claude Code. Review source, dependencies, build artifacts, requested events, side effects, and updates before installation.

Why is individual mod testing insufficient?

Wrappers nest. A later mod can transform an already-approved event; an outer mod can change the returned result; two mods can retry or emit the same side effect. Only the exact ordered stack exposes those interactions.

Does sec-default make third-party mods safe?

No. It supplies a managed first layer and useful baseline behavior. Administrators still need allowlisting, order control, full-stack tests, external controls, effective-stack evidence, and rollback.

When should I use a skill or MCP server instead?

Use a skill for reusable guidance and an MCP server for explicit remote tools or resources. Use a mod when synchronous in-process event observation or transformation is necessary and the added coupling is justified.

Sources and research notes

Current facts were checked October 4, 2026. The mod API is early access; verify the current repository types and documentation before implementing examples.

  1. Anthropic, “Claude Code Mods” — launch description, event interception, built-in examples, and trust warning.
  2. Anthropic Claude Code repository, Mods README — setup, lifecycle, ordering, managed settings, validation, and early-access status.
  3. Claude Code mod type definitions — current event and API contract.
  4. Claude Code v2.1.287 release — release context for Mods and the built-in onboarding material.
  5. Anthropic Claude Code issue #91870 — public implementation discussion and limitations.
  6. Awesome Claude Code Mods — community catalog and capability scanner snapshot; used as discovery evidence, not a safety endorsement.
  7. Anthropic documentation, Claude Code plugins — package and distribution boundary.
  8. Anthropic documentation, hooks — comparison with process-based hooks.
  9. Model Context Protocol specification — comparison with protocol-defined tools and resources.

Signal scan: the unattended 30-day scan queried Reddit, Hacker News, GitHub, Polymarket, and the open web with explicit Claude Code and repository targeting. X could not be accessed in the local unattended environment, and YouTube/TikTok/Instagram were unavailable. Polymarket results referred to unrelated Claude model markets and were excluded. Social counts are contextual signals, not adoption statistics.