Enterprise AI architecture | October 1, 2026

An enterprise knowledge agent is a platform, not a RAG chatbot

Stripe's Kai case study shows the shape of a useful internal agent: persistent workspaces, isolated code execution, progressive skills, domain ownership, and a shared harness. The reusable lesson is not the chat surface. It is the contract between identity, knowledge, skills, execution, artifacts, and evidence.

Identity-scoped retrieval Domain-owned skills Artifact evidence Sources checked Oct 1
Enterprise knowledge agent architecture with identity, sources, skills, sandbox, artifacts, and evidence

The hard problem begins after retrieval works

A demo answers a question from a vector index. A production knowledge agent acts inside a company: it inherits a person's access, selects instructions written by many teams, queries changing systems, runs generated code, creates durable artifacts, and may propose writes. Those are different systems. Treating the second as a larger version of the first hides the authority and evidence boundaries that determine whether anyone can trust the result.

The renewed discussion around Stripe's Knowledge AI Platform, also called Kai, is useful because the public implementation goes beyond a search box. Stripe and LangChain describe an agent used for multi-turn analysis, document and dashboard creation, company-specific tools, and more than 1,000 skills contributed by over 100 teams. LangChain reports that 83 percent of Stripe employees use Kai weekly across more than 60,000 sessions. Those are vendor and case-study figures, not independently measured productivity or safety results. They still reveal the architectural pressure created when an internal assistant becomes common infrastructure.

At that scale, the objective cannot be “return a plausible answer.” The platform must preserve who asked, which sources were eligible, what versions were read, which skill supplied procedure, where code ran, what artifact changed, and who approved any consequence. A company can have accurate retrieval and still fail because an index flattened access controls, an old policy outranked a current one, two skills disagreed, a sandbox output lost provenance, or a polished document circulated without review.

Enterprise knowledge is not one corpus. It is a changing set of claims, permissions, procedures, and owners that an agent must join without erasing their boundaries.

What the Stripe and Deep Agents case study actually exposes

Stripe's public design has four layers. Deep Agents supplies the tool loop, state, middleware, filesystem, skills, and context management. A Stripe-specific harness connects that foundation to internal security, infrastructure, and services. Teams configure specialized agents without rewriting the harness. The Kai interface sits on top. This is a valuable separation: general agent mechanics remain below company-specific controls, while domain behavior remains above both.

The workspace is also more than conversation history. LangChain says Stripe uses an S3-backed virtual filesystem. Before a sandbox execution, relevant files are synchronized into the isolated environment; changed files are synchronized back afterward. The agent remains outside the sandbox and calls it as a tool. That distinction limits generated code to a specific execution boundary while giving long sessions a coherent workspace.

Skills solve a different pressure. With more than 500 internal MCP tools and a skill catalog beyond 1,000 entries, loading everything into the model is neither economical nor reliable. Stripe reportedly uses a two-pass pattern: choose a skill, then expose the tools allowed by that skill. Foundational skills remain pinned. Domain teams own their procedural modules. LangChain also reports quality degradation when more than roughly 150 skill descriptions compete with the system prompt, which is why Stripe is exploring hybrid retrieval or classification before final model selection.

The late-September Hacker News thread drew 189 points and 118 comments. Its skepticism is part of the evidence, not noise. Commenters questioned polish, generalizability, and how much of the success depends on Stripe's unusually mature internal infrastructure. A responsible architecture guide should preserve that limitation. A team that copies the interface without source ownership, service identity, sandbox operations, and evaluation infrastructure has not copied the system.

Public mechanismReusable lessonWhat it does not prove
Shared agent harnessCentralize execution, middleware, tracing, and platform policy.That one harness fits every risk class.
Federated skillsKeep domain procedure with accountable owners.That model-selected procedure is always correct.
Virtual filesystemMake context and artifacts inspectable across turns.That every file is current, authorized, or approved.
Sandbox as a toolSeparate generated code from the orchestrator.That output is safe to publish or write back.
High weekly usageThe interaction model may fit real work.Independent productivity, accuracy, or control effectiveness.

Separate six planes before connecting more sources

human identity + task purpose
          |
          v
1 identity plane ---- current user, role, case, delegated scope
          |
2 knowledge plane --- source version, ACL, owner, freshness, provenance
          |
3 skill plane ------- procedure, allowed tools, tests, owner, expiry
          |
4 execution plane --- sandbox, network, secrets, limits, file sync
          |
5 artifact plane ---- draft, citations, calculations, change set, approval
          |
6 evidence plane ---- trace, hashes, policy decision, reviewer, effect receipt

Identity plane: resolve the user and task before retrieval. Authorization should follow the current human or approved service identity, not a broad index credential. A finance analyst asking for a deal brief and a security engineer investigating an incident may use the same interface but must not inherit the same tools, sources, or write scope. Delegation needs an expiry, sponsor, and purpose.

Knowledge plane: retain source ACLs, document identity, version, effective date, owner, and freshness objective. Indexing is not permission. Retrieval should filter before content reaches the model, then cite the exact source object used. When a document changes, the platform should know which embeddings, cached summaries, skills, fixtures, and artifacts may now be stale.

Skill plane: treat every skill as versioned production configuration. It should declare its trigger, precedence, allowed tools, required inputs, stop conditions, output schema, owner, review date, and tests. Progressive disclosure reduces prompt load, but selection itself becomes a model decision that needs negative fixtures: tasks for which the skill must not load.

Execution plane: a sandbox limits generated code; it does not authorize network calls, secrets, or downstream effects. Materialize only the files required for the task, use short-lived credentials, default-deny egress, cap CPU/time/output, and scan synchronized artifacts before they re-enter the persistent workspace. NIST's least-privilege definition applies directly to processes acting for users.

Artifact plane: reports, dashboards, spreadsheets, and documents need a visible transition from draft to reviewed output. Store source citations, calculation inputs, code version, unresolved assumptions, and the exact artifact hash. A chat answer can disappear; an artifact can influence a decision weeks later.

Evidence plane: record the policy decision as well as the tool trace. A trace shows what the model attempted. The evidence record should also show why a source or tool was allowed, which identity was bound, which version ran, what changed, which checks passed, and whether a named reviewer released the result.

Make knowledge and skills deployable contracts

A compact manifest gives platform and domain owners something testable. It does not replace documentation; it binds the fields whose drift can change behavior. Store the canonical manifest with the skill, sign releases through the normal code-review path, and make the runtime emit the resolved version.

apiVersion: enterprise-agent.example/v1
kind: KnowledgeSkill
metadata:
  name: quarterly-deal-brief
  version: 7.2.1
  owner: sales-operations
  review_due: 2026-11-15
spec:
  trigger: "prepare an internal deal review"
  required_inputs: [account_id, review_date, requesting_user]
  sources:
    - id: crm-account
      access: inherit_requesting_user
      freshness_slo: 15m
    - id: approved-pricing-policy
      required_version: policy-2026-09
      freshness_slo: 24h
  allowed_tools: [crm.read, warehouse.query_readonly, sandbox.python]
  denied_effects: [crm.write, email.send, pricing.override]
  output_schema: deal-brief-v4
  stop_if: [missing_account_owner, stale_pricing_policy, acl_mismatch]
  approval: named_account_owner
  fixtures: [happy_path, stale_policy, cross-region_acl, missing_source]
  rollback: 7.1.4

Source precedence must be explicit. “Use the latest document” is unsafe when a future policy is published before its effective date. “Use the most relevant document” is unsafe when an obsolete playbook has stronger lexical similarity. Prefer an authority graph: official policy over team guidance, effective version over publication date, approved record over meeting note, live system state over cached summary. When sources conflict, return the conflict rather than averaging it away.

Skills need collision rules too. Two domain modules may claim the same task while assuming different geographies or systems. Resolve by declared applicability and precedence, not by whichever description the model scores highest. The related agent instruction debt guide explains how stale instructions accumulate; an enterprise catalog should make expiry and ownership part of deployment.

Re-authorize every consequential step at runtime

Do not convert successful skill selection into blanket authority. Retrieval, sandbox execution, artifact publication, and system writes are different effects. The runtime should check the current identity, task contract, resource, policy version, and proposed arguments at each boundary.

def authorize(call, run):
    skill = resolve_signed_skill(run.skill_id, run.skill_version)
    assert run.user.is_active and run.delegation.not_expired()
    assert call.tool in skill.allowed_tools
    assert call.tool not in skill.denied_effects

    resource = resolve_current_resource(call.resource_id)
    assert acl_allows(run.user, resource, call.operation)
    assert resource.version == call.expected_version

    if call.has_external_effect or call.changes_system_of_record:
        require_named_approval(run, call, exact_args_hash(call))

    return policy_receipt(run, skill, resource, call)

This policy makes revocation meaningful. If a user's role changes, a document becomes restricted, a skill expires, or a source owner withdraws approval, a long-running session should not preserve earlier access indefinitely. Re-evaluate at tool call time. Summaries and files already in the workspace may also need quarantine or redaction because authorization can change after retrieval.

Generated code needs two checks. First, execution policy limits the sandbox: image, network, secrets, mounts, duration, output size, and prohibited syscalls. Second, artifact policy evaluates what leaves it: formulas, files, plots, embedded links, macros, external references, and write targets. The execution receipts guide provides a compatible evidence model; the new layer here is knowledge and skill provenance.

Measure accepted knowledge work, not sessions

Usage tells you whether people return. It does not tell you whether sources were correct, permissions held, calculations tied, or artifacts were accepted without rework. Freeze realistic fixtures by domain and include both positive and negative cases. A sales brief fixture should test correct account scope, current pricing policy, missing CRM ownership, conflicting notes, revoked access, and a request to email the customer without approval.

MetricQuestion answeredBad shortcut
Source precisionDid every material claim cite an eligible authoritative source?“The answer sounds right.”
Freshness complianceWere source and skill SLOs satisfied at execution time?Index updated recently.
Authorization correctnessWere denied and revoked cases blocked?No access incident reported.
Accepted-artifact rateDid a qualified reviewer accept the artifact under frozen criteria?Artifact was generated.
Material correction rateHow often did reviewers change facts, calculations, or decisions?Average edit distance.
Skill-selection accuracyDid the correct procedure load and conflicting ones stay out?Tool call completed.
Cost per accepted taskWhat did useful work cost across retries and review?Cost per session.
Evidence retrieval timeCan an investigator reproduce a claim or effect quickly?Trace exists somewhere.

Segment results by domain, source version, skill version, user role, model, and risk class. A global average can hide a finance skill that fails whenever a quarter closes or a sales skill that crosses regional ACLs. Re-run fixtures whenever a model, retriever, ACL adapter, source connector, skill, sandbox image, or summarizer changes.

Common failures come from boundary collapse

FailureWhy it hidesControl
Permission flatteningIndexer credentials can read everything.Preserve resource ACLs and enforce them before model context.
Stale authority winsOld policy ranks higher than current policy.Effective-date and authority precedence, freshness stop condition.
Skill collisionTwo procedures look semantically relevant.Applicability rules, precedence, negative fixtures, named owners.
Context launderingA summary drops provenance and later looks authoritative.Carry source IDs, versions, and uncertainty through compaction.
Sandbox output escapes reviewThe code ran in isolation, so the result feels safe.Scan, validate, cite, and approve artifacts separately.
Long session keeps revoked dataThe original retrieval was authorized.Re-authorize calls and quarantine persisted context on revocation.
Adoption substitutes for qualitySession growth is easy to report.Accepted-task, correction, authorization, and incident metrics.
Domain owner disappearsThe skill still loads successfully.Expiry, ownership health, automatic disable, and rollback.

NIST's AI RMF organizes risk work around Govern, Map, Measure, and Manage, and its generative AI profile calls out provenance as a lifecycle control. Those frameworks do not prescribe this exact platform. They reinforce the operating principle: provenance, access, evaluation, and accountability must travel with the system, not appear as a final policy review.

Use a 30-day bounded rollout

Week 1 - define one domain: choose a high-frequency, read-heavy workflow with a named source owner and reviewer. Map identities, sources, effective dates, ACLs, skills, tools, artifacts, prohibited effects, and acceptance fixtures. Do not begin with company-wide search.

Week 2 - build the evidence path: preserve source versions and ACLs, implement one signed skill manifest, create read-only tool adapters, emit policy receipts, and make every material artifact carry citations, assumptions, unresolved conflicts, and the resolved skill version.

Week 3 - attack the boundaries: seed stale policy, revoked user, cross-region record, skill collision, malicious document instruction, missing source, sandbox egress attempt, oversized artifact, post-approval mutation, and a request for an unapproved write. Confirm the system blocks or exposes each case.

Week 4 - parallel production: run beside the existing process. Measure source precision, reviewer corrections, skill selection, freshness, policy decisions, acceptance rate, review time, and cost per accepted task. Expand only when a reviewer can reproduce a material claim from the evidence without reading raw chat history.

The build-versus-buy decision should follow these boundaries. Buy commodity agent mechanics when the provider exposes state, middleware, sandbox, and evaluation hooks. Build the identity, source authority, skill ownership, effect policy, and artifact-release layers that encode your company. If a platform hides those layers, speed at prototype time becomes investigation debt later.

FAQ

Is an enterprise knowledge agent just RAG plus tools?

No. RAG retrieves candidate context. A production platform also binds current identity, source authority, procedure, execution constraints, artifact state, approvals, and evidence. The difference appears when access changes, sources conflict, code runs, or an output becomes a business record.

Should every internal tool be available to the agent?

No. Load tools progressively from approved skills, pin only foundational policy, and authorize each call against the current task and identity. Large catalogs also degrade selection quality and make descriptions compete for context.

Does high weekly adoption prove safety or value?

No. It shows use. Pair adoption with source correctness, permission tests, accepted artifacts, material correction rate, incidents, reviewer effort, and total cost. Stripe's public figures are self-reported case-study evidence, not a universal benchmark.

What is the safest first pilot?

Choose one bounded domain with read-only sources, a named source and skill owner, frozen fixtures, cited artifacts, and a human release gate. Keep system-of-record and external writes out of the first phase.

Sources and further reading

Sources were checked on October 1, 2026. Stripe and LangChain metrics are treated as first-party or case-study claims, not independent benchmarks. The proposed six-plane contract is an architectural synthesis, not a claim about unpublished Stripe controls.