Source: https://internal-agents.com/infrastructure

# Internal Agents Map — Infrastructure

A source-backed catalog of AI systems organizations build for their own teams.

- Infrastructure records: 13
- Organizations: 12
- Sources: 24
- Claims: 233
- Latest entry review: 2026-09-17. Individual source dates vary.

This file holds the infrastructure collection with all of its claims, qualifications, and
sources. The compact index is at https://internal-agents.com/agents/index.json
and the complete dataset is at https://internal-agents.com/agents.json.

[Agents](https://internal-agents.com/index.md) · [Infrastructure](https://internal-agents.com/infrastructure.md). Historical JSON endpoints contain both collections.

Company-reported metrics and catalog judgments are not independent verification.
Keep the qualifications and the dates with the statements that they belong to.

## Airbnb — Airchat (airchat-cli)

Airchat is Airbnb's internal agentic-coding harness, built by its Dev AI team as a wrapper over vendor coding agents such as Claude Code and Codex. The team first built its own orchestrator from scratch, never shipped it, and delegated to Airchat with a thin shim instead.

- Company: [Airbnb](https://internal-agents.com/organizations/airbnb)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Scaled
- Evidence strength: Mixed
- Status: Internal
- First reported year: 2025
- Work: Coding, Code review
- Interfaces: Cli, Web
- Entry reviewed: 2026-09-16

Page: https://internal-agents.com/agents/airbnb-airchat

### Purpose

#### Summary

Airchat is Airbnb's internal agentic-coding harness, built by its Dev AI team as a wrapper over vendor coding agents such as Claude Code and Codex. The team first built its own orchestrator from scratch, never shipped it, and delegated to Airchat with a thin shim instead.

Fact · Reported · Medium confidence · `airbnb-airchat--summary`

Confidence reason: The build is described by Airbnb engineers in talks and podcasts, not in a first-party engineering blog.

Evidence:

- Contextualizes · [1] [Agentic coding at Airbnb (DPE.org)](https://dpe.org/sessions/szczepan-faber-mike-nakhimovich/agentic-coding-at-airbnb/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-1/content.md)
- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md)
- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md)

### Capabilities and architecture

#### Harness

Wrapper over vendor coding agents with a unified gateway, an internal plugin marketplace, and AirDev parallel workspaces

Fact · Reported · Medium confidence · `airbnb-airchat--architecture-harness`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md)

#### Model

Vendor coding agents (Claude Code and Codex are named in use), wrapped by Airbnb

Fact · Reported · Medium confidence · `airbnb-airchat--architecture-model`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md) · Preserved content.md, lines 114, 214

#### Tool access

More than a dozen internal MCP servers connect agents to internal systems

Fact · Reported · Medium confidence · `airbnb-airchat--architecture-tool-access`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md)
- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, line 166

#### Interfaces

cli, web

Fact · Reported · Medium confidence · `airbnb-airchat--architecture-interfaces`

Confidence reason: The CLI is the stated core abstraction and the Remote web UI is live as an internal early-access program.

Qualifications:

- Observation date: 2026

Evidence:

- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md) · Preserved content.md, lines 188–192, 214

#### Sandbox

AirDev remote workspaces run agent sessions network-isolated from production services and user data

Fact · Reported · Medium confidence · `airbnb-airchat--architecture-sandbox`

Confidence reason: An Airbnb engineering manager described the workspace isolation in the podcast transcript.

Qualifications:

- Observation date: 2026

Evidence:

- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md) · Preserved content.md, line 232

#### Credentials

AirChat handles authentication and permissioning for internal agent use

Fact · Reported · Medium confidence · `airbnb-airchat--architecture-credentials`

Confidence reason: An Airbnb product lead stated that AirChat handles authentication and permissioning.

Qualifications:

- Observation date: 2026

Evidence:

- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md) · Preserved content.md, line 118

#### Context management

AirChat loads default MCP servers into sessions, and the agent loop reads configuration files such as AGENTS.md and CLAUDE.md

Fact · Reported · Medium confidence · `airbnb-airchat--architecture-context-mgmt`

Confidence reason: Both direct-participant captures describe default MCP loading and configuration files the loop reads.

Qualifications:

- Observation date: 2026

Evidence:

- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md) · Preserved content.md, line 118
- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, line 82

- **Knowledge:** unreported — The captures describe internal context through MCP servers and configuration files, not a separate knowledge store.

### Documented uses

Documented use example: Engineer-written spec to reviewed pull request on the AirChat platform.

#### Start an agentic coding session

An engineer writes a short spec or prompt that starts an autonomous coding session through AirChat

Fact · Reported · Medium confidence · `airbnb-airchat--primitives-0`

Confidence reason: The reviewed talk defines the agentic loop as starting from an engineer-written spec or prompt.

Qualifications:

- Observation date: 2025-10

Evidence:

- Supports · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md) · Preserved content.md, line 118
- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, lines 82, 104

#### Materialize the code change

The session calls LLMs and internal tools in a loop until it produces a complete code change

Fact · Reported · Medium confidence · `airbnb-airchat--primitives-1`

Confidence reason: The reviewed talk's loop diagram states repeated LLM and tool calls produce the materialized change.

Qualifications:

- Observation date: 2025-10

Evidence:

- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, line 82

#### Review every line before merge

The loop ends at the diff; a human reviews every line of the change before it merges as a pull request

Fact · Reported · High confidence · `airbnb-airchat--primitives-2`

Confidence reason: A speaker quote states that every line is reviewed by a human before merge.

Qualifications:

- Observation date: 2025-10

Evidence:

- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, lines 178–182

### Access and controls

See the credential and access boundaries in architecture (`airbnb-airchat--architecture-credentials`). Scoped human-review assessments for individual uses remain in research details.

### Reliability and validation

The documented check is the human line-by-line diff review before merge; no separate automated gate is described for coding runs.

### Adoption and operating evidence

Observation: Adoption output · Reported measurement · Share of Airbnb pull requests materialized through agentic coding by the October 2025 talk

#### Headline claim

About 64% of pull requests materialized through agentic coding

Metric · Reported · Low confidence · `airbnb-airchat--headline-metric`

Confidence reason: The 64% figure comes from a third-party newsletter that quotes the engineers, not from a first-party Airbnb source.

Qualifications:

- Reported by: Airbnb
- Scope: Airbnb PRs materialized through agentic coding by the October 2025 talk
- Denominator: Airbnb pull requests; exact count not supplied
- Observation date: 2025-10

Evidence:

- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, lines 16, 24, 194

### Lessons

#### Lesson

Wrap vendor coding agents with a thin shim instead of building a full orchestrator from scratch; Airbnb's own orchestrator never shipped

Opinion · Reported · Medium confidence · `airbnb-airchat--lessons-learned-0`

Confidence reason: A direct participant described the pivot from an unshipped orchestrator to a thin shim as the team's choice, not a measured comparison.

Qualifications:

- Observation date: 2025-10

Evidence:

- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, lines 202–210

### Duplicate observation representations

Duplicate of `airbnb-airchat--headline-metric`: Identical text, scope, period, denominator, and confidence as the headline; adds nothing the canonical lacks.

#### Key observation

About 64% of pull requests materialized through agentic coding

Metric · Reported · Low confidence · `airbnb-airchat--key-metrics-0`

Confidence reason: The figure comes from a third-party newsletter, not a first-party Airbnb source.

Qualifications:

- Reported by: Airbnb
- Scope: Airbnb PRs materialized through agentic coding by the October 2025 talk
- Denominator: Airbnb pull requests; exact count not supplied
- Observation date: 2025-10

Evidence:

- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, lines 16, 24, 194

### Research details and scoped use assessments

#### Operating model assessment

Level 3 for coding task → reviewed pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence · `airbnb-airchat--operating-models-0`

Confidence reason: Airbnb engineers describe agents producing pull requests that engineers review, which locates human attention at work-product review.

Qualifications:

- Observation date: 2025

Evidence:

- Contextualizes · [2] [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md)
- Supports · [3] [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md) · Preserved content.md, lines 82, 250

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported — The documented check is the human line-by-line diff review before merge; no separate automated gate is described for coding runs.
- **observations:** Reported
- **lessons:** Reported

### Related reading

- Uses this infrastructure: [Airbnb — Datako](https://internal-agents.com/agents/airbnb-datako)
- Uses this infrastructure: [Airbnb — Pascal](https://internal-agents.com/agents/airbnb-pascal)

### Sources

1. [Agentic coding at Airbnb (DPE.org)](https://dpe.org/sessions/szczepan-faber-mike-nakhimovich/agentic-coding-at-airbnb/)
   - Talk · Direct participant · Evidence
   - Original URL: <https://dpe.org/sessions/szczepan-faber-mike-nakhimovich/agentic-coding-at-airbnb/>
   - Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-1/content.md>
2. [Beyond the CLI (DX podcast)](https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/)
   - Podcast · Direct participant · Evidence
   - Original URL: <https://getdx.com/podcast/beyond-the-cli-agentic-ai-for-async-workloads-and-non-developers/>
   - Published: 2026-06-22 · Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-2/content.md>
3. [How to get your team past the AI (The AI Thinker)](https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai)
   - News · Independent secondary · Evidence
   - Original URL: <https://www.theaithinker.com/p/how-to-get-your-team-past-the-ai>
   - Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/airbnb-airchat-source-3/content.md>

## Brex — Internal Agent Platform

Brex built a Retool-based platform where operations staff build, test and deploy agents for KYC, disputes, quality assurance and collections. Its systems engineering team maintains the builder; human involvement differs by workflow.

- Company: [Brex](https://internal-agents.com/organizations/brex)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Scaled
- Evidence strength: Secondary only
- Status: Internal
- First reported year: 2025
- Work: Finance ops, Support, Customer success
- Interfaces: Slack, Internal ui
- Entry reviewed: 2026-09-16

Page: https://internal-agents.com/agents/brex-agent-platform

### Purpose

#### Summary

Brex built a Retool-based platform where operations staff build, test and deploy agents for KYC, disputes, quality assurance and collections. Its systems engineering team maintains the builder; human involvement differs by workflow.

Fact · Reported · Medium confidence · `brex-agent-platform--summary`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)
- Contextualizes · [2] [The end of the trade-off: How AI agents broke the onboarding trilemma](https://www.brex.com/journal/rebuilding-onboarding-ai-native) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-onboarding-source/content.md) · Separate documented implementation: see dedicated agent record; not evidence of identical builder version.

### Capabilities and architecture

#### Sandbox

Retool-hosted runtime (no bespoke execution env described)

Fact · Reported · Medium confidence · `brex-agent-platform--architecture-sandbox`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

#### Harness

Retool-based builder with prompt management and multi-model testing/evaluation; built by a ~25-person systems-engineering team

Fact · Reported · Medium confidence · `brex-agent-platform--architecture-harness`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

#### Interfaces

slack, internal-ui

Fact · Reported · Medium confidence · `brex-agent-platform--architecture-interfaces`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

#### Tool access

An MCP server exposes product capabilities to internal agents. Separately, Slack /c1 requests access to AI tools through ConductorOne; it is not evidence of invoking an operations agent.

Fact · Reported · Medium confidence · `brex-agent-platform--architecture-tool-access`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

#### Knowledge

Standard operating procedures uploaded as a knowledge base, such as 100-page dispute guides; customer account data

Fact · Reported · Medium confidence · `brex-agent-platform--architecture-knowledge`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

#### Credentials

SSO via internal Retool proxies (no per-user accounts); ConductorOne access management; Okta auth; data classified by risk; ≤30-day retention, no training on inputs

Fact · Reported · Medium confidence · `brex-agent-platform--architecture-credentials`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

#### Prompt + eval studio

Non-technical ops staff design prompts, test across models, deploy with QA oversight

Fact · Reported · Medium confidence · `brex-agent-platform--primitives-0`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

#### MCP tool bridge

Product features exposed to internal agents through one MCP server

Fact · Reported · Medium confidence · `brex-agent-platform--primitives-1`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

- **Model:** unreported — Deployed-agent models are not documented; multi-model testing is build time. The placeholder claim stays in research details.
- **Context management:** unreported — No mechanism for maintaining agent context over time is documented; the QA feedback loop improves performance rather than managing context.

### Documented uses

Documented use example: Two documented agent runs on the platform: dispute-submission preparation, and collections response drafting for the servicing team.

#### Prepare the dispute submission

The dispute agent applies the 100+ page dispute guide and the BPO operating procedure to produce a more robust, better-articulated submission

Fact · Reported · Medium confidence · `brex-agent-platform--primitives-2`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 179–189

#### Draft the collections response

The agent analyzes the customer reply and layers in account context and negotiating guidelines, offering the servicing team drafts in several tones to select, customize, or send

Fact · Reported · Medium confidence · `brex-agent-platform--primitives-3`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, line 173

### Access and controls

See the credential and access boundaries in architecture (`brex-agent-platform--architecture-credentials`). Scoped human-review assessments for individual uses remain in research details.

### Reliability and validation

#### Apply the quality rubric

A QA agent reviews every support interaction against the quality rubric and feeds response trends into a closed loop improving agent and human performance

Fact · Reported · Medium confidence · `brex-agent-platform--primitives-4`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 80–82

### Adoption and operating evidence

Observation: Effectiveness · Reported measurement · Dispute-submission preparation time on the internal agent platform

#### Headline claim

Dispute-agent use example: Dispute processing time fell from three hours to three seconds

Metric · Reported · Medium confidence · `brex-agent-platform--headline-metric`

Confidence reason: A linked participant or independent source reports the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Brex
- Scope: Dispute-submission preparation using the internal agent platform; not end-to-end chargeback resolution
- Observation date: 2025

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 179–189

Observation: Adoption output · Reported measurement · Customer-support cases resolved by the chatbot at first touch (customer-facing product AI, adjacent to this platform)

#### Key observation

50%+ of customer-support cases resolved by chatbot as first touch

Metric · Reported · Medium confidence · `brex-agent-platform--key-metrics-0`

Confidence reason: A linked participant or independent source reports the claim.

Qualifications:

- Reported by: Brex
- Scope: Customer-support cases resolved by the chatbot at first touch (Brex customer-facing product AI, adjacent to this internal platform)
- Denominator: Customer-support cases; exact sample size not provided
- Observation date: 2025

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 76

Observation: Effectiveness · Reported measurement · Support-interaction QA coverage and the staffing change from five specialists to one

#### Key observation

QA covers every support interaction; one person using AI instead of five QA specialists

Metric · Reported · Medium confidence · `brex-agent-platform--key-metrics-2`

Confidence reason: A linked participant or independent source reports the claim.

Qualifications:

- Reported by: Brex
- Scope: Quality assurance of customer-support interactions
- Denominator: Every support interaction
- Method: Agent applies the quality rubric to every response; one person oversees instead of five QA specialists
- Observation date: 2025

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 80

Observation: Effectiveness · Reported measurement · KYC adverse-media classification accuracy, agent versus human

#### Key observation

KYC adverse-media accuracy 85% → 88%

Metric · Reported · Medium confidence · `brex-agent-platform--key-metrics-3`

Confidence reason: A linked participant or independent source reports the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Brex
- Scope: KYC adverse-media classification; human versus agent accuracy
- Observation date: 2025

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 163–167

### Lessons

#### Lesson

Brex's CTO recommends automating a useful portion of a workflow before pursuing complete automation.

Opinion · Reported · Medium confidence · `brex-agent-platform--lessons-learned-0`

Confidence reason: The interview contrasts a 40% target with 100% and describes fraud cases still routed to analysts. These are the CTO's priorities, not measured optimal targets.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 153-159

#### Lesson

Brex assigns a systems engineering team to maintain its internal agent platform and shares capabilities with the customer product.

Fact · Reported · Medium confidence · `brex-agent-platform--lessons-learned-1`

Confidence reason: The interview identifies the platform team and explains how new external MCP tools become available internally.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 38-40

#### Lesson

Brex automates high-confidence fraud cases and sends the remaining cases to analysts with AI-generated findings.

Fact · Reported · Medium confidence · `brex-agent-platform--lessons-learned-2`

Confidence reason: The fraud example names the automation boundary. It does not establish that all complete automation fails on accuracy.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 159

#### Lesson

Brex operations staff and engineers translated KYC procedures into individual decisions and instructions, retaining human quality checks.

Fact · Reported · Medium confidence · `brex-agent-platform--lessons-learned-3`

Confidence reason: The interview describes a joint, line-by-line review of operating procedures and subsequent human review of LLM work.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 195-205

#### Lesson

Brex groups legal approvals by data retention, training use, and segregation requirements, with controls agreed by product and legal teams.

Fact · Reported · Medium confidence · `brex-agent-platform--lessons-learned-4`

Confidence reason: The approval tracks and defined controls are described in the interview; vendor identity alone is not the stated basis for those tracks.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 54-60

### Duplicate observation representations

Duplicate of `brex-agent-platform--headline-metric`: Identical subject, values, period, scope, and locator as the headline; the shorthand adds nothing.

#### Key observation

Dispute processing: 3 hours → 3 seconds

Metric · Reported · Medium confidence · `brex-agent-platform--key-metrics-1`

Confidence reason: A linked participant or independent source reports the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Brex
- Scope: Dispute-submission preparation using the internal agent platform; not end-to-end chargeback resolution
- Observation date: 2025

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 179–189

### Research details and scoped use assessments

#### Model

Not documented for deployed agents; the platform provides multi-model testing and evaluation at build time

Fact · Reported · Medium confidence · `brex-agent-platform--architecture-model`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md)

#### Operating model assessment

Level 5 for high-confidence operations case (support first touch, fraud recommendation) → automated completion; low-confidence remainder → analyst review; human attention boundary: exception-only.

Inference · Catalog judgment · Medium confidence · `brex-agent-platform--operating-models-0`

Confidence reason: The report documents confidence-threshold automation for fraud recommendations and describes L1 process work as fully automated in most cases, with the remainder reviewed by an analyst.

Qualifications:

- Observation date: 2025-09-25

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 76, 159

#### Operating model assessment

Level 3 for KYC application → completed verification steps; human attention boundary: work-product-review.

Inference · Catalog judgment · Low confidence · `brex-agent-platform--operating-models-1`

Confidence reason: The report documents human checks inside the KYC flow and humans reviewing LLM work, but a later passage shifts humans toward reviewing trends, so the per-run boundary is mixed in the evidence.

Qualifications:

- Observation date: 2025-09-25

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, lines 143, 205
- Contextualizes · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, line 211 (humans shift toward reviewing trends)

#### Operating model assessment

Level 3 for collections follow-up interaction → drafted response options the servicing team selects, customizes, or sends; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence · `brex-agent-platform--operating-models-2`

Confidence reason: The report documents the servicing team picking among agent-drafted response options before anything is sent; initial outreach itself is automated.

Qualifications:

- Observation date: 2025-09-25

Evidence:

- Supports · [1] [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md) · Preserved content.md, line 173

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported — The prompt studio and the MCP bridge are build-time mechanisms. Platform-wide capability and the L1/L2/L3 restructure stay out of the run description.
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported — The quality-rubric agent checks every support interaction; the KYC human checks are cited with human involvement.
- **observations:** Reported
- **lessons:** Reported

### Related reading

- Related implementation: [Brex — Collections response agent](https://internal-agents.com/agents/brex-collections)
- Uses this infrastructure: [Brex — Dispute preparation agent](https://internal-agents.com/agents/brex-disputes)
- Related implementation: [Brex — Onboarding decision system](https://internal-agents.com/agents/brex-onboarding)
- Related implementation: [Brex — Support quality agent](https://internal-agents.com/agents/brex-support-qa)

### Sources

1. [Agent, Human, Ops: How Brex Is Changing Roles and Workflows](https://www.firstround.com/ai/brex)
   - Case study · Independent secondary · Evidence
   - Original URL: <https://www.firstround.com/ai/brex>
   - Published: 2025-09-25 · Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-agent-platform-source-1/content.md>
2. [The end of the trade-off: How AI agents broke the onboarding trilemma](https://www.brex.com/journal/rebuilding-onboarding-ai-native)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.brex.com/journal/rebuilding-onboarding-ai-native>
   - Published: 2026-01-20 · Accessed: 2026-09-17 · Last verified: 2026-09-17
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/brex-onboarding-source/content.md>

## Cloudflare — Internal AI engineering stack

Cloudflare's Dev Productivity team runs an internal AI engineering stack built on the company's own products. It puts MCP servers behind one OAuth portal, routes every model request through a gateway, and generates context files across thousands of repos. Every merge request gets an automated multi-agent review.

- Company: [Cloudflare](https://internal-agents.com/organizations/cloudflare)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Scaled
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2026
- Work: Coding, Code review
- Interfaces: Cli, Ci, Web
- Entry reviewed: 2026-09-16

Page: https://internal-agents.com/agents/cloudflare-ai-stack

### Purpose

#### Summary

Cloudflare's Dev Productivity team runs an internal AI engineering stack built on the company's own products. It puts MCP servers behind one OAuth portal, routes every model request through a gateway, and generates context files across thousands of repos. Every merge request gets an automated multi-agent review.

Fact · Reported · High confidence · `cloudflare-ai-stack--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)
- Contextualizes · [2] [Hacker News discussion of Cloudflare's internal AI engineering stack](https://news.ycombinator.com/item?id=47837240) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-2/content.md)

### Capabilities and architecture

#### Sandbox

Dynamic Workers for sandboxed code execution; Sandbox SDK to clone/build/test

Fact · Reported · High confidence · `cloudflare-ai-stack--architecture-sandbox`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### Harness

OpenCode + Windsurf clients; Agents SDK (McpAgent + Durable Objects) for stateful sessions

Fact · Reported · High confidence · `cloudflare-ai-stack--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### Model

Workers AI (open-weight, on-platform) + frontier models (Opus, GPT), routed by task

Fact · Reported · High confidence · `cloudflare-ai-stack--architecture-model`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### Interfaces

cli, ci, web

Fact · Reported · High confidence · `cloudflare-ai-stack--architecture-interfaces`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### Tool access

MCP Server Portal; one OAuth point aggregating 182+ tools from 13 servers; AI Gateway for routing, cost, BYOK, ZDR

Fact · Reported · High confidence · `cloudflare-ai-stack--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### Knowledge

Backstage catalog (2,055 services) + AGENTS.md generated across ~3,900 repos

Fact · Reported · High confidence · `cloudflare-ai-stack--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### Credentials

Zero API keys on client machines; a Worker injects keys server-side; Cloudflare Access (Zero Trust) auth

Fact · Reported · High confidence · `cloudflare-ai-stack--architecture-credentials`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### Context management

Code Mode collapses upstream tool schemas into search + execute, holding token overhead constant at scale

Fact · Reported · High confidence · `cloudflare-ai-stack--architecture-context-mgmt`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### MCP Server Portal

One OAuth aggregation point for all MCP tools

Fact · Reported · High confidence · `cloudflare-ai-stack--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

#### Code Mode

Searches tool schemas through a portal and exposes a separate execution function

Fact · Reported · High confidence · `cloudflare-ai-stack--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 179-189

#### AGENTS.md

Structured, generated repo context (runtime, nav, conventions, boundaries, deps)

Fact · Reported · High confidence · `cloudflare-ai-stack--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

### Documented uses

Documented use example: Merge request → AI Code Reviewer findings and approval decision.

#### AI Code Reviewer

Multi-agent CI review: risk tiering, specialist agents, Codex-rule citations

Fact · Reported · High confidence · `cloudflare-ai-stack--primitives-3`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md)

### Access and controls

See the credential and access boundaries in architecture (`cloudflare-ai-stack--architecture-credentials`). Scoped human-review assessments for individual uses remain in research details.

### Reliability and validation

#### AGENTS.md staleness flag

The AI Code Reviewer flags merge requests whose repository changes suggest the repo's AGENTS.md is outdated

Fact · Reported · High confidence · `cloudflare-ai-stack--primitives-4`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 249–251

### Adoption and operating evidence

Observation: Adoption output · Reported measurement · AI requests across the internal AI engineering system

#### Headline claim

47.95 million AI requests in 30 days across the internal AI engineering system

Metric · Reported · High confidence · `cloudflare-ai-stack--headline-metric`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Cloudflare
- Scope: Internal AI engineering requests in the 30 days preceding the report
- Method: Company-reported total internal AI request count for the preceding 30 days
- Observation date: 2026

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 14–20

Observation: Adoption output · Reported measurement · Active internal AI coding-tool users

#### Key observation

3,683 internal users (60% of company, 93% of R&D)

Metric · Reported · High confidence · `cloudflare-ai-stack--key-metrics-0`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- Reported by: Cloudflare
- Scope: Active internal AI coding-tool users in the preceding 30 days
- Denominator: Approximately 6,100 employees for company share; R&D organization for R&D share
- Observation date: 2026

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 14–16

Observation: Runtime capacity · Reported measurement · Internal AI request, AI Gateway request, and AI Gateway token volume

#### Key observation

47.95M AI requests, 20.18M AI Gateway requests per month, and 241.37B tokens routed through AI Gateway

Metric · Reported · High confidence · `cloudflare-ai-stack--key-metrics-1`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Cloudflare
- Scope: Total internal AI requests in the preceding 30 days; AI Gateway requests reported per month; AI Gateway token volume
- Method: Reported total AI request count, AI Gateway request count, and AI Gateway token count
- Observation date: 2026

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 14–20

Observation: Adoption output · Reported measurement · Company-wide merge requests during AI tool adoption

#### Key observation

10,952 merge requests in the week of March 23, 2026, nearly double the Q4 baseline; four-week average above 8,700

Metric · Reported · High confidence · `cloudflare-ai-stack--key-metrics-2`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Cloudflare
- Scope: Company merge requests in the week of March 23, 2026, versus Q4 baseline; not agent-authored PRs
- Method: Weekly merge-request count; distinct from the four-week rolling average
- Observation date: 2026-03

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 25–29

Observation: Adoption output · Reported measurement · Teams with agentic AI tool activity

#### Key observation

295 teams using agentic AI tools

Metric · Reported · High confidence · `cloudflare-ai-stack--key-metrics-3`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Cloudflare
- Scope: Teams using agentic AI tools and coding assistants in the reported 30-day snapshot
- Observation date: 2026

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 14–18

### Lessons

#### Lesson

Cloudflare added user attribution, model discovery, and permission checks at a shared proxy without changing client configurations.

Fact · Reported · Medium confidence · `cloudflare-ai-stack--lessons-learned-0`

Confidence reason: The proxy retrospective names these later additions and explains why the existing Worker provided one place to introduce them.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 105

#### Lesson

Cloudflare exposes service ownership and dependencies through Backstage and writes repository conventions into AGENTS.md files.

Fact · Reported · Medium confidence · `cloudflare-ai-stack--lessons-learned-1`

Confidence reason: The report connects missing ownership and repository context to these two mechanisms, including incorrect test commands and local conventions.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 195-212

#### Lesson

Cloudflare replaced upfront tool schemas with portal search and execution functions after measuring their context overhead.

Fact · Reported · Medium confidence · `cloudflare-ai-stack--lessons-learned-2`

Confidence reason: The Code Mode section reports 15,000 tokens for 34 GitLab schemas and names the two replacement portal functions; it does not measure an accuracy gain.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 179-189

#### Lesson

Cloudflare uses frontier models for most complex coding work and Workers AI for selected workloads, including documentation review and context-file generation.

Fact · Reported · Medium confidence · `cloudflare-ai-stack--lessons-learned-3`

Confidence reason: The workload breakdown and named Workers AI tasks support this division; future growth of open-model traffic is presented as an expectation.

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, lines 84-103

### Research details and scoped use assessments

#### Operating model assessment

Unclassified for merge request → AI review that can approve, revoke prior bot approval or block merging; a human can override; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `cloudflare-ai-stack--operating-models-0`

Confidence reason: The reviewer changes approval state and supports human override; this does not establish the required human attention boundary for every successful merge.

Qualifications:

- Observation date: 2026-09-17

Evidence:

- Supports · [3] [Orchestrating AI Code Review at scale](https://blog.cloudflare.com/ai-code-review/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-code-review-source/content.md) · The coordinator helps keep things focused; approval rubric and break glass override

#### Operating model assessment

Level 3 for AGENTS.md generation → merge request the owning team reviews and refines; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence · `cloudflare-ai-stack--operating-models-1`

Confidence reason: The source states the system opens a merge request so the owning team can review and refine the generated AGENTS.md document.

Qualifications:

- Observation date: 2026-04-20

Evidence:

- Supports · [1] [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md) · Preserved content.md, line 247 (owning team reviews and refines the generated document)

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported
- **observations:** Reported
- **lessons:** Reported

### Related reading

- Uses this infrastructure: [Cloudflare — AI Code Reviewer](https://internal-agents.com/agents/cloudflare-code-reviewer)

### Sources

1. [The AI engineering stack we built internally](https://blog.cloudflare.com/internal-ai-engineering-stack/)
   - Engineering blog · First party · Evidence
   - Original URL: <https://blog.cloudflare.com/internal-ai-engineering-stack/>
   - Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-1/content.md>
2. [Hacker News discussion of Cloudflare's internal AI engineering stack](https://news.ycombinator.com/item?id=47837240)
   - Hn thread · Community · Commentary
   - Original URL: <https://news.ycombinator.com/item?id=47837240>
   - Published: 2026-04-20 · Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-ai-stack-source-2/content.md>
3. [Orchestrating AI Code Review at scale](https://blog.cloudflare.com/ai-code-review/)
   - Engineering blog · First party · Evidence
   - Original URL: <https://blog.cloudflare.com/ai-code-review/>
   - Accessed: 2026-09-17 · Last verified: 2026-09-17
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/cloudflare-code-review-source/content.md>

## Coinbase — Mux

Mux gives Coinbase employees a shared interface for working with multiple coding agents concurrently.

- Company: [Coinbase](https://internal-agents.com/organizations/coinbase)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Deployed
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2026
- Work: Coding
- Entry reviewed: 2026-09-17

Page: https://internal-agents.com/agents/coinbase-mux

### Purpose

#### Summary

Mux gives Coinbase employees a shared interface for working with multiple coding agents concurrently.

Fact · Reported · High confidence · `coinbase-mux--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux workspace and concurrent-agent sections

### Capabilities and architecture

#### Sandbox

Separate worktrees isolate edits; they are not evidence of a security sandbox.

Fact · Reported · High confidence · `coinbase-mux--architecture-sandbox`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux workspace and concurrent-agent sections

#### Harness

Interface for concurrent coding-agent sessions.

Fact · Reported · High confidence · `coinbase-mux--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux workspace and concurrent-agent sections

- **Model:** unreported — Not documented for this subject in the reviewed source.
- **Tool access:** unreported — Not documented for this subject in the reviewed source.
- **Knowledge:** unreported — Not documented for this subject in the reviewed source.
- **Context management:** unreported — Not documented for this subject in the reviewed source.
- **Credentials:** unreported — Not documented for this subject in the reviewed source.
- **Interfaces:** unreported — Not documented for this subject in the reviewed source.

### Documented uses

Documented use example: parallel coding sessions → changes reviewed by an employee.

#### Create workspaces

Each concurrent agent receives its own git worktree, branch and terminal.

Fact · Reported · High confidence · `coinbase-mux--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux workspace and concurrent-agent sections

#### Coordinate work

Employees move between agent sessions, review results and provide feedback.

Fact · Reported · High confidence · `coinbase-mux--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux workspace and concurrent-agent sections

### Access and controls

No separate access-control assessment is recorded.

### Reliability and validation

The reviewed source does not report a distinct validation process for this subject.

### Adoption and operating evidence

Observation: Adoption output · Reported measurement · Mux: 600+ users including engineers, PMs, and designers (335 active, 197 power users)

#### Key observation

Mux: 600+ users including engineers, PMs, and designers (335 active, 197 power users)

Metric · Reported · High confidence · `coinbase-mux--key-metrics-0`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Coinbase
- Scope: Registered Mux users including engineers, PMs, and designers; 335 active and 197 power users

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux adoption and merged-PR statistics; self-selection caveat

Observation: Adoption output · Reported measurement · Mux: 5,068 merged PRs across 461 repositories and 10 orgs

#### Key observation

Mux: 5,068 merged PRs across 461 repositories and 10 orgs

Metric · Reported · Medium confidence · `coinbase-mux--key-metrics-1`

Confidence reason: Coinbase reported the aggregate in its own engineering blog without independent verification or a measurement window.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Coinbase
- Scope: Mux merged pull requests across repositories and organizations

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux adoption and merged-PR statistics; self-selection caveat

Observation: Adoption output · Reported measurement · Mux users show 3.5x more merged PRs per engineer than baseline (39.6 vs 11.4); the post cautions users may skew toward already-high-output engineers

#### Key observation

Mux users show 3.5x more merged PRs per engineer than baseline (39.6 vs 11.4); the post cautions users may skew toward already-high-output engineers

Metric · Reported · Medium confidence · `coinbase-mux--key-metrics-2`

Confidence reason: Self-reported comparison whose own text cautions that Mux users likely skew toward engineers who were already high-output and AI-forward; no cohort method is documented.

Qualifications:

- Reported by: Coinbase
- Scope: Merged pull requests per engineer for Mux users compared with a baseline of non-users
- Denominator: Engineers not using Mux at the time of the comparison (11.4 merged PRs)

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux adoption and merged-PR statistics; self-selection caveat

The reviewed source does not document this for the named subject.

### Research details and scoped use assessments

#### Operating model assessment

Level 3 for parallel coding sessions → changes reviewed by an employee; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence · `coinbase-mux--operating-models-0`

Confidence reason: The source documents this scoped human attention boundary; it is not a claim about every system at the company.

Qualifications:

- Observation date: 2026

Evidence:

- Supports · [1] [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it) · Mux workspace and concurrent-agent sections

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported — Documented use example; no universal platform run is implied.
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Unreported — The reviewed source does not report a distinct validation process for this subject.
- **observations:** Reported
- **lessons:** Unreported — The reviewed source does not document this for the named subject.

### Related reading

- Related implementation: [Coinbase — Forge](https://internal-agents.com/agents/coinbase-forge-mux)

### Sources

1. [Coding had a concurrency problem: how Mux helped solve it](https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.coinbase.com/de/blog/coding-had-a-concurrency-problem-how-mux-helped-solve-it>
   - Accessed: 2026-08-12 · Last verified: 2026-08-31

## Databricks — coSTAR

Databricks runs internal engineering agents for work such as on-call support and automated code review. coSTAR ships and tests them, using LLM judges as the test suite and a coding assistant to refine the agent until the judges pass. Databricks' Omnigent is a separate product.

- Company: [Databricks](https://internal-agents.com/organizations/databricks)
- Collection: Infrastructure
- Approach type: Component
- Deployment stage: Scaled
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2025
- Work: Coding, Code review, On-call
- Entry reviewed: 2026-09-17

Page: https://internal-agents.com/agents/databricks-costar

### Purpose

#### Summary

Databricks runs internal engineering agents for work such as on-call support and automated code review. coSTAR ships and tests them, using LLM judges as the test suite and a coding assistant to refine the agent until the judges pass. Databricks' Omnigent is a separate product.

Fact · Reported · High confidence · `databricks-costar--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md)
- Supports · [2] [Benchmarking coding agents on a multi-million line codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-2/content.md)

### Capabilities and architecture

#### Harness

coSTAR framework for shipping and testing internal agents

Fact · Reported · High confidence · `databricks-costar--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md)

#### Knowledge

Private benchmark built from the Databricks multi-million line codebase

Fact · Reported · High confidence · `databricks-costar--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Benchmarking coding agents on a multi-million line codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-2/content.md)

#### Tool access

Internal agents call MCP tools for data access, code execution, and environment setup

Fact · Reported · High confidence · `databricks-costar--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md) · Preserved content.md, line 133

#### Align judges against a human-graded Golden Set

Judge alignment refines the judges against a Golden Set of engineer-assessed outputs using MLflow alignment techniques

Fact · Reported · High confidence · `databricks-costar--primitives-3`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md) · Preserved content.md, lines 113–117

- **Model:** unreported — Neither capture states the models behind the internal engineering agents; the benchmark evaluates third-party models and harnesses, which is separate from the shipped agents.
- **Sandbox:** unreported — No runtime isolation boundary is documented for the agents; the benchmark seals git history during evaluation runs, which is not a production sandbox.
- **Context management:** unreported — MLflow traces record execution for scoring, but no capture describes how agent context is maintained or organized over time.
- **Credentials:** unreported — Credential handling is not documented in either capture.
- **Interfaces:** unreported — The interfaces list is empty; captures name harnesses only as benchmark evaluation subjects, not as surfaces of the internal agents.

### Documented uses

Documented use example: coSTAR scenario run — execute the agent, score with judges, refine until pass.

#### Run the scenario against the agent under test

The coSTAR test harness sends each scenario prompt to the agent under test and captures the execution as an MLflow trace

Fact · Reported · High confidence · `databricks-costar--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md) · Preserved content.md, lines 62–64

#### Score the trace with aligned judges

Agentic judges inspect the trace for properties such as code validity, best-practice adherence, and tool sequencing instead of asserting exact outputs

Fact · Reported · High confidence · `databricks-costar--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md) · Preserved content.md, lines 70–74

#### Refine the agent until the judges pass

A coding assistant reads judge failures, diagnoses root causes, and patches the agent while the engineer remains the reviewer and final arbiter of the proposed changes

Fact · Reported · High confidence · `databricks-costar--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md) · Preserved content.md, lines 119–121

### Access and controls

No separate access-control assessment is recorded.

### Reliability and validation

#### Run the judges in production and CI

The same judges run on production traffic, in CI/CD pipelines, and on nightly builds to catch regressions from agent or infrastructure changes

Fact · Reported · High confidence · `databricks-costar--primitives-4`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md) · Preserved content.md, lines 12, 135

### Adoption and operating evidence

Observation: Cost latency · Reported measurement · Time to verify agent changes under coSTAR, previously manual review cycles

#### Key observation

coSTAR reduced the time to verify agent changes from two weeks down to hours

Metric · Reported · Medium confidence · `databricks-costar--key-metrics-0`

Confidence reason: Databricks reported the reduction in its own engineering blog; the measurement method is not described.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Databricks
- Scope: Time to verify agent changes on the Databricks codebase under coSTAR
- Observation date: 2025

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md) · Preserved content.md, line 11

Observation: Implementation scale · Qualitative · Codebase scale behind the private coding benchmark

#### Key observation

Private benchmark built from a multi-million line codebase

Fact · Reported · Medium confidence · `databricks-costar--key-metrics-1`

Confidence reason: Databricks described the benchmark in its own engineering blog.

Evidence:

- Supports · [2] [Benchmarking coding agents on a multi-million line codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-2/content.md)

### Lessons

#### Lesson

Give judges tools, not traces: Databricks reports that agentic judges which call targeted tools on a trace scale better than feeding the full trace into a model

Opinion · Reported · Medium confidence · `databricks-costar--lessons-learned-0`

Confidence reason: Databricks authors attribute the recommendation to their own testing experience; it is an attributed preference rather than an independently verified result.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md) · Preserved content.md, lines 72–74, 180

### Research details and scoped use assessments

#### Operating model assessment

Unclassified for internal engineering workflows → agent-produced changes; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `databricks-costar--operating-models-0`

Confidence reason: The record covers several internal engineering agents with different workflows, so no single human-attention boundary applies.

Qualifications:

- Observation date: 2025

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md)

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported
- **observations:** Reported
- **lessons:** Reported

### Sources

1. [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things>
   - Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md>
2. [Benchmarking coding agents on a multi-million line codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase>
   - Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-2/content.md>

## DoorDash — Analytics AI Marketplace

DoorDash’s analytics platform hosts agents and reporting workflows over internal data and operational knowledge.

- Company: [DoorDash](https://internal-agents.com/organizations/doordash)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Deployed
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2025
- Work: Data
- Interfaces: Web, Slack, Cursor
- Entry reviewed: 2026-09-17

Page: https://internal-agents.com/agents/doordash-ai-marketplace

### Purpose

#### Summary

DoorDash’s analytics platform hosts agents and reporting workflows over internal data and operational knowledge.

Fact · Reported · High confidence · `doordash-ai-marketplace--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

### Capabilities and architecture

#### Harness

LangGraph-based execution graphs; deep-agent systems in preview and A2A swarms exploratory.

Fact · Reported · High confidence · `doordash-ai-marketplace--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

#### Knowledge

Hybrid keyword and semantic retrieval across internal knowledge.

Fact · Reported · High confidence · `doordash-ai-marketplace--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

#### Tool access

MCP tools; DescribeTable provides schema and cached examples.

Fact · Reported · High confidence · `doordash-ai-marketplace--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

#### Interfaces

web, slack, cursor

Fact · Reported · High confidence · `doordash-ai-marketplace--architecture-interfaces`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

- **Model:** unreported — Not documented for this subject in the reviewed source.
- **Sandbox:** unreported — Not documented for this subject in the reviewed source.
- **Context management:** unreported — Not documented for this subject in the reviewed source.
- **Credentials:** unreported — Not documented for this subject in the reviewed source.

### Documented uses

Documented use example: Analytics AI Marketplace request → completed work.

#### Discover an agent

Employees find specialized agents in a shared marketplace.

Fact · Reported · High confidence · `doordash-ai-marketplace--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

#### Answer with tools

Agents retrieve context and use schema-aware SQL tools.

Fact · Reported · High confidence · `doordash-ai-marketplace--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

### Access and controls

No separate access-control assessment is recorded.

### Reliability and validation

#### Validate queries

SQL linting and EXPLAIN checks precede execution; model judges evaluate response quality.

Fact · Reported · High confidence · `doordash-ai-marketplace--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

### Adoption and operating evidence

The reviewed source does not document this for the named subject.

The reviewed source does not document this for the named subject.

### Research details and scoped use assessments

#### Operating model assessment

Unclassified for Analytics AI Marketplace request → completed work; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `doordash-ai-marketplace--operating-models-0`

Confidence reason: The source describes Analytics AI Marketplace and its work, but does not establish whether a person must approve or inspect each successful output.

Qualifications:

- Observation date: 2025-11-11

Evidence:

- Supports · [1] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · Taking a high-level look; Moving forward

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported — Documented use example; no universal platform run is implied.
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported
- **observations:** Unreported — The reviewed source does not document this for the named subject.
- **lessons:** Unreported — The reviewed source does not document this for the named subject.

### Related reading

- Uses this infrastructure: [DoorDash — DataExplorer](https://internal-agents.com/agents/doordash-dataexplorer)

### Sources

1. [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/)
   - Engineering blog · First party · Evidence
   - Original URL: <https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/>
   - Published: 2025-11-11 · Accessed: 2026-08-12 · Last verified: 2026-08-31

## DoorDash — Flux

Flux is DoorDash’s cloud platform for engineering agents. It supplies isolated sandboxes, a governed MCP gateway, reusable playbooks and invocation surfaces. DoorDash’s analytics AI Marketplace is a separate documented platform.

- Company: [DoorDash](https://internal-agents.com/organizations/doordash)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Scaled
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2026
- Work: Code review, Coding, CI triage, On-call, Maintenance
- Interfaces: Slack, Github, Scheduled, Cli, Skill
- Entry reviewed: 2026-09-17

Page: https://internal-agents.com/agents/doordash-flux

### Purpose

#### Summary

Flux is DoorDash’s cloud platform for engineering agents. It supplies isolated sandboxes, a governed MCP gateway, reusable playbooks and invocation surfaces. DoorDash’s analytics AI Marketplace is a separate documented platform.

Fact · Reported · High confidence · `doordash-flux--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md) · Introduction; Primitives, not workflows

### Capabilities and architecture

#### Sandbox

Firecracker microVMs; <5s p95 end-to-end setup (boot, clone repos, install tools, configure harness)

Fact · Reported · High confidence · `doordash-flux--architecture-sandbox`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### Harness

Separate analytics AI Marketplace context (not a Flux capability): Maturity model: deterministic workflows → ReAct agents → hierarchical deep agents → experimental swarms

Fact · Reported · High confidence · `doordash-flux--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md)

#### Interfaces

slack, github, scheduled, cli, skill

Fact · Reported · High confidence · `doordash-flux--architecture-interfaces`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### Tool access

Flux Agent Gateway brokers playbook-declared tools with scoped, logged permissions. The separately documented analytics AI Marketplace uses LangGraph and explores A2A.

Fact · Reported · High confidence · `doordash-flux--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md)
- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### Knowledge

Flux playbooks package DoorDash-specific context. The separate analytics AI Marketplace hosts DataExplorer.

Fact · Reported · High confidence · `doordash-flux--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md)
- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### Credentials

Scoped per playbook; brokered through the gateway, never on the laptop; provenance on every action

Fact · Reported · High confidence · `doordash-flux--architecture-credentials`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### Context management

Separate analytics AI Marketplace context (not a Flux capability): Hybrid retrieval: BM25 + dense semantic + reciprocal-rank fusion → RAG; schema-aware SQL with EXPLAIN validation

Fact · Reported · High confidence · `doordash-flux--architecture-context-mgmt`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md)

#### Sandbox

Isolated Firecracker microVM with repos, tools, secrets, runtime deps

Fact · Reported · High confidence · `doordash-flux--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### MCP Gateway

Governed, audited access to CI, observability, issue trackers, deploy, code search

Fact · Reported · High confidence · `doordash-flux--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### Playbook

YAML unit of agentic work: task, inputs, skills, tools, permissions, validation, outputs

Fact · Reported · High confidence · `doordash-flux--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### DataExplorer

Separate analytics AI Marketplace context (not a Flux capability): Identifies schemas, generates grounded SQL, validates via EXPLAIN before execution

Fact · Reported · High confidence · `doordash-flux--primitives-3`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md)

#### Maturity model

Separate analytics AI Marketplace context (not a Flux capability): Deterministic workflows and single agents are in use; deep agents are being developed and tested; swarms remain research

Fact · Reported · High confidence · `doordash-flux--primitives-4`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md) · Preserved content.md, lines 66

- **Model:** unreported — No model or provider is named for either platform. The modular-primitives design claim stays in research details.

### Documented uses

Documented use example: Business question to grounded, validated answer on the analytics platform.

#### Answer a grounded data question

Separate analytics AI Marketplace context (not a Flux capability): From a business question, DataExplorer identifies candidate tables through DescribeTable with cached column examples, generates starter SQL grounded in the schema, and EXPLAIN validation with autocorrection precedes execution

Fact · Reported · High confidence · `doordash-flux--primitives-5`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md) · Preserved content.md, lines 42, 78–80

### Access and controls

See the credential and access boundaries in architecture (`doordash-flux--architecture-credentials`). Scoped human-review assessments for individual uses remain in research details.

### Reliability and validation

The analytics platform validates generated SQL with linting, EXPLAIN checks, and statistical-metadata checks with autocorrection, plus a separate LLM-as-judge and DeepEval framework. Flux playbooks package validation per playbook.

### Adoption and operating evidence

Observation: Adoption output · Reported measurement · Engineering tasks automated on Flux in one reported month; the calendar month is unspecified

#### Headline claim

130,000 engineering tasks automated in one month

Metric · Reported · Medium confidence · `doordash-flux--headline-metric`

Confidence reason: Dated August 11, 2026 report of a one-month count; this is the observation date, not the measurement window.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: DoorDash
- Scope: Engineering tasks automated in one reported month; calendar measurement month unspecified
- Observation date: 2026-08-11

Evidence:

- Supports · [1] [Delegating Engineering Work To Cloud-Based Agents (Flux)](https://x.com/AIatDoorDash/status/2087285008906240193) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-1/content.md) · Preserved content.md, lines 10–12

Observation: Adoption output · Reported measurement · Weekly automated code reviews powered by Flux; the reviewer agent itself is the separate doordash-code-review record

#### Key observation

25,000+ automated code reviews per week

Metric · Reported · Medium confidence · `doordash-flux--key-metrics-1`

Confidence reason: Dated August 11, 2026 report; weekly measurement boundaries are not supplied.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: DoorDash
- Scope: Weekly automated code reviews powered by Flux
- Observation date: 2026-08-11

Evidence:

- Supports · [1] [Delegating Engineering Work To Cloud-Based Agents (Flux)](https://x.com/AIatDoorDash/status/2087285008906240193) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-1/content.md) · Preserved content.md, lines 10–12

Observation: Adoption output · Reported measurement · Unique Flux playbooks and weekly playbook invocations

#### Key observation

300+ playbooks; 10,000+ invocations per week

Metric · Reported · High confidence · `doordash-flux--key-metrics-2`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: DoorDash
- Scope: Unique playbooks and weekly invocations on Flux

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md) · Preserved content.md, lines 10

### Lessons

#### Lesson

DoorDash first used Flux for code review, then expanded into CI triage, on-call work, maintenance, and ticket-driven development.

Fact · Reported · Medium confidence · `doordash-flux--lessons-learned-0`

Confidence reason: The Flux retrospective lists the initial workflow and later uses; the sequence alone does not prove that this rollout order caused adoption.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md) · Preserved content.md, lines 85

#### Lesson

DoorDash credits a move from private run channels to public Slack threads with helping teams learn to use Flux.

Opinion · Reported · Medium confidence · `doordash-flux--lessons-learned-1`

Confidence reason: The authors describe the visibility change and its observed adoption pattern, without a controlled comparison of channel choices.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md) · Preserved content.md, lines 86

#### Lesson

DoorDash ran workshops and hackathons to help teams turn recurring operational tasks into Flux playbooks.

Fact · Reported · Medium confidence · `doordash-flux--lessons-learned-2`

Confidence reason: The retrospective names these activities and the work they supported, rather than reporting that a playbook library appeared automatically.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md) · Preserved content.md, lines 87

#### Lesson

DoorDash reports deterministic workflows and single agents in use, deep agents under development, and swarms as research on its analytical platform.

Fact · Reported · Medium confidence · `doordash-flux--lessons-learned-3`

Confidence reason: The architecture article distinguishes these deployment stages. Its swarm discussion identifies governance and explainability as unresolved work.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md) · Preserved content.md, lines 58-66

#### Lesson

DoorDash's analytical platform uses SQL linting and EXPLAIN checks, with separate model judges for response quality.

Fact · Reported · Medium confidence · `doordash-flux--lessons-learned-4`

Confidence reason: The validation sections distinguish query checks from LLM-as-judge and DeepEval assessments; a valid query is not proof of a correct business answer.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md) · Preserved content.md, lines 80-82

#### Lesson

DoorDash describes recording the queries, documents, and agent interactions behind an analytical answer.

Fact · Reported · Medium confidence · `doordash-flux--lessons-learned-5`

Confidence reason: The provenance discussion names these records as the audit trail; recording them does not itself establish that the answer is correct.

Evidence:

- Supports · [2] [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md) · Preserved content.md, lines 94

### Duplicate observation representations

Duplicate of `doordash-flux--headline-metric`: Identical statement, scope, period, and locators as the headline: the same 130,000-task one-month count.

#### Key observation

130,000 engineering tasks automated in one month

Metric · Reported · Medium confidence · `doordash-flux--key-metrics-0`

Confidence reason: Dated August 11, 2026 report of a one-month count; this is the observation date, not the measurement window.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: DoorDash
- Scope: Engineering tasks automated in one reported month; calendar measurement month unspecified
- Observation date: 2026-08-11

Evidence:

- Supports · [1] [Delegating Engineering Work To Cloud-Based Agents (Flux)](https://x.com/AIatDoorDash/status/2087285008906240193) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-1/content.md) · Preserved content.md, lines 10–12

### Research details and scoped use assessments

#### Model

Sources name no model or provider; the documented primitives are modular, allowing the best third-party tool for each job or in-house builds, with multiple supported coding agent harnesses

Fact · Reported · High confidence · `doordash-flux--architecture-model`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

#### Operating model assessment

Level 3 for engineering task → reviewed agent output; human attention boundary: work-product-review.

Inference · Catalog judgment · Medium confidence · `doordash-flux--operating-models-0`

Confidence reason: The platform spans several workflows; the cited engineering examples retain human review of agent output.

Qualifications:

- Observation date: 2025-11-11

Evidence:

- Supports · [3] [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md)

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported — Historical analytics example retained for anchor compatibility; it belongs to DataExplorer on the separate AI Marketplace, not Flux.
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported — The analytics platform validates generated SQL with linting, EXPLAIN checks, and statistical-metadata checks with autocorrection, plus a separate LLM-as-judge and DeepEval framework. Flux playbooks package validation per playbook.
- **observations:** Reported
- **lessons:** Reported

### Related reading

- Includes component: [DoorDash — AI Code Review Agent](https://internal-agents.com/agents/doordash-code-review)

### Sources

1. [Delegating Engineering Work To Cloud-Based Agents (Flux)](https://x.com/AIatDoorDash/status/2087285008906240193)
   - Social post · Direct participant · Evidence
   - Original URL: <https://x.com/AIatDoorDash/status/2087285008906240193>
   - Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-1/content.md>
2. [Beyond single agents: DoorDash's collaborative AI ecosystem](https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/)
   - Engineering blog · First party · Evidence
   - Original URL: <https://careersatdoordash.com/blog/beyond-single-agents-doordash-building-collaborative-ai-ecosystem/>
   - Published: 2025-11-11 · Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-2/content.md>
3. [Delegating Engineering Work To Cloud-Based Agents](https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/)
   - Engineering blog · First party · Evidence
   - Original URL: <https://careersatdoordash.com/blog/delegating-engineering-work-to-cloud-based-agents/>
   - Published: 2026-08-11 · Accessed: 2026-08-31 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/doordash-flux-source-3/content.md>

## Dropbox — Nova

Nova is Dropbox's internal service for running coding agents in its cloud. Engineers launch parallel sessions from a web UI, CLI, or API, and internal systems call agents inside automated workflows like CI triage and migrations. Callers attach validation commands that Nova runs after each attempt.

- Company: [Dropbox](https://internal-agents.com/organizations/dropbox)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Scaled
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2026
- Work: Coding, CI triage, On-call, Maintenance
- Interfaces: Web, Cli, Api, Slack
- Entry reviewed: 2026-09-16

Page: https://internal-agents.com/agents/dropbox-nova

### Purpose

#### Summary

Nova is Dropbox's internal service for running coding agents in its cloud. Engineers launch parallel sessions from a web UI, CLI, or API, and internal systems call agents inside automated workflows like CI triage and migrations. Callers attach validation commands that Nova runs after each attempt.

Fact · Reported · High confidence · `dropbox-nova--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

### Capabilities and architecture

#### Sandbox

Isolated env with a codebase snapshot at a specific commit; full Dropbox monorepo via Bazel; hermetic remote execution + caching

Fact · Reported · High confidence · `dropbox-nova--architecture-sandbox`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Harness

Validation loop (propose → validate → feed back) with continue_on_validation_failure and max_iterations (~5); branch management kept outside the agent

Fact · Reported · High confidence · `dropbox-nova--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Model

Multiple coding agents behind one interface; prompt-evaluation tooling; helpers that make it easier to add AI-powered steps without rebuilding surrounding infrastructure; sources name no specific model or provider

Fact · Reported · High confidence · `dropbox-nova--architecture-model`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Interfaces

web, cli, api, slack

Fact · Reported · High confidence · `dropbox-nova--architecture-interfaces`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Tool access

Skills/plugins to gather evidence, read logs, inspect failures; MCP integrations; Bazel-aware selectivity tools

Fact · Reported · High confidence · `dropbox-nova--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Knowledge

Localized AGENTS.md per service; Dash (Dropbox context engineering); passing + failing test logs

Fact · Reported · High confidence · `dropbox-nova--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Context management

Session history (notes/logs) carried across retry attempts

Fact · Reported · High confidence · `dropbox-nova--architecture-context-mgmt`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Validation loop

Bounded iteration with feedback on failure; deterministic systems control test execution

Fact · Reported · High confidence · `dropbox-nova--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Dash

Nova is expanding its context sources, including Dash and MCP-based integrations

Fact · Reported · High confidence · `dropbox-nova--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 30,82

- **Credentials:** unreported — The capture states agents operate within existing infrastructure and validation paths but names no authentication or authorization mechanism; the legacy same-as-engineers claim stays in research details.

### Documented uses

Documented use example: Athena flaky-test alert through the Deflaker fix-and-validate loop to a landed fix or capped attempts.

#### Ingest flaky-test evidence

Athena detects a flaky test; the Deflaker workflow sends its passing and failing logs to Nova as context and asks the agent to identify a likely root cause

Fact · Reported · High confidence · `dropbox-nova--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 58, 62

#### Propose a root-cause fix

The Nova agent proposes a fix for the flaky test

Fact · Reported · High confidence · `dropbox-nova--primitives-3`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, line 58

#### Retry with carried-forward notes

A flake starts another attempt with new logs and notes from the previous one, until a working fix lands or attempts reach the cap of five

Fact · Reported · High confidence · `dropbox-nova--primitives-5`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 58, 62

### Access and controls

No separate access-control assessment is recorded.

### Reliability and validation

#### Validate the fix in CI

CI runs the test 100 or more times depending on its failure rate to check the proposed change

Fact · Reported · High confidence · `dropbox-nova--primitives-4`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, line 58

### Adoption and operating evidence

Observation: Runtime capacity · Qualitative · Parallel agents a migration owner launches and manages from one runbook

#### Headline claim

Dozens of agents can run in parallel from one runbook

Metric · Reported · High confidence · `dropbox-nova--headline-metric`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Dropbox
- Scope: Migration-owner orchestration of dozens of agents from a shared runbook; qualitative capacity description

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 65–69

Observation: Runtime capacity · Reported measurement · CI validation runs per proposed Deflaker flaky-test fix

#### Key observation

Flaky-test remediation (Deflaker): 100+ validation runs

Metric · Reported · High confidence · `dropbox-nova--key-metrics-0`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Dropbox
- Scope: CI validation runs per proposed Deflaker flaky-test fix
- Method: Run the test 100 or more times depending on its failure rate; retry capped at five fix attempts

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 58–62

Observation: Adoption output · Reported measurement · Migration entries processed by the predecessor Goose-based migrator before workflows moved onto Nova

#### Key observation

Predecessor Goose-based migrator used across thousands of migration entries before workflows moved onto Nova

Metric · Reported · High confidence · `dropbox-nova--key-metrics-1`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Dropbox
- Scope: Predecessor Goose-based migrator, before workflows moved onto Nova

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 65–69

### Lessons

#### Lesson

Dropbox attributes much of Nova's usefulness to its surrounding execution and validation infrastructure.

Opinion · Reported · Medium confidence · `dropbox-nova--lessons-learned-0`

Confidence reason: The retrospective explicitly makes this assessment and names context files, tests, caching, isolation, and retry loops; it provides no controlled component comparison.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 78

#### Lesson

Nova can return failed validation results to the agent and continue the session to repair the proposed change.

Fact · Reported · Medium confidence · `dropbox-nova--lessons-learned-1`

Confidence reason: The session example describes optional validation commands and continuation on failure, making the feedback mechanism explicit.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 26

#### Lesson

Dropbox moved CI triggering outside the agent after seeing waits of hours and validation against the wrong tests.

Fact · Reported · Medium confidence · `dropbox-nova--lessons-learned-2`

Confidence reason: The retrospective names these failure modes and explains that surrounding workflows trigger CI and call the agent back for repair.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 80

#### Lesson

Nova reuses Dropbox's Bazel and on-premise validation paths because those systems support its monorepo development workflow.

Fact · Reported · Medium confidence · `dropbox-nova--lessons-learned-3`

Confidence reason: Dropbox gives its repository shape and existing build infrastructure as reasons for integration, rather than a separate agent-only workflow.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 20-22,55

### Duplicate observation representations

Duplicate of `dropbox-nova--headline-metric`: Same dozens-of-agents capacity from the same runbook passage; identical scope with no separate period or qualification.

#### Key observation

Dozens of agents launchable from one runbook

Metric · Reported · High confidence · `dropbox-nova--key-metrics-2`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Dropbox
- Scope: Migration-owner orchestration of dozens of agents from a shared runbook; qualitative capacity description

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 65–69

### Research details and scoped use assessments

#### Credentials

Operates within Dropbox's existing infra and validation paths; same auth/authz as engineers

Fact · Reported · High confidence · `dropbox-nova--architecture-credentials`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md)

#### Operating model assessment

Unclassified for event-triggered validation-gated remediation (Deflaker, crash-alert candidate fixes) → landed fix or candidates routed to service teams; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `dropbox-nova--operating-models-0`

Confidence reason: Deflaker automates bounded repair and CI validation, while publication stays outside the agent. The source does not specify a required human approval or exception-handling boundary.

Qualifications:

- Observation date: 2026-05-22

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 18, 58, 62 (async workflows; capped fix-and-validate loop with no person in the loop)
- Contextualizes · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, line 32 (publication kept outside the agent)

#### Operating model assessment

Level 2 for interactive developer session from web, CLI, or Slack → agent-assisted change; human attention boundary: continuous-steering.

Inference · Catalog judgment · Medium confidence · `dropbox-nova--operating-models-1`

Confidence reason: The source describes standard interactive chat sessions in which engineers launch and steer agents; the exact review boundary of those sessions is not further specified.

Qualifications:

- Observation date: 2026-05-22

Evidence:

- Supports · [1] [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md) · Preserved content.md, lines 18, 26 (standard interactive chat; Nova continues the session on validation failure)

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported — The Deflaker run is the documented operational workflow; interactive sessions and the crash-alert experiment are other modes on the platform, outside this scope.
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported — Caller-attached validation commands run after each attempt; the Deflaker CI runs are the documented instance.
- **observations:** Reported
- **lessons:** Reported

### Related reading

- Uses this infrastructure: [Dropbox — Deflaker](https://internal-agents.com/agents/dropbox-deflaker)

### Sources

1. [Introducing Nova, our internal platform for coding agents](https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents)
   - Engineering blog · First party · Evidence
   - Original URL: <https://dropbox.tech/machine-learning/introducing-nova-our-internal-platform-for-coding-agents>
   - Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-1/content.md>
2. [Hacker News submission for Nova](https://news.ycombinator.com/item?id=48235065)
   - Hn thread · Community · Discovery
   - Original URL: <https://news.ycombinator.com/item?id=48235065>
   - Published: 2026-05-22 · Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/dropbox-nova-source-2/content.md>

## Harvey — Spectre

Harvey's internal collaborative cloud agent platform; turns requests from Slack, the web app, or automations into durable runs that return reviewable summaries, diffs, branches, and PRs.

- Company: [Harvey](https://internal-agents.com/organizations/harvey)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Deployed
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2026
- Work: Coding, Code review, On-call, Security
- Interfaces: Slack, Web, Automation, Cli
- Entry reviewed: 2026-09-16

Page: https://internal-agents.com/agents/harvey-spectre

### Purpose

#### Summary

Harvey's internal collaborative cloud agent platform; turns requests from Slack, the web app, or automations into durable runs that return reviewable summaries, diffs, branches, and PRs.

Fact · Reported · High confidence · `harvey-spectre--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

### Capabilities and architecture

#### Sandbox

Isolated ephemeral execution environments; durable runs

Fact · Reported · High confidence · `harvey-spectre--architecture-sandbox`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

#### Harness

Collaborative cloud agent platform with explicit boundaries around GitHub, Datadog, Linear, and other connected systems

Fact · Reported · High confidence · `harvey-spectre--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

#### Interfaces

slack, web, automation, cli

Fact · Reported · High confidence · `harvey-spectre--architecture-interfaces`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md) · Preserved content.md, lines 16, 40

#### Tool access

Explicit tool boundaries; tool configuration injected at run start; requests start from Slack, the web app, or automations; connects to systems like GitHub, Datadog, and Linear

Fact · Reported · High confidence · `harvey-spectre--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

#### Credentials

Short-lived, scoped credentials injected at run start; no ambient access to the control plane or a user's machine

Fact · Reported · High confidence · `harvey-spectre--architecture-credentials`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

#### Context management

Durable runs over disposable execution environments

Fact · Reported · High confidence · `harvey-spectre--architecture-context-mgmt`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

#### Durable runs / ephemeral execution

Long-lived run state over throwaway compute

Fact · Reported · High confidence · `harvey-spectre--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

#### Reviewable artifacts

Returns summaries, diffs, branches, and pull requests alongside the run history

Fact · Reported · High confidence · `harvey-spectre--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md) · Preserved content.md, lines 16,24

- **Model:** unreported — The post names providers and multi-provider support but no model; the placeholder claim stays in research details.
- **Knowledge:** unreported — No knowledge store for Spectre runs is documented; the SOC's memory system belongs to the separate security platform.

### Documented uses

Documented use example: Request from Slack, the web app, an automation, or the CLI through a sandboxed durable run to reviewable artifacts.

#### Start a durable run

A request from Slack, the web app, an automation, or the CLI creates a durable run and starts a fresh sandbox that hydrates the repository

Fact · Reported · High confidence · `harvey-spectre--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md) · Preserved content.md, lines 16, 24, 40

#### Execute the agent in the sandbox

The worker runs the agent, streams output, and reaches approved external systems through tool configuration injected at run start

Fact · Reported · High confidence · `harvey-spectre--primitives-3`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md) · Preserved content.md, lines 24, 44, 46

#### Return reviewable artifacts

The run persists summaries, diffs, branches, and pull requests; a failed run returns the result to a person who decides what happens next

Fact · Reported · High confidence · `harvey-spectre--primitives-4`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md) · Preserved content.md, lines 16, 24

### Access and controls

See the credential and access boundaries in architecture (`harvey-spectre--architecture-credentials`). Scoped human-review assessments for individual uses remain in research details.

### Reliability and validation

#### Scheduled verification loops

Cron-based automations materialize ordinary runs for cleanup passes, test generation, dependency checks, and verification loops

Fact · Reported · High confidence · `harvey-spectre--primitives-5`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md) · Preserved content.md, lines 68, 74

### Adoption and operating evidence

The platform post reports no measured outcomes. The security article's figures describe the separate agentic SOC, which runs on its own infrastructure by design.

### Lessons

#### Lesson

Spectre returns branches, diffs, summaries, and run history so another person can inspect or continue the work.

Fact · Reported · Medium confidence · `harvey-spectre--lessons-learned-0`

Confidence reason: The platform introduction and execution flow name reviewable artifacts and a human decision point when a run fails.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md) · Preserved content.md, lines 16,24

#### Lesson

Harvey identifies coordination, ownership, and retained context as challenges when many people and agents share work.

Opinion · Reported · Medium confidence · `harvey-spectre--lessons-learned-1`

Confidence reason: The authors frame these as engineering challenges and describe a durable run as their response; they do not quantify a shift in review bottlenecks.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md) · Preserved content.md, lines 88

#### Lesson

Harvey keeps Spectre and its security agents on separate infrastructure to limit the consequences of a compromised product agent.

Opinion · Reported · Medium confidence · `harvey-spectre--lessons-learned-2`

Confidence reason: The security article explains that sharing detection knowledge and investigation memory would expose information valuable to an attacker; this is Harvey's design rationale.

Evidence:

- Supports · [2] [Building an agentic security operations center](https://www.harvey.ai/blog/building-an-agentic-security-operations-center) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-2/content.md) · Preserved content.md, lines 115

### Research details and scoped use assessments

#### Model

Not specified

Fact · Reported · High confidence · `harvey-spectre--architecture-model`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

#### Operating model assessment

Level 3 for request from Slack, the web app, or an automation → reviewable diff or pull request; human attention boundary: work-product-review.

Inference · Catalog judgment · High confidence · `harvey-spectre--operating-models-0`

Confidence reason: The source explicitly frames diffs, branches, and pull requests as reviewable outputs.

Qualifications:

- Observation date: 2026

Evidence:

- Supports · [1] [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md)

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported — The named run is the first-version flow the post documents: create a run, start a fresh sandbox, hydrate the repository, run the agent, stream output, persist artifacts.
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported — Verification loops run as scheduled ordinary runs; the post documents no separate check inside an interactive run beyond the review boundary.
- **observations:** Unreported — The platform post reports no measured outcomes. The security article's figures describe the separate agentic SOC, which runs on its own infrastructure by design.
- **lessons:** Reported

### Related reading

- Related implementation: [Harvey — Security operations agents](https://internal-agents.com/agents/harvey-security-operations)

### Sources

1. [Building Spectre; internal collaborative cloud agent platform](https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.harvey.ai/blog/building-spectre-internal-collaborative-cloud-agent-platform>
   - Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-1/content.md>
2. [Building an agentic security operations center](https://www.harvey.ai/blog/building-an-agentic-security-operations-center)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.harvey.ai/blog/building-an-agentic-security-operations-center>
   - Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/harvey-spectre-source-2/content.md>

## Notion — Custom Agents

Custom Agents is Notion's platform for building agents in a Notion workspace, and Notion ships it internally first. Staff in IT ticketing, supply chain procurement, and recruiting build their own agents. One early internal agent triages bugs posted in Slack into task-database entries.

- Company: [Notion](https://internal-agents.com/organizations/notion)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Scaled
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2026
- Work: Support, Finance ops, Recruitment, Security
- Interfaces: Slack
- Entry reviewed: 2026-09-16

Page: https://internal-agents.com/agents/notion-custom-agents

### Purpose

#### Summary

Custom Agents is Notion's platform for building agents in a Notion workspace, and Notion ships it internally first. Staff in IT ticketing, supply chain procurement, and recruiting build their own agents. One early internal agent triages bugs posted in Slack into task-database entries.

Fact · Reported · High confidence · `notion-custom-agents--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space)](https://latent.space/p/notion) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-1/content.md)
- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md)
- Contextualizes · [3] [Meet Scruff, Security’s New AI Teammate](https://www.notion.com/blog/meet-scruff-securitys-new-ai-teammate) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-scruff-source/content.md) · Separate documented implementation: see dedicated agent record; not evidence of identical builder version.

### Capabilities and architecture

#### Harness

Shared Custom Agents harness with team-owned tools and evaluations

Fact · Reported · Medium confidence · `notion-custom-agents--architecture-harness`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space)](https://latent.space/p/notion) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-1/content.md) · Preserved content.md, lines 405, 609

#### Tool access

Agents start without access; owners grant resource and tool permissions for the intended work

Fact · Reported · High confidence · `notion-custom-agents--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md) · Preserved content.md, lines 22–37

#### Interfaces

slack

Fact · Reported · Medium confidence · `notion-custom-agents--architecture-interfaces`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space)](https://latent.space/p/notion) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-1/content.md) · Preserved content.md, line 609

#### Knowledge

Workspace pages and connected resources granted to the individual agent

Fact · Reported · High confidence · `notion-custom-agents--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md) · Preserved content.md, lines 30–37

#### Credentials

Deterministic resource and action permissions; risky runtime actions require owner confirmation

Fact · Reported · High confidence · `notion-custom-agents--architecture-credentials`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md) · Preserved content.md, lines 22–45

#### Context management

Tool definitions and progressive disclosure expose capabilities as needed

Fact · Reported · Medium confidence · `notion-custom-agents--architecture-context-mgmt`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space)](https://latent.space/p/notion) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-1/content.md) · Preserved content.md, lines 707–717

#### Default-deny permission model

Each Custom Agent starts without access to most resources and receives explicit page

Fact · Reported · High confidence · `notion-custom-agents--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md) · Preserved content.md, lines 22–37

- **Model:** unreported — The captures do not identify one model for this representative workflow.
- **Sandbox:** unreported — Execution isolation for the representative workflow is not documented.

### Documented uses

Documented use example: Internal Slack bug report → routed task-database entry.

#### Slack bug-triage agent

A documented internal agent receives a bug posted in Slack, routes it to the responsible team, creates a task-database entry, and replies in the triggering channel

Fact · Reported · Medium confidence · `notion-custom-agents--primitives-0`

Confidence reason: A linked participant or independent source reports the claim.

Evidence:

- Supports · [1] [Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space)](https://latent.space/p/notion) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-1/content.md) · Preserved content.md, line 609

### Access and controls

See the credential and access boundaries in architecture (`notion-custom-agents--architecture-credentials`). Scoped human-review assessments for individual uses remain in research details.

### Reliability and validation

#### Risk confirmation and remediation

Potentially risky runtime actions pause for confirmation from the agent owner

Fact · Reported · High confidence · `notion-custom-agents--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md) · Preserved content.md, lines 41–47

### Adoption and operating evidence

Observation: Implementation scale · Reported measurement · Internal Custom Agents at the end of alpha testing

#### Headline claim

More than 3,000 internal Custom Agents by end of alpha testing

Metric · Reported · Medium confidence · `notion-custom-agents--headline-metric`

Confidence reason: Notion reported the agent count in its own engineering blog without independent verification.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Notion
- Scope: Internal Notion Custom Agents at the end of alpha; excludes the separate customer alpha count
- Observation date: 2026-04

Evidence:

- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md) · Preserved content.md, lines 51

Observation: Adoption output · Qualitative · Notion security team's internal Custom Agent use

#### Key observation

Notion's security team is one of the most active internal users

Fact · Reported · Medium confidence · `notion-custom-agents--key-metrics-1`

Confidence reason: Qualitative first-party description; no comparative activity count is supplied.

Qualifications:

- Scope: Notion security team internal Custom Agent use

Evidence:

- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md) · Preserved content.md, lines 57

Observation: Implementation scale · Estimate · Rebuilds of Notion's agent harness and framework

#### Key observation

Agent harness rebuilt three to five times as models improved

Metric · Reported · Medium confidence · `notion-custom-agents--key-metrics-2`

Confidence reason: Participants give differing approximate rebuild counts for harness, framework, and feature; not a precise engineering inventory.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Notion
- Scope: Notion agent/framework rebuilds recalled by participants; estimates range from three to five

Evidence:

- Supports · [1] [Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space)](https://latent.space/p/notion) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-1/content.md) · Preserved content.md, lines 115, 275–277, 819

The record contains no separate lessons claim after review of both captures.

### Duplicate observation representations

Duplicate of `notion-custom-agents--headline-metric`: Same count, population, period, and qualification as the headline observation.

#### Key observation

More than 3,000 internal Custom Agents by end of alpha testing

Metric · Reported · Medium confidence · `notion-custom-agents--key-metrics-0`

Confidence reason: Notion reported the agent count in its own engineering blog.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Notion
- Scope: Internal Notion Custom Agents at the end of alpha; excludes the separate customer alpha count

Evidence:

- Supports · [2] [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md) · Preserved content.md, lines 51

### Research details and scoped use assessments

#### Operating model assessment

Unclassified for cross-team internal tasks → Custom Agents output; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `notion-custom-agents--operating-models-0`

Confidence reason: Custom Agents spans many teams and workflows, so no single human-attention boundary applies.

Qualifications:

- Observation date: 2026-04

Evidence:

- Supports · [1] [Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space)](https://latent.space/p/notion) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-1/content.md)

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported
- **observations:** Reported
- **lessons:** Unreported — The record contains no separate lessons claim after review of both captures.

### Related reading

- Uses this infrastructure: [Notion — Internal bug-triage agent](https://internal-agents.com/agents/notion-bug-triage)
- Uses this infrastructure: [Notion — Scruff](https://internal-agents.com/agents/notion-scruff)

### Sources

1. [Notion's Token Town: 5 Rebuilds, 100+ Tools (Latent Space)](https://latent.space/p/notion)
   - Podcast · Direct participant · Evidence
   - Original URL: <https://latent.space/p/notion>
   - Published: 2026-04-14 · Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-1/content.md>
2. [How we built security into Custom Agents](https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.notion.com/en-gb/blog/how-we-built-security-into-custom-agents>
   - Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-custom-agents-source-2/content.md>
3. [Meet Scruff, Security’s New AI Teammate](https://www.notion.com/blog/meet-scruff-securitys-new-ai-teammate)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.notion.com/blog/meet-scruff-securitys-new-ai-teammate>
   - Published: 2026-01-12 · Accessed: 2026-09-17 · Last verified: 2026-09-17
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/notion-scruff-source/content.md>

## Plaid — Internal MCP server

Plaid's internal Model Context Protocol server is supporting infrastructure rather than an agent. The central server connects engineers' AI tools to internal systems such as Jira, application logs, and data schemas that third-party MCP servers cannot reach.

- Company: [Plaid](https://internal-agents.com/organizations/plaid)
- Collection: Infrastructure
- Approach type: Component
- Deployment stage: Scaled
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2025
- Work: Coding
- Entry reviewed: 2026-09-16

Page: https://internal-agents.com/agents/plaid-internal-mcp-server

### Purpose

#### Summary

Plaid's internal Model Context Protocol server is supporting infrastructure rather than an agent. The central server connects engineers' AI tools to internal systems such as Jira, application logs, and data schemas that third-party MCP servers cannot reach.

Fact · Reported · High confidence · `plaid-internal-mcp-server--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, lines 38–45, 65–89

### Capabilities and architecture

#### Harness

Central internal MCP server that fronts vendor AI clients

Fact · Reported · High confidence · `plaid-internal-mcp-server--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md)

#### Tool access

More than 20 tools and several internal services (Jira, logs, schemas)

Fact · Reported · High confidence · `plaid-internal-mcp-server--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md)

#### Credentials

Behind Plaid's identity-aware proxy and centralized authorization

Fact · Reported · High confidence · `plaid-internal-mcp-server--architecture-credentials`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md)

- **Model:** not-applicable — The server connects AI clients to internal data and runs no model itself; the capture describes no model component.
- **Sandbox:** unreported — The capture describes the identity-aware proxy and managed-device checks as access control, not an execution isolation boundary.
- **Knowledge:** unreported — Documentation and internal data are integrated as tools; the capture describes no separate knowledge or context store.
- **Context management:** unreported — The capture describes delivering context to AI clients, not how context is maintained or organized over time.
- **Interfaces:** unreported — The capture names AI clients such as Claude Code and Cursor that connect as MCP clients, not a human-facing surface of the server.

### Documented uses

Documented use example: AI-client request → authenticated internal tool access and context.

#### Connect an AI client

An engineer's AI tool such as Claude Code or Cursor reaches internal data through the one central internal MCP server instead of a locally managed arrangement of MCP servers

Fact · Reported · High confidence · `plaid-internal-mcp-server--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, lines 38–40, 67

#### Authenticate the engineer

CLI-based authentication uses the Device Authorization Grant flow with DPoP and short-lived bearer tokens through a locally running proxy

Fact · Reported · High confidence · `plaid-internal-mcp-server--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, line 59

#### Authorize the request

The identity-aware proxy validates a Plaid managed device and identity-provider login; a signed identity token is parsed and the centralized authorization server checks the employee's access to the target service

Fact · Reported · High confidence · `plaid-internal-mcp-server--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, lines 57, 59

#### Return internal context

The tool returns data from internal systems such as Jira, application logs, and data schemas to the AI client

Fact · Reported · High confidence · `plaid-internal-mcp-server--primitives-4`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, lines 14–16

### Access and controls

See the credential and access boundaries in architecture (`plaid-internal-mcp-server--architecture-credentials`). Scoped human-review assessments for individual uses remain in research details.

### Reliability and validation

#### Enforce data policy at call time

Tool-level controls enforce the LLM data access policy when a tool is called, and the centralized design adds caller identity inspection and audit logging

Fact · Reported · High confidence · `plaid-internal-mcp-server--primitives-3`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, lines 61–63, 82, 87

### Adoption and operating evidence

Observation: Adoption output · Estimate · Claude Code and Cursor usage among Plaid engineers and agents relying on the internal MCP server

#### Headline claim

Dozens of agents rely on the internal MCP server; Claude Code and Cursor are used by over 80% of engineers

Metric · Reported · Medium confidence · `plaid-internal-mcp-server--headline-metric`

Confidence reason: Plaid reported the figures in its own engineering blog without independent verification.

Qualifications:

- Reported by: Plaid
- Scope: Claude Code and Cursor adoption among Plaid engineers; separate from internal MCP server adoption
- Denominator: Plaid engineers for the AI-client usage share
- Observation date: 2025

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, lines 38, 89

Observation: Adoption output · Estimate · Claude Code and Cursor adoption among Plaid engineers

#### Key observation

Claude Code and Cursor are used by over 80% of Plaid engineers; the source does not report internal MCP server adoption share

Metric · Reported · Medium confidence · `plaid-internal-mcp-server--key-metrics-0`

Confidence reason: Plaid reported the figures in its own engineering blog.

Qualifications:

- Reported by: Plaid
- Scope: Claude Code and Cursor adoption among Plaid engineers; separate from internal MCP server adoption
- Denominator: Plaid engineers for the AI-client usage share

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, lines 38, 89

Observation: Adoption output · Estimate · Tool calls through and agents built on the internal MCP server across engineering, product, and support

#### Key observation

Thousands of tool calls and dozens of agents built on it

Metric · Reported · Medium confidence · `plaid-internal-mcp-server--key-metrics-1`

Confidence reason: Plaid reported the figures in its own engineering blog.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Plaid
- Scope: Tool calls and agents relying on the internal MCP server across engineering, product, and support

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md) · Preserved content.md, lines 89

The capture offers design rationale, such as the view that existing access-control frameworks can be naturally replicated, but states no evaluated transferable lesson.

### Research details and scoped use assessments

#### Operating model assessment

Unclassified for engineer request → internal tool access; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `plaid-internal-mcp-server--operating-models-0`

Confidence reason: The MCP server is a tool-access layer, not a workflow with a single human-attention boundary.

Qualifications:

- Observation date: 2025

Evidence:

- Supports · [1] [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md)

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported
- **observations:** Reported
- **lessons:** Unreported — The capture offers design rationale, such as the view that existing access-control frameworks can be naturally replicated, but states no evaluated transferable lesson.

### Sources

1. [The Plaid internal MCP server](https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb)
   - Engineering blog · First party · Evidence
   - Original URL: <https://engineering.plaid.com/the-plaid-internal-mcp-server-8eff08bb6bdb>
   - Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/plaid-internal-mcp-server-source-1/content.md>

## Shopify — Aquifer

The Aquifer agent platform (session/harness/sandbox split) powers River, a Slack-native coding agent, plus PR-review and headless Vanilla agent profiles; other teams have requested research, migration, and compliance-scan agents.

- Company: [Shopify](https://internal-agents.com/organizations/shopify)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Scaled
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2026
- Work: Coding, Code review
- Interfaces: Slack
- Entry reviewed: 2026-09-17

Page: https://internal-agents.com/agents/shopify-internal-agents

### Purpose

#### Summary

The Aquifer agent platform (session/harness/sandbox split) powers River, a Slack-native coding agent, plus PR-review and headless Vanilla agent profiles; other teams have requested research, migration, and compliance-scan agents.

Fact · Reported · High confidence · `shopify-internal-agents--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

### Capabilities and architecture

#### Sandbox

An execution environment (filesystem, shell, repo, build/test) separated from the harness; Shopify credits the brain-and-hands framing to Anthropic

Fact · Reported · High confidence · `shopify-internal-agents--architecture-sandbox`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### Harness

Session (durable; Postgres append-only event log) + Harness (cheap agent loop) + Cell (ephemeral Go runtime); cells die, sessions persist

Fact · Reported · High confidence · `shopify-internal-agents--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### Model

The design lets Shopify change the model without changing the sandbox

Fact · Reported · High confidence · `shopify-internal-agents--architecture-model`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### Interfaces

slack

Fact · Reported · High confidence · `shopify-internal-agents--architecture-interfaces`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### Tool access

Repo + tests + data warehouse + production traces + PR creation; a gateway credentials proxy

Fact · Reported · High confidence · `shopify-internal-agents--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### Knowledge

Monorepo 'World' (code + skills + conventions + intent docs + runbooks + AGENTS.md); Nix reproducible envs; Slack-transcript corpus mining

Fact · Reported · High confidence · `shopify-internal-agents--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### Credentials

Gateway credentials proxy and per-profile sandbox policies; the source does not name an SSO mechanism.

Fact · Reported · High confidence · `shopify-internal-agents--architecture-credentials`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### Context management

Skills loaded on-demand as files, updatable per session; session survival across cell/sandbox/machine death

Fact · Reported · High confidence · `shopify-internal-agents--architecture-context-mgmt`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### Session/Harness/Sandbox split

Durable identity, disposable loop, isolated execution; swap any layer independently

Fact · Reported · High confidence · `shopify-internal-agents--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

#### River (public-by-default)

River works in shared Slack threads whose history can be searched and used to update skills

Fact · Reported · High confidence · `shopify-internal-agents--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 71-79,87

#### Skills as files

Written-down knowledge mined from successful patterns and public transcripts

Fact · Reported · High confidence · `shopify-internal-agents--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md)

### Documented uses

Documented use example: One River thread: public @river request through worked findings, a human redirect, and a searchable transcript.

#### Take a request in a public channel

A person @-mentions River with a question or task in a public Slack channel; River works only in the open

Fact · Reported · High confidence · `shopify-internal-agents--primitives-3`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 69, 75

#### Work the problem in the open

River reads files, runs queries and tests, and posts partial findings to the thread

Fact · Reported · High confidence · `shopify-internal-agents--primitives-4`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 65, 76

#### Absorb a mid-thread redirect

A second person drops in with a constraint or redirect; River incorporates it without losing the conversation and continues

Fact · Reported · High confidence · `shopify-internal-agents--primitives-5`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 77-78

#### Leave a reusable transcript

The finished thread stays searchable so the next person starts from it, and observed patterns feed back into skills

Fact · Reported · High confidence · `shopify-internal-agents--primitives-6`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 71, 79

### Access and controls

See the credential and access boundaries in architecture (`shopify-internal-agents--architecture-credentials`). Scoped human-review assessments for individual uses remain in research details.

### Reliability and validation

Tests River runs are part of its work, and CI improvements serve human developers; no check of River's own output is documented.

### Adoption and operating evidence

Observation: Adoption output · Reported measurement · Company-wide merged pull requests coauthored by River

#### Headline claim

River use example: 1 in 8 merged PRs company-wide coauthored by River

Metric · Reported · High confidence · `shopify-internal-agents--headline-metric`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- Reported by: Shopify
- Scope: Merged River-coauthored PRs across Shopify
- Denominator: All Shopify merged pull requests

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 12

Observation: Adoption output · Reported measurement · River sessions and distinct Slack channels over a recent 30-day period

#### Key observation

River: 59,918 sessions / 30 days across 5,170 Slack channels

Metric · Reported · High confidence · `shopify-internal-agents--key-metrics-0`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Shopify
- Scope: River sessions and distinct Slack channels in a recent 30-day period
- Method: river_sessions domain table, written by River every session

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 83

Observation: Adoption output · Reported measurement · River-coauthored merged pull requests and their company-wide share in a recent 30-day period

#### Key observation

3,536 River-coauthored PRs merged; 1 in 8 merged PRs company-wide

Metric · Reported · High confidence · `shopify-internal-agents--key-metrics-1`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- Reported by: Shopify
- Scope: River-coauthored merged PRs in the recent 30-day period; company-wide PR share
- Denominator: All Shopify merged pull requests for the one-in-eight share
- Method: river_sessions domain table for reported session-linked counts

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 12, 83

Observation: Cost latency · Reported measurement · Median River session duration and median tool calls per session

#### Key observation

River: Median session 19 min; median 50 tool calls/session

Metric · Reported · High confidence · `shopify-internal-agents--key-metrics-2`

Confidence reason: A linked first-party source states the claim.

Qualifications:

- The source does not report the denominator of this figure.
- Reported by: Shopify
- Scope: Median River session duration and tool calls per session

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 65

### Lessons

#### Lesson

Shopify says its monorepo, reproducible environments, written skills, and CI improvements benefited both engineers and agents.

Opinion · Reported · Medium confidence · `shopify-internal-agents--lessons-learned-0`

Confidence reason: The retrospective identifies these as existing human-development priorities made more urgent by agents; it does not quantify each contribution.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 24-25,38-47

#### Lesson

Shopify uses shared River threads and the run corpus to update skills, prompts, and defaults for later work.

Fact · Reported · Medium confidence · `shopify-internal-agents--lessons-learned-1`

Confidence reason: The article describes mining earlier runs and preserving searchable threads, replacing the vague claim that private agents cannot teach anyone.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 57,71,79,87

#### Lesson

Aquifer stores session identity and its event log in Postgres so a fresh worker can continue the same conversation.

Fact · Reported · Medium confidence · `shopify-internal-agents--lessons-learned-2`

Confidence reason: The session and lifecycle sections explicitly separate the durable log from disposable cells; they do not claim every external side effect is automatically recoverable.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 116-134

#### Lesson

Shopify adds agent variants as profiles containing prompts, skills, extensions, sandbox policy, and model defaults on Aquifer.

Fact · Reported · Medium confidence · `shopify-internal-agents--lessons-learned-3`

Confidence reason: The profile section enumerates the bundle contents and explains how a new bundle becomes an internal agent product.

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 142

### Research details and scoped use assessments

#### Operating model assessment

Unclassified for River coding request → River-opened, River-coauthored pull request; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `shopify-internal-agents--operating-models-0`

Confidence reason: The source documents River-opened, River-coauthored pull requests that are later merged, but no passage establishes that a person reviews the work product on each run; PR review runs as a separate Aquifer profile and is described as often having no human in the loop.

Qualifications:

- Observation date: 2026

Evidence:

- Supports · [1] [Under the River](https://shopify.engineering/under-the-river) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md) · Preserved content.md, lines 65, 83, 140

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported — The article documents exactly one representative workflow, the well-going public River thread; PR review and the headless Vanilla profile are separate Aquifer profiles, not steps of this run.
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Unreported — Tests River runs are part of its work, and CI improvements serve human developers; no check of River's own output is documented.
- **observations:** Reported
- **lessons:** Reported

### Related reading

- Uses this infrastructure: [Shopify — River](https://internal-agents.com/agents/shopify-river)

### Sources

1. [Under the River](https://shopify.engineering/under-the-river)
   - Engineering blog · First party · Evidence
   - Original URL: <https://shopify.engineering/under-the-river>
   - Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/shopify-internal-agents-source-1/content.md>

## Y Combinator — Internal agent infrastructure

Y Combinator built its own agent harness and infrastructure, led by partner Pete Koomen. It is an agent loop over a shared tool registry, now over 350 tools, including one that runs read-only SQL against YC's single Postgres database. Agent conversations are broadcast to a Slack channel.

- Company: [Y Combinator](https://internal-agents.com/organizations/y-combinator)
- Collection: Infrastructure
- Approach type: Platform
- Deployment stage: Deployed
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2026
- Work: Coding, Ops
- Interfaces: Slack
- Entry reviewed: 2026-09-16

Page: https://internal-agents.com/agents/ycombinator-agent-infra

### Purpose

#### Summary

Y Combinator built its own agent harness and infrastructure, led by partner Pete Koomen. It is an agent loop over a shared tool registry, now over 350 tools, including one that runs read-only SQL against YC's single Postgres database. Agent conversations are broadcast to a Slack channel.

Fact · Reported · Low confidence · `ycombinator-agent-infra--summary`

Confidence reason: Only limited public evidence supports this claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md)

### Capabilities and architecture

#### Harness

Own harnesses built from the ground up for internal AI use

Fact · Reported · Low confidence · `ycombinator-agent-infra--architecture-harness`

Confidence reason: Only limited public evidence supports this claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md)

#### Tool access

Shared registry of more than 350 YC-specific tools, including read-only SQL access

Fact · Reported · High confidence · `ycombinator-agent-infra--architecture-tool-access`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 124, 154

#### Interfaces

slack

Fact · Reported · High confidence · `ycombinator-agent-infra--architecture-interfaces`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 190–192

#### Knowledge

A shared database and common context layer expose YC's internal organizational context

Fact · Reported · High confidence · `ycombinator-agent-infra--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, line 154

#### Context management

An agent loop uses a shared tool registry, skill registry, and model router

Fact · Reported · High confidence · `ycombinator-agent-infra--architecture-context-mgmt`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 154–158

- **Model:** unreported — A model router is mentioned, but no model is identified for the representative workflow.
- **Sandbox:** unreported — The legacy unknown placeholder is retained in research details; isolation is not documented.
- **Credentials:** unreported — The capture describes read-only SQL access but not credential handling.

### Documented uses

Documented use example: Finance-team natural-language question → read-only database research.

#### Finance question

A finance team member asks a real operational question in natural language instead of writing SQL or requesting purpose-built software

Fact · Reported · High confidence · `ycombinator-agent-infra--primitives-0`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 120–124

#### Read-only data access

The general agent loop selects tools from the shared registry and can run read-only SQL against YC's database and read model files

Fact · Reported · High confidence · `ycombinator-agent-infra--primitives-1`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 124, 154

#### Employee-visible conversation

Agent conversations are broadcast to an internal Slack channel where employees can inspect and learn from them

Fact · Reported · High confidence · `ycombinator-agent-infra--primitives-2`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 190–192

### Access and controls

No separate access-control assessment is recorded.

### Reliability and validation

#### Nightly skill review

A general agent reads employee-agent conversations nightly to identify failures and missing context that could improve shared skills

Fact · Reported · High confidence · `ycombinator-agent-infra--primitives-3`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 162–166

### Adoption and operating evidence

The reviewed capture reports capabilities and examples but no measured outcome for the named finance workflow.

### Lessons

#### Lesson

YC makes a shared database and tool registry available to its internal agents, with teams adding tools for their own work.

Fact · Reported · Low confidence · `ycombinator-agent-infra--lessons-learned-0`

Confidence reason: The interview describes the common schema and expanding registry, including finance and scheduling examples; the operating-system metaphor adds no implementation detail.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 154

#### Lesson

YC began its own agent harness while working with its finance team on top of existing internal software.

Fact · Reported · Low confidence · `ycombinator-agent-infra--lessons-learned-1`

Confidence reason: The interview recounts that origin and YC's longstanding custom software. It does not compare the build with a hosted-agent alternative.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md) · Preserved content.md, lines 120

### Research details and scoped use assessments

#### Sandbox

unknown

Inference · Catalog judgment · Medium confidence · `ycombinator-agent-infra--architecture-sandbox`

Confidence reason: The preserved sources do not document an execution sandbox; unknown does not mean absent.

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md)

#### Operating model assessment

Unclassified for internal request → agent-assisted organizational work; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `ycombinator-agent-infra--operating-models-0`

Confidence reason: The source describes internal agent infrastructure but not a sufficiently specific human attention boundary.

Qualifications:

- Observation date: 2026

Evidence:

- Supports · [1] [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md)

### Question coverage and scope

- **purpose:** Reported
- **workflow:** Reported
- **human involvement:** Not applicable — There is no single platform-wide agent review boundary. Scoped downstream operating-model claims remain in research details; access controls are described separately.
- **implementation:** Reported
- **validation:** Reported
- **observations:** Unreported — The reviewed capture reports capabilities and examples but no measured outcome for the named finance workflow.
- **lessons:** Reported

### Related reading

- Uses this infrastructure: [Y Combinator — Internal operations agent](https://internal-agents.com/agents/ycombinator-operations)

### Sources

1. [Inside YC's AI Playbook (Lightcone podcast, with Pete Koomen)](https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook)
   - Podcast · First party · Evidence
   - Original URL: <https://www.ycombinator.com/library/Qh-inside-yc-s-ai-playbook>
   - Accessed: 2026-08-12 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/ycombinator-agent-infra-source-1/content.md>
