
Agentic software factory
An agent system built around OpenAI's internal Codex. A person defines the outcome; Codex implements the change and babysits CI until green; domain-specialist agents review each change behind risk-tiered routing; and a per-change deploy agent handholds approved changes to production and builds its own monitoring dashboards.
- Company
- OpenAI
- Approach type
- Agent system
- Work
- Coding, Code review, CI triage, Ops
- Human involvement
- Human in loop
- Invocation
- Interactive, Background, Scheduled, Event driven
- Interfaces
- Desktop, Cli, Slack, Github, Skill
- Deployment stage
- Scaled
- Evidence strength
- Mixed
- Entry reviewed
How it works
The workflow the sources report for this implementation.
Multiple agents, each configured as a domain specialist, review every change; the article compares this to a review by a domain expert from each relevant infrastructure team
Changes are classified by risk; high-risk changes can trigger more agent reviews or a mandated human review, while opted-in low-risk areas use an agent that auto-approves pull requests
After a human approves production, an assigned agent handholds the change to full rollout; it decides which signals mean success or failure, builds its own monitoring dashboard, and watches production signals
A perf harness sends problematic pull requests to the Synthetics A/B framework to evaluate performance implications
Agents sift through alerts and dashboards, de-duplicate signals, identify real latency regressions, root-cause them, and propose fixes
Where people stay involved
Each scope pairs its normal attention boundary with supporting evidence. See thesupervision definitions for the level mapping and limits.
low-risk pull request -> merge in opted-in codebase areas
Exception only · Level 5code change -> production rollout on the general path
Work product review · Level 3production alert -> proposed performance fix
Unknown · Level unknown
Catalog interpretation
Level 5 for low-risk pull request -> merge in opted-in codebase areas; human attention boundary: exception-only.
Observed in September 2026
Catalog interpretation
Level 3 for code change -> production rollout on the general path; human attention boundary: work-product-review.
Observed in September 2026
Catalog interpretation
Unclassified for production alert -> proposed performance fix; human attention boundary: unknown.
Observed in September 2026
Implementation details
- Model
- unknown
- Harness
- Internal Codex, described as much more advanced than the external product because it is plugged into almost every OpenAI system; ChatGPT Work runs on the same harness
- Sandbox
- unknown
- Tool access
- Git repositories and GitHub, Slack, Notion, Databricks, Datadog, and internal logs and data sources
- Knowledge
- OpenAI moved its documentation inside the source code; internal Codex skills, some maintained by Codex itself; new engineers are directed to ask Codex during onboarding
- Context management
- The /goal setting lets an agent work until a goal is complete; threads run for days and spin off other agents
- Credentials
- N/A
- Interfaces
- desktop, cli, slack, github, skill
Reported results and limitations
The catalog records what the sources report, with the scope and the denominator of every figure. A qualification below limits the figure it sits under.
“Almost all OpenAI employees used Codex and ChatGPT Work weekly as of the September 2026 report (self-reported).”
- Reported by
- OpenAI
- Scope
- Weekly Codex and ChatGPT Work use across OpenAI employees
- Method
- Internal token-usage tracking by department, chart sourced to OpenAI
The source does not report the denominator of this figure.
Observed in September 2026
Non-engineering orgs such as finance, recruitment, and legal went from about 0% to 90% Codex usage within a four-month period (self-reported)
- Reported by
- OpenAI
- Scope
- Weekly Codex use in non-engineering orgs such as finance, recruitment, and legal
The source does not report the denominator of this figure.
Observed in September 2026
Some build-test-deploy systems saw about a 10x load increase within roughly six months (self-reported)
- Reported by
- OpenAI
- Scope
- Load on build-test-deploy pipeline systems
The source does not report the denominator of this figure.
Observed in September 2026
Lessons and interpretation
Reported opinion
Longer-running /goal tasks drove adoption; one long-running agent that spins off other agents reduces the surface a person manages
Reported opinion
Role-specific and team-specific plugins spread adoption beyond a generic coding agent
Sources and research details
Citations link to the original publisher. Each source also keeps a preserved copy in the repository, so a changed or removed page stays checkable.
- Inside OpenAI's agentic software factoryhttps://newsletter.pragmaticengineer.com/p/openai-software-factory
- The most important OpenAI announcement you probably missed at DevDay 2025https://venturebeat.com/infrastructure/the-most-important-openai-announcement-you-probably-missed-at-devday-2025
- Harness engineering: leveraging Codex in an agent-first worldhttps://openai.com/index/harness-engineering/
- Harness engineering: Leveraging Codex in an agent-first world (discussion)https://news.ycombinator.com/item?id=48416264
- zbrock: the Codex app began as an internal prototypehttps://news.ycombinator.com/item?id=48435137
- zbrock: many internal teams adopted the same practiceshttps://news.ycombinator.com/item?id=48435213
- Introducing the Agents APIhttps://openai.com/index/introducing-the-agents-api/
- ChatGPT is now a partner for your most ambitious workhttps://openai.com/index/chatgpt-for-your-most-ambitious-work/
- Transcript: 'How OpenAI's Codex Team Uses Their Coding Agent'https://every.to/podcast/transcript-how-openai-s-codex-team-uses-their-coding-agent
Research details for every claim on this page
- Summary
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A linked participant or independent source reports the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 138-186 (pipeline steps 1-8)
- ContextualizesHarness engineering: leveraging Codex in an agent-first worldPreserved content.md, lines 46-48 (one OpenAI team's repository runs review predominantly agent-to-agent; a single team's account, not the company-wide pipeline)
- Headline claim
- Statement type
- Metric
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- Company usage claim relayed by an independent reporter; no method or denominator published. OpenAI's own 2026-09-10 launch post states nearly 100% of teams inside OpenAI, including finance and sales, use ChatGPT Work and Codex (a teams denominator, not employees).
- Reported by
- OpenAI
- Scope
- Weekly Codex and ChatGPT Work use across OpenAI employees
- Denominator
- Not reported
- Method
- Internal token-usage tracking by department, chart sourced to OpenAI
- Observation date
- 2026-09
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 46
- SupportsChatGPT is now a partner for your most ambitious workPreserved content.md, line 34 (OpenAI's own launch post: nearly 100% of teams inside OpenAI, including finance and sales, use ChatGPT Work and Codex; teams denominator, OpenAI self-report)
- ContextualizesTranscript: 'How OpenAI's Codex Team Uses Their Coding Agent'Preserved content.md, line 486 (Codex lead, February 2026: 'almost everyone technical at the company uses Codex')
- Sandbox
- Statement type
- Inference
- Provenance
- Catalog judgment
- Confidence
- Medium
- Confidence reason
- The article does not document an execution sandbox; unknown does not mean absent.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 138-186 describe the pipeline without documenting an execution sandbox
- Harness
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- High
- Confidence reason
- A linked first-party source states the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 48 (internal Codex is a lot more advanced than its external counterpart; ChatGPT Work is powered by the Codex harness)
- Supportszbrock: the Codex app began as an internal prototypePreserved content.md, line 12 (harness-engineering co-author: 'It was an internal prototype that looked very much like the current Codex app')
- SupportsIntroducing the Agents APIPreserved content.md, lines 12-14 (OpenAI scaled Codex and ChatGPT for Work on the same harness; the Agents API exposes 'that same harness and infrastructure that powers Codex')
- Model
- Statement type
- Inference
- Provenance
- Catalog judgment
- Confidence
- Medium
- Confidence reason
- The article names no underlying models for the internal harness; unknown does not mean absent.
- SupportsInside OpenAI's agentic software factoryThe preserved article names no underlying models for the internal Codex harness
- Interfaces
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A linked participant or independent source reports the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 48, 80, and 140-145 (Codex app and CLI, GitHub, Slack, internal skills)
- Tool access
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A linked participant or independent source reports the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 140-145
- Knowledge
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- High
- Confidence reason
- A linked first-party source states the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 140-147 (documentation inside the source code; internal skills; onboarding)
- SupportsHarness engineering: leveraging Codex in an agent-first worldPreserved content.md, lines 30, 73, and 117 (structured docs/ directory as the system of record; a roughly 100-line AGENTS.md serving as a map; a recurring doc-gardening agent; the initial AGENTS.md was written by Codex)
- SupportsTranscript: 'How OpenAI's Codex Team Uses Their Coding Agent'Preserved content.md, lines 236, 258, and 274 (automations that keep PRs mergeable by resolving merge conflicts, a random-file bug-finder that runs multiple times a day, and a bot that quietly fixes bugs in recently merged PRs)
- Context management
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A linked participant or independent source reports the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 52-56 (/goal setting; long-running threads)
- ContextualizesHarness engineering: leveraging Codex in an agent-first worldPreserved content.md, line 58 (single Codex runs 'upwards of six hours' while humans sleep; hours, not the days the article reports)
- Supporting component
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- High
- Confidence reason
- A linked first-party source states the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 155-157 (domain-specialist review agents)
- SupportsThe most important OpenAI announcement you probably missed at DevDay 2025Preserved content.md, line 44 (DevDay 2025, October 2025: nearly every pull request at OpenAI is reviewed by Codex, catching hundreds of issues daily; OpenAI's own figure relayed by press)
- SupportsHarness engineering: leveraging Codex in an agent-first worldPreserved content.md, lines 46 and 48 (Codex requests additional specific agent reviews locally and in the cloud until all agent reviewers are satisfied; 'almost all review effort' handled agent-to-agent; one repository)
- Supporting component
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A linked participant or independent source reports the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 159 (risk classification and auto-approve)
- Supporting component
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A linked participant or independent source reports the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 165-175 (agentic deploy)
- Contextualizeszbrock: many internal teams adopted the same practicesPreserved content.md, lines 12 and 16 (harness-engineering co-author: many internal teams adopted the same practices; some run centralized agent-mediated integration queues and local Codex threads that monitor CI; no self-built deploy dashboards described)
- Supporting component
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A linked participant or independent source reports the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 153 (perf harness sends problematic PRs to the Synthetics A/B framework)
- Supporting component
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A linked participant or independent source reports the claim.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 186 (Perf Factory)
- Key observation
- Statement type
- Metric
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- Company adoption figures relayed by an independent reporter; the four-month window carries no calendar dates. OpenAI's launch post corroborates finance-org use (month-end close reduced from days to hours) but not the 0%-to-90% trajectory.
- Reported by
- OpenAI
- Scope
- Weekly Codex use in non-engineering orgs such as finance, recruitment, and legal
- Denominator
- Not reported
- Method
- Not reported
- Observation date
- 2026-09
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 46 (non-engineering orgs went from ~0% to 90% usage in a four-month period)
- ContextualizesChatGPT is now a partner for your most ambitious workPreserved content.md, line 37 (OpenAI's finance org uses ChatGPT Work for month-end close and forecasting, reduced from days to hours; corroborates non-engineering adoption, not the 0%-to-90% trajectory)
- Key observation
- Statement type
- Metric
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- A VP of Engineering load estimate relayed by an independent reporter.
- Reported by
- OpenAI
- Scope
- Load on build-test-deploy pipeline systems
- Denominator
- Not reported
- Method
- Not reported
- Observation date
- 2026-09
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 92 (roughly 10x load increase on some systems in about six months)
- Lesson
- Statement type
- Opinion
- Provenance
- Reported
- Confidence
- High
- Confidence reason
- Stated explanation by the Codex desktop lead in the reported interviews.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 52-56 (Andrew Ambrosino on long-running tasks and agents spinning off agents)
- Lesson
- Statement type
- Opinion
- Provenance
- Reported
- Confidence
- High
- Confidence reason
- Stated explanation by the Codex desktop lead in the reported interviews.
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 64 (role-specific and team-specific plugins are created and distributed)
- Operating model assessment
- Statement type
- Inference
- Provenance
- Catalog judgment
- Confidence
- Medium
- Confidence reason
- The article documents agent auto-approve for low-risk pull requests in opted-in areas and risk classification that routes high-risk changes to stricter review; the run-time exception path is not described. The harness-engineering post corroborates agent automerge of low-friction pull requests in one repository but describes no risk tiering, so the routing mechanism itself stays single-sourced.
- Observation date
- 2026-09
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 159 (risk classification; low-risk areas can opt in to an agent that auto-approves pull requests)
- ContextualizesHarness engineering: leveraging Codex in an agent-first worldPreserved content.md, line 199 (background Codex tasks 'reviewed in under a minute and automerged' in one repository; no risk-tiered routing described)
- Operating model assessment
- Statement type
- Inference
- Provenance
- Catalog judgment
- Confidence
- High
- Confidence reason
- Pipeline step 6 states that a human approves a change to go to production before the deploy agent takes over.
- Observation date
- 2026-09
- SupportsInside OpenAI's agentic software factoryPreserved content.md, lines 165-175 (after a human approves production, an assigned agent handholds the rollout)
- Operating model assessment
- Statement type
- Inference
- Provenance
- Catalog judgment
- Confidence
- Unverified
- Confidence reason
- The article documents Perf Factory only up to agents that de-duplicate signals, identify real latency regressions, root-cause them, and propose fixes; whether a human reviews, applies, or lands those proposals is not stated.
- Observation date
- 2026-09
- SupportsInside OpenAI's agentic software factoryPreserved content.md, line 186 (Perf Factory proposes fixes)