Databricks · Agent system
coSTAR and internal engineering agents
Databricks runs internal engineering agents for work such as on-call support and automated code review. coSTAR ships and tests them, using LLM judges as the test suite and a coding assistant to refine the agent until the judges pass. Databricks' open-source Omnigent is a separate product.
- Approach type
- Agent system
- Work
- Coding, Code review, On-call
- Human involvement
- Unknown
- Invocation
- Interactive, Background
- Deployment stage
- Scaled
- Evidence strength
- Detailed primary
- Entry reviewed
Where people stay involved
- internal engineering workflows → agent-produced changesUnknown · Level unknown
Unclassified for internal engineering workflows → agent-produced changes; human attention boundary: unknown.
- Observation date
- 2025
Implementation details
coSTAR framework for shipping and testing internal agents
Private benchmark built from the Databricks multi-million line codebase
Reported results and limitations
The catalog records what the sources report, with the scope and the denominator of every figure. A qualification below limits the figure it sits under.
Reported outcomes and statements
Internal agents serve as daily coding drivers on the Databricks codebase
Private benchmark built from a multi-million line codebase
Sources and research details
Citations link to the original publisher. Each source also keeps a preserved copy in the repository, so a changed or removed page stays checkable.
- coSTAR: how we ship AI agents at Databricks fasthttps://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things
- Benchmarking coding agents on a multi-million line codebasehttps://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase
Research details for every claim on this page
- Summary
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- High
- Confidence reason
- A linked first-party source states the claim.
- Harness
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- High
- Confidence reason
- A linked first-party source states the claim.
- Knowledge
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- High
- Confidence reason
- A linked first-party source states the claim.
- Key observation
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- Databricks described internal use in its own engineering blog.
- Key observation
- Statement type
- Fact
- Provenance
- Reported
- Confidence
- Medium
- Confidence reason
- Databricks described the benchmark in its own engineering blog.
- Operating model assessment
- Statement type
- Inference
- Provenance
- Catalog judgment
- Confidence
- Unverified
- Confidence reason
- The record covers several internal engineering agents with different workflows, so no single human-attention boundary applies.
- Observation date
- 2025