Source: https://internal-agents.com/agents/databricks-costar

# Databricks — coSTAR and internal engineering agents

Databricks runs internal engineering agents for work such as on-call support and automated code review. coSTAR ships and tests them, using LLM judges as the test suite and a coding assistant to refine the agent until the judges pass. Databricks' open-source Omnigent is a separate product.

- Approach type: Agent system
- Deployment stage: Scaled
- Autonomy: Unknown
- Evidence strength: Detailed primary
- Status: Internal
- First reported year: 2025
- Work: Coding, Code review, On-call
- Invocation: Interactive, Background
- Entry reviewed: 2026-09-15

## Where people stay involved

- **internal engineering workflows → agent-produced changes** — Unknown · Level unknown

## Overview

### Summary

Databricks runs internal engineering agents for work such as on-call support and automated code review. coSTAR ships and tests them, using LLM judges as the test suite and a coding assistant to refine the agent until the judges pass. Databricks' open-source Omnigent is a separate product.

Fact · Reported · High confidence · `databricks-costar--summary`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md)
- Supports · [2] [Benchmarking coding agents on a multi-million line codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-2/content.md)

## Supervision evidence

### Operating model assessment

Unclassified for internal engineering workflows → agent-produced changes; human attention boundary: unknown.

Inference · Catalog judgment · Unverified confidence · `databricks-costar--operating-models-0`

Confidence reason: The record covers several internal engineering agents with different workflows, so no single human-attention boundary applies.

Qualifications:

- Observation date: 2025

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md)

## Implementation details

### Harness

coSTAR framework for shipping and testing internal agents

Fact · Reported · High confidence · `databricks-costar--architecture-harness`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md)

### Knowledge

Private benchmark built from the Databricks multi-million line codebase

Fact · Reported · High confidence · `databricks-costar--architecture-knowledge`

Confidence reason: A linked first-party source states the claim.

Evidence:

- Supports · [2] [Benchmarking coding agents on a multi-million line codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-2/content.md)

## Reported outcomes and statements

### Key observation

Internal agents serve as daily coding drivers on the Databricks codebase

Fact · Reported · Medium confidence · `databricks-costar--key-metrics-0`

Confidence reason: Databricks described internal use in its own engineering blog.

Evidence:

- Supports · [1] [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md)

### Key observation

Private benchmark built from a multi-million line codebase

Fact · Reported · Medium confidence · `databricks-costar--key-metrics-1`

Confidence reason: Databricks described the benchmark in its own engineering blog.

Evidence:

- Supports · [2] [Benchmarking coding agents on a multi-million line codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase) · [Preserved copy](https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-2/content.md)

## Sources

1. [coSTAR: how we ship AI agents at Databricks fast](https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-things>
   - Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-1/content.md>
2. [Benchmarking coding agents on a multi-million line codebase](https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase)
   - Engineering blog · First party · Evidence
   - Original URL: <https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase>
   - Accessed: 2026-08-13 · Last verified: 2026-08-31
   - Preserved copy in the repository: <https://github.com/steel-experiments/internal-agents-map/blob/main/archive/sources/databricks-costar-source-2/content.md>
