← All implementations

Databricks · Agent system

coSTAR and internal engineering agents

Databricks runs internal engineering agents for work such as on-call support and automated code review. coSTAR ships and tests them, using LLM judges as the test suite and a coding assistant to refine the agent until the judges pass. Databricks' open-source Omnigent is a separate product.

1 Supports2 Supports

Approach type
Agent system
Work
Coding, Code review, On-call
Human involvement
Unknown
Invocation
Interactive, Background
Deployment stage
Scaled
Evidence strength
Detailed primary
Entry reviewed

Where people stay involved

  • internal engineering workflows → agent-produced changesUnknown · Level unknown

Unclassified for internal engineering workflows → agent-produced changes; human attention boundary: unknown.

1

Observation date
2025

Implementation details

Harness

coSTAR framework for shipping and testing internal agents

1

Knowledge

Private benchmark built from the Databricks multi-million line codebase

2

Reported results and limitations

The catalog records what the sources report, with the scope and the denominator of every figure. A qualification below limits the figure it sits under.

Reported outcomes and statements

Key observationFact

Internal agents serve as daily coding drivers on the Databricks codebase

1

Key observationFact

Private benchmark built from a multi-million line codebase

2

Sources and research details

Citations link to the original publisher. Each source also keeps a preserved copy in the repository, so a changed or removed page stays checkable.

  1. coSTAR: how we ship AI agents at Databricks fasthttps://www.databricks.com/blog/costar-how-we-ship-ai-agents-databricks-fast-without-breaking-thingsEngineering blog · First party · Last source verification: 2026-08-31
  2. Benchmarking coding agents on a multi-million line codebasehttps://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebaseEngineering blog · First party · Last source verification: 2026-08-31
Research details for every claim on this page
  1. Summary
    Statement type
    Fact
    Provenance
    Reported
    Confidence
    High
    Confidence reason
    A linked first-party source states the claim.
  2. Harness
    Statement type
    Fact
    Provenance
    Reported
    Confidence
    High
    Confidence reason
    A linked first-party source states the claim.
  3. Knowledge
    Statement type
    Fact
    Provenance
    Reported
    Confidence
    High
    Confidence reason
    A linked first-party source states the claim.
  4. Key observation
    Statement type
    Fact
    Provenance
    Reported
    Confidence
    Medium
    Confidence reason
    Databricks described internal use in its own engineering blog.
  5. Key observation
    Statement type
    Fact
    Provenance
    Reported
    Confidence
    Medium
    Confidence reason
    Databricks described the benchmark in its own engineering blog.
  6. Operating model assessment
    Statement type
    Inference
    Provenance
    Catalog judgment
    Confidence
    Unverified
    Confidence reason
    The record covers several internal engineering agents with different workflows, so no single human-attention boundary applies.
    Observation date
    2025