02 / Human attention
More comments can mean more work
A review agent also creates work for its reader.
What the teams report
Uber found that a single prompt produced false alarms and valid comments with little value. Its uReview system checks confidence and removes duplicate comments. It also suppresses categories that engineers rarely use. [1]
“Precision Is More Valuable than Volume”
A section title in Uber’s uReview report. [1]
HubSpot made its reviewer faster, but review quality remained a problem. It added a judge agent to check comments before publication. [2]
Our observation
Measure the work after the output
Comment count shows how much an agent writes. It does not show how much useful work the team completes.
A useful evaluation can track accepted findings, review time, and missed defects. A filter can reduce noise and still remove a real issue.
A model judge can also make mistakes. The reports do not establish that an extra agent always improves a review.
A question for your buildDoes each additional comment save more work than it creates?
Sources
- Uber: uReviewQuality filters, duplicate removal, and feedback from engineers.
- HubSpot: Sidekick code reviewThe change to Aviator and the judge agent.