Skip to content

Evaluating Coding Agents

Evaluate more than whether a task appears completed.

Dimension Example evidence
Task success Acceptance tests and user-visible behaviour
Engineering quality Review of architecture, readability and maintainability
Process quality Appropriate exploration, tool use and adherence to instructions
Safety No unauthorised access, secret exposure or unsafe action
Efficiency Time, tokens, tool calls and unnecessary churn
Robustness Success across varied repositories and ambiguous cases
Collaboration Clear handoffs, bounded diffs and useful progress communication

Trace and provenance record task assignment, actions, results and artifact ownership. They support debugging, audit and workflow evaluation.