The benchmark for real work

CorpBench

Most benchmarks score answers. CorpBench scores the work itself: the final state, the actions taken, and whether the agent completed or escalated correctly.

14k
Tool calls
100
Workflows
5
Departments
15
Models

01Overall results

Which models perform best overall

Models are ranked on the quality of their work, how often they pass, and how efficiently they run.

CorpBench Score by model

Bars show the overall CorpBench Score from 0 to 100. The marker shows the share of scored runs that passed every grading requirement. Time per task shows how long a task took, averaged across scored runs. Select a model to see its full result.

CorpBench Score, with pass-rate marker
0255075100CorpBench Score (0-100) and pass rate (%)

What this measures. CorpBench Work Score is a provisional composite combining average outcome score (60%), strict pass rate (25%), and efficiency against fixed, versioned baselines (15%). The marker shows pass rate across scored runs. Provider failures are excluded from scored runs and reported separately.

02Cost and speed

What each task costs in money and time

Model quality is only part of the deployment decision. This view compares outcome with the cost and time required to run each task.

Outcome against cost per task

Up is better work. Left is cheaper. The green corner highlights outcomes of 85+ at lower cost.

Outcome score (0-100), vertical axis

Smart and cheaper

Smart, but expensive

10092.58542.50
$0.0000$0.0020$0.0077$0.0247$0.0750

Cost per task (USD) · logarithmic scale

The vertical scale expands scores from 85–100 into the upper half of the chart.

What this measures. Cost per task is total measured benchmark cost divided by scheduled runs. Time per task shows how long a task took, averaged across scored runs. Outcome score remains on the vertical axis in both views so cost and time are compared against the same result measure.

03Departments

Where the work gets harder

Overall scores can hide large differences between departments. A model that performs well across the benchmark may not be the best fit for every team.

Department outcome score

Average outcome score across scored runs in each department, on a 0 to 100 scale. Select a department to sort the models by its results.

Outcome score<6060+70+80+90+95+99+

Sorted by overall ranking

99
99
100
99
98
99
90
100
100
94
100
100
99
99
95
100
96
99
98
97

What this measures. Department outcome score is the average grader score across scored runs in that department. It includes partial credit, so it is not a pass rate or a recalculated CorpBench Score.

04Failure modes

How models fail at real work

A single score hides very different weaknesses. This view shows whether models stop early, skip required steps, give up, or get the handoff wrong.

Failure rate by behavior

Percentage of applicable scored runs exhibiting each failure. Lower is better. A run can show more than one, so the columns do not sum to 100%.

Failure rate · fixed colour bands0%≤1%≤3%≤5%≤10%≤20%>20%
Incomplete
workflow
Missing required
calls
Misordered required
calls
Abandoned
recoverable retry
Missing required
final state
Incorrect human
handoff
4%
2%
3%
9.1%
0%
0%
6.1%
6.1%
0%
0%
4.0%
1.0%
1%
1%
2%
0%
1%
0%
14.0%
14.0%
3%
22.7%
2%
0%

What this measures. Each value is the share of applicable scored runs in which the grader recorded that failure. Lower is better. A run can be recorded against more than one failure, so the columns overlap and do not sum to 100%. A blank cell means the failure was not scored for that model in this release.

05Reliability

How consistently each model succeeds

A strong average can hide inconsistent runs. Pass rate shows how often a model met every grading requirement.

Pass rate across scored runs

Share of scored runs that passed every grading requirement. Provider failures are excluded from scored runs and shown separately.

Pass rate
0%20%40%60%80%100%Pass rate (%)

What this measures. Pass rate is passed runs divided by scored runs. A run passes only when it meets every grading requirement. Provider failures after configured retries are excluded from scored runs and reported separately. Small score differences alone do not establish a meaningful capability difference.

06Methodology

How the results were produced

Each workflow runs in a fresh, deterministic workspace. Grading checks the final state, the required calls in order, and whether it completed or escalated correctly.

Workflows are observable work

Each workflow leaves evidence the grader can inspect: state changes, ordered tool calls, and a completion or escalation decision.

Grading is deterministic

The expected state and tool trace are fixed before the run. The same recorded evidence always receives the same grade.

Failures are separated

A provider failure after retries is excluded from scored runs and reported separately. Model errors, timeouts, and incorrect work remain in the scored results.

Repetitions define coverage

The release records workflow definitions and repetitions separately. A single run per workflow measures completion on that fixture; it does not establish consistency across variants.

About this release

CorpBench Work v1.3

Benchmark version
1.3
Workflows
100
Models
15
Published
Sept 7, 2026
Runs / workflow / model
1
Scheduled runs / model
100
Pricing snapshot
Sept 7, 2026
Model cohort
Custom research

This release uses one run per workflow per model. Pass rates describe success across these workflows; they do not measure repeatability on the same task.

Build Capacity To Grow. Become AI Native.

Bring the workflow costing you the most. We'll map the bottleneck and show you what the AI system could look like.