The benchmark for real work
Most benchmarks score answers. CorpBench scores the work itself: the final state, the actions taken, and whether the agent completed or escalated correctly.
- 14k
- Tool calls
- 100
- Workflows
- 5
- Departments
- 15
- Models
01Overall results
Which models perform best overall
Models are ranked on the quality of their work, how often they pass, and how efficiently they run.
CorpBench Score by model
Bars show the overall CorpBench Score from 0 to 100. The marker shows the share of scored runs that passed every grading requirement. Time per task shows how long a task took, averaged across scored runs. Select a model to see its full result.
What this measures. CorpBench Work Score is a provisional composite combining average outcome score (60%), strict pass rate (25%), and efficiency against fixed, versioned baselines (15%). The marker shows pass rate across scored runs. Provider failures are excluded from scored runs and reported separately.
02Cost and speed
What each task costs in money and time
Model quality is only part of the deployment decision. This view compares outcome with the cost and time required to run each task.
Outcome against cost per task
Up is better work. Left is cheaper. The green corner highlights outcomes of 85+ at lower cost.
Smart and cheaper
Smart, but expensive
Cost per task (USD) · logarithmic scale
The vertical scale expands scores from 85–100 into the upper half of the chart.
What this measures. Cost per task is total measured benchmark cost divided by scheduled runs. Time per task shows how long a task took, averaged across scored runs. Outcome score remains on the vertical axis in both views so cost and time are compared against the same result measure.
03Departments
Where the work gets harder
Overall scores can hide large differences between departments. A model that performs well across the benchmark may not be the best fit for every team.
Department outcome score
Average outcome score across scored runs in each department, on a 0 to 100 scale. Select a department to sort the models by its results.
Sorted by overall ranking
What this measures. Department outcome score is the average grader score across scored runs in that department. It includes partial credit, so it is not a pass rate or a recalculated CorpBench Score.
04Failure modes
How models fail at real work
A single score hides very different weaknesses. This view shows whether models stop early, skip required steps, give up, or get the handoff wrong.
Failure rate by behavior
Percentage of applicable scored runs exhibiting each failure. Lower is better. A run can show more than one, so the columns do not sum to 100%.
workflow
calls
calls
recoverable retry
final state
handoff
What this measures. Each value is the share of applicable scored runs in which the grader recorded that failure. Lower is better. A run can be recorded against more than one failure, so the columns overlap and do not sum to 100%. A blank cell means the failure was not scored for that model in this release.
05Reliability
How consistently each model succeeds
A strong average can hide inconsistent runs. Pass rate shows how often a model met every grading requirement.
Pass rate across scored runs
Share of scored runs that passed every grading requirement. Provider failures are excluded from scored runs and shown separately.
What this measures. Pass rate is passed runs divided by scored runs. A run passes only when it meets every grading requirement. Provider failures after configured retries are excluded from scored runs and reported separately. Small score differences alone do not establish a meaningful capability difference.
06Methodology
How the results were produced
Each workflow runs in a fresh, deterministic workspace. Grading checks the final state, the required calls in order, and whether it completed or escalated correctly.
Workflows are observable work
Each workflow leaves evidence the grader can inspect: state changes, ordered tool calls, and a completion or escalation decision.
Grading is deterministic
The expected state and tool trace are fixed before the run. The same recorded evidence always receives the same grade.
Failures are separated
A provider failure after retries is excluded from scored runs and reported separately. Model errors, timeouts, and incorrect work remain in the scored results.
Repetitions define coverage
The release records workflow definitions and repetitions separately. A single run per workflow measures completion on that fixture; it does not establish consistency across variants.
About this release
CorpBench Work v1.3
- Benchmark version
- 1.3
- Workflows
- 100
- Models
- 15
- Published
- Sept 7, 2026
- Runs / workflow / model
- 1
- Scheduled runs / model
- 100
- Pricing snapshot
- Sept 7, 2026
- Model cohort
- Custom research
This release uses one run per workflow per model. Pass rates describe success across these workflows; they do not measure repeatability on the same task.
Build Capacity To Grow. Become AI Native.
Bring the workflow costing you the most. We'll map the bottleneck and show you what the AI system could look like.