Workflow evals

Structured workflows for automation tasks

Real world tasks can be executed via structured workflows or standalone prompts. Structure is always better.

Lens
Mean accuracy vs cost and time
TypeSafeOpenAIAnthropicFireworksworkflowprompt
frontier: nothing is both cheaper and more accurate
40%50%60%70%80%$0.0001$0.001$0.01$0.1$1cost per case, USD (log)accuracyhaiku 4.5 · workflow · 53.6% · $0.0195 · 12.5 shaiku 4.5 · prompt · 18.1% · $0.0363 · 21.2 sopus 5 · workflow · 73.1% · $0.1761 · 37.8 sopus 5 · prompt · 64.8% · $0.3417 · 70.5 ssonnet 5 · workflow · 67.8% · $0.1174 · 78.1 ssonnet 5 · prompt · 60.4% · $0.2251 · 149.2 sDS v4 flash · workflow · 64.4% · $0.0059 · 51.9 sDS v4 flash · prompt · 59.3% · $0.0132 · 120.1 sDS v4 pro · workflow · 65.5% · $0.0413 · 86.5 sDS v4 pro · prompt · 59.7% · $0.0907 · 192.1 sluna · workflow · 66.8% · $0.0033 · 12.9 sluna · prompt · 51.9% · $0.0079 · 27.3 ssol · workflow · 74.1% · $0.0836 · 23.3 ssol · prompt · 63.4% · $0.2005 · 48.6 sterra · workflow · 67.9% · $0.0304 · 10.1 sterra · prompt · 61.6% · $0.0750 · 25.1 sJev · workflow · 67.8% · $0.0004 · 0.4 shaiku 4.5haiku 4.5 ↓ 18%opus 5opus 5sonnet 5sonnet 5DS v4 flashDS v4 flashDS v4 proDS v4 prolunalunasolsolterraterraJev

Each point averages one model configuration's accuracy, cost and time over the four workflows with equal weight, against the consensus labels. Every model runs at its provider's default reasoning setting. Up and to the left is better.

How we evaluate

Decompose the work, build a harness

To automate a task, we decompose the decisions into programmatic rules and intelligent judgments. Rather than ask a model to solve the entire problem in one shot (like the prompt examples in the plot), we ask independent narrow questions and defer to code where possible. We use three types: Noul, yes or no; Choice, one option among several; Score, a level on a scale. The results are then used programmatically to produce the output actions. Averaged across the four example tasks, every model is more accurate, cheaper and faster in the workflow than it is with the same policy as a prompt.

A toy example. The policy on the left is the kind of paragraph a team writes down; the chart on the right is the same policy as a workflow. Each sentence became either a question for the model, with a type, or a rule for the code.

Expense claims

1Every claim comes with a receipt. If the receipt cannot be read, ask the employee for a new one.

2Work out what kind of expense it is: a meal, travel, or equipment.

3A meal over $75 needs a manager's sign-off when the description on the claim does not clearly match the receipt.

4Everything else is approved.

One expense claimthe receipt · theclaim formThe model reads the claim:reads the receipt and the claim formthe receipt can be readwhat kind of expense: a meal, travel, orequipmenthow clearly the claim's description matchesthe receipt, on four levelsreadable?noNEW RECEIPTyesThe code decides:a meal over $75 whose description doesnot clearly matchMANAGER REVIEWanything elseAPPROVE1THE MODEL ANSWERS2THE CODE DECIDES
QuestionsNoulScoreChoice

Assume the harness is correct

Instead of debating the correctness of the harness and labels, we assume that the code is correct, and measure against the current smartest large models. For this eval, the reference labels are generated via an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking, answering every question in the harness. All other models are evaluated using the provider's default reasoning settings.

Example workflows

Security Incidents

A security alert fires on a laptop or a server. Given the alert and everything on file about that machine, we decide whether to close it, pass it to an analyst, or contain it now.

10%30%50%70%$0.0001$0.001$0.01$0.1haiku 4.5 · workflow · 58.8% · $0.0047 · 3.4 shaiku 4.5 · prompt · 17.1% · $0.0068 · 4.6 sopus 5 · workflow · 66.2% · $0.0574 · 15.1 sopus 5 · prompt · 62.9% · $0.0619 · 13.5 ssonnet 5 · workflow · 60.8% · $0.0271 · 18.9 ssonnet 5 · prompt · 51.7% · $0.0506 · 37.5 sDS v4 flash · workflow · 37.9% · $0.0032 · 37.4 sDS v4 flash · prompt · 44.6% · $0.0037 · 41.9 sDS v4 pro · workflow · 41.7% · $0.0234 · 60.0 sDS v4 pro · prompt · 37.1% · $0.0267 · 66.3 sluna · workflow · 52.1% · $0.0013 · 7.0 sluna · prompt · 29.6% · $0.0021 · 13.4 ssol · workflow · 62.5% · $0.0295 · 8.5 ssol · prompt · 45.8% · $0.0472 · 14.0 sterra · workflow · 51.2% · $0.0119 · 5.7 sterra · prompt · 45.4% · $0.0140 · 7.1 sJev · workflow · 61.7% · $0.0001 · 0.3 s

Agent Trace Observability

A support agent has just finished with a customer. Given the whole run, every tool call included, we decide whether a person needs to look at it, and how soon.

20%40%60%80%$0.0001$0.001$0.01$0.1$1haiku 4.5 · workflow · 57.2% · $0.0100 · 7.1 shaiku 4.5 · prompt · 29.3% · $0.0163 · 9.3 sopus 5 · workflow · 75.2% · $0.1033 · 27.4 sopus 5 · prompt · 63.5% · $0.1642 · 37.5 ssonnet 5 · workflow · 68.0% · $0.0545 · 38.0 ssonnet 5 · prompt · 65.8% · $0.0863 · 53.0 sDS v4 flash · workflow · 73.0% · $0.0043 · 51.7 sDS v4 flash · prompt · 68.5% · $0.0063 · 62.7 sDS v4 pro · workflow · 71.6% · $0.0357 · 90.1 sDS v4 pro · prompt · 72.1% · $0.0424 · 92.6 sluna · workflow · 76.1% · $0.0025 · 14.5 sluna · prompt · 66.2% · $0.0038 · 15.1 ssol · workflow · 76.6% · $0.0575 · 40.3 ssol · prompt · 73.9% · $0.0907 · 37.0 sterra · workflow · 73.0% · $0.0209 · 11.4 sterra · prompt · 72.5% · $0.0350 · 14.9 sJev · workflow · 71.6% · $0.0003 · 0.5 s

Invoice Processing

A vendor's bill arrives. Given the bill, the order behind it, and what was actually delivered, we decide whether it gets paid, held, or sent back.

0%20%40%60%80%$0.001$0.01$0.1$1haiku 4.5 · workflow · 42.9% · $0.0558 · 30.8 shaiku 4.5 · prompt · 6.0% · $0.0763 · 63.4 sopus 5 · workflow · 78.4% · $0.4856 · 92.1 sopus 5 · prompt · 66.9% · $0.7699 · 188.7 ssonnet 5 · workflow · 72.9% · $0.3616 · 241.3 ssonnet 5 · prompt · 62.4% · $0.5548 · 409.4 sDS v4 flash · workflow · 69.8% · $0.0133 · 84.1 sDS v4 flash · prompt · 58.2% · $0.0288 · 281.7 sDS v4 pro · workflow · 72.7% · $0.0830 · 137.4 sDS v4 pro · prompt · 63.8% · $0.2039 · 474.4 sluna · workflow · 67.8% · $0.0081 · 21.4 sluna · prompt · 55.1% · $0.0160 · 61.5 ssol · workflow · 79.1% · $0.2152 · 34.3 ssol · prompt · 65.1% · $0.4338 · 118.4 sterra · workflow · 74.7% · $0.0778 · 17.3 sterra · prompt · 64.0% · $0.1618 · 62.0 sJev · workflow · 61.8% · $0.0011 · 0.5 s

Customer Service

A customer writes in. Given the thread so far and the state of their account, we decide what the assistant should say and do next.

10%30%50%70%90%$0.0001$0.001$0.01$0.1$1haiku 4.5 · workflow · 55.4% · $0.0074 · 8.8 shaiku 4.5 · prompt · 19.9% · $0.0459 · 7.6 sopus 5 · workflow · 72.4% · $0.0579 · 16.6 sopus 5 · prompt · 66.0% · $0.3710 · 42.2 ssonnet 5 · workflow · 69.3% · $0.0264 · 14.3 ssonnet 5 · prompt · 61.6% · $0.2085 · 96.7 sDS v4 flash · workflow · 76.8% · $0.0029 · 34.6 sDS v4 flash · prompt · 65.8% · $0.0139 · 93.9 sDS v4 pro · workflow · 76.1% · $0.0232 · 58.6 sDS v4 pro · prompt · 65.8% · $0.0898 · 134.9 sluna · workflow · 71.4% · $0.0013 · 8.8 sluna · prompt · 56.9% · $0.0094 · 19.2 ssol · workflow · 78.3% · $0.0323 · 10.1 ssol · prompt · 69.0% · $0.2303 · 24.9 sterra · workflow · 72.7% · $0.0111 · 6.0 sterra · prompt · 64.4% · $0.0894 · 16.3 sJev · workflow · 76.0% · $0.0001 · 0.4 s