Workflow evals

Structured workflows for automation tasks

Real world tasks can be executed via structured workflows or standalone prompts. Structure is always better.

Lens
Mean accuracy vs cost and time
TypeSafeOpenAIAnthropicFireworksworkflowprompt
frontier: nothing is both cheaper and more accurate
40%50%60%70%80%$0.0001$0.001$0.01$0.1$1cost per case, USD (log)accuracyhaiku 4.5haiku 4.5 ↓ 18%opus 5opus 5sonnet 5sonnet 5DS v4 flashDS v4 flashDS v4 proDS v4 prolunalunasolsolterraterraJev

Each point averages one model configuration's accuracy, cost and time over the four workflows with equal weight, against the consensus labels. Every model runs at its provider's default reasoning setting. Up and to the left is better.

How we evaluate

Decompose the work, build a harness

To automate a task, we decompose the decisions into programmatic rules and intelligent judgments. Rather than ask a model to solve the entire problem in one shot (like the prompt examples in the plot), we ask independent narrow questions and defer to code where possible. We use three types: Noul, yes or no; Choice, one option among several; Score, a level on a scale. The results are then used programmatically to produce the output actions. Averaged across the four example tasks, every model is more accurate, cheaper and faster in the workflow than it is with the same policy as a prompt.

A toy example. The policy on the left is the kind of paragraph a team writes down; the chart on the right is the same policy as a workflow. Each sentence became either a question for the model, with a type, or a rule for the code.

Expense claims

1Every claim comes with a receipt. If the receipt cannot be read, ask the employee for a new one.

2Work out what kind of expense it is: a meal, travel, or equipment.

3A meal over $75 needs a manager's sign-off when the description on the claim does not clearly match the receipt.

4Everything else is approved.

One expense claimthe receipt · theclaim formThe model reads the claim:reads the receipt and the claim formthe receipt can be readwhat kind of expense: a meal, travel, orequipmenthow clearly the claim's description matchesthe receipt, on four levelsreadable?noNEW RECEIPTyesThe code decides:a meal over $75 whose description doesnot clearly matchMANAGER REVIEWanything elseAPPROVE1THE MODEL ANSWERS2THE CODE DECIDES
QuestionsNoulScoreChoice

Assume the harness is correct

Instead of debating the correctness of the harness and labels, we assume that the code is correct, and measure against the current smartest large models. For this eval, the reference labels are generated via an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking, answering every question in the harness. All other models are evaluated using the provider's default reasoning settings.