Open specification v0.2

Agent Effectiveness
Index.

How well does an agent understand the business,
do the work, and learn?

We have seen the bills for an agent. How do we know how much it helps? AEI defines a standard way to measure an agentic system against the work a business needs it to do.

01 / THE INDEX

Understand. Do. Learn.

Three questions we would ask about anyone doing work for the business.

B40%

Business understanding

Know the business, its facts, and its rules.

Entities · relationships · evidence · policy
O40%

Operational execution

Do the work correctly, within its authority.

Final state · permissions · side effects
L20%

Learning persistence

Retain, transfer, and adapt what it learns.

Gain · transfer · retention · adaptation
THE COMPOSITE
AEI = 100 × B0.40 × O0.40 × L0.20
0–100 scale
Weighted geometric mean of B, O, and L (each 0–1). Safety and resource costs are reported separately.
What makes two AEI scores comparable?

The same work and conditions

Declare the task suite, system version, tools, permissions, teaching profile, and budgets. Access to documents and a business graph is part of the comparison, not something left implicit.

Access profiles: K0 starting system; K1 documents; K2 graph with a base navigator; K3 the same graph with a trained navigator.

Evidence of what happened

Use answers and citations for understanding, environment state and action records for execution, and matched tests over time for learning. A transcript helps explain the work; it does not replace checking the outcome.

A score with its context

Report B, O, L, the safety gate, the evaluation scope, and 95% confidence intervals. The reference method uses 10,000 stratified paired block-bootstrap resamples, conditional on the declared benchmark population. A critical prohibited action sets episode execution to zero and fails the reference safety gate. If a required learning panel is missing, L and the full AEI are unavailable.

HOW SHOW & TELL CONNECTS TO AEI

What stays after the lesson?

This repository tests Teach → Comprehend: what an agent picks up from one demonstration.

Cold testThe starting point
SHOW & TELL · THIS REPOSITORY

TeachOne narrated workflow

ComprehendRules, exceptions, authority

CheckpointPermitted learned state
Comprehend diagnoses acquisition. The full AEI learning score requires the matched panel below. Cold testing precedes a reset; the teaching intervention then starts from the pre-cold state.
The full AEI learning panel
Independent tests from the same checkpoint
ImmediateDid it improve?
TransferCan it handle a new case?
Delayed · 24hDoes it retain the lesson?
DriftCan it adapt to a change?

A demonstration is one declared teaching profile. Written procedures, feedback, practice, and training are also supported by AEI.

TASKS / SHOW & TELL

Show the work. Test the understanding.

Business processes with rules, boundaries, exceptions, and consequences.

Task inventory grouped by business process; counts and category names are shown in the figure and linked data.
Case inventory from the workbook; groupings are based on process names. Inspect the classification ↗
Another way to group tasks: TCI

TCI is a useful heuristic for choosing tasks for the arena. It helps us spread the benchmark across expected difficulty levels and look for tasks we expect to be harder. It groups tasks by their rules, evidence, steps, and demonstration. This view uses 49 historical task definitions from 26 July 2026, separate from the workbook cases above.

Task definitions grouped into five heuristic TCI buckets; the figure labels counts and the linked export contains the inputs.
See the components for eight example tasks ↗ · Inspect the inputs ↗ · Grouping method ↗

EXAMPLE / INSIDE A TASK

The rule. The boundary.
The exception.

Customer Returns Inbox Triage · One inbox, several rules, and an exception that overrides them.

Same inbox. Different decisions.

Frames from the human demonstration · Open an image to read it at full size

Recorded Roundcube reply to Priya approving an £18.40 refund to the original card within three to five working days.

(a) Approve directly

£18.40 · within the return window

Refund to the original card; three to five working days; no return needed.

Recorded Roundcube forward to finance with the £142 amount, nine-day delivery age, and non-faulty return reason.

(b) Ask finance to approve

£142 · above the direct approval limit

Forward the amount, delivery age, and reason to finance.

Recorded Roundcube reply to Tom declining a non-faulty return delivered on 21 February, outside the 30-day window.

(c) Decline under the normal policy

21 February delivery · outside 30 days

Explain why the normal return window has passed.

Original frames from the human demonstration; prepared decisions, not agent results. Source: Brackett AI, Show & Tell dataset ↗
Customer Returns Inbox Triage · Roundcube · 8m 37s · application view
Browser chrome cropped; opening and closing frames held.
Recording, narration, screenshots, and questions ↗

ONE NARRATED WORKFLOW

More than knowing
where to click.

Review the morning returns inbox. Work out what qualifies, what needs approval, and what should be left alone.

Normal return window
≤ 30 days from delivery
Direct approval limit
≤ £50
Arrived damaged
Overrides both normal gates

These are the rules of this benchmark scenario. The question set contains 11 questions; three are explored below.

Try three questions: a boundary, an exception, and a wrong automation

QUESTION FROM THE DATASET · CLOSED ANSWER

On 9 April, Sam requested £41 for an item delivered on 10 March, while Nadia requested £50 for an item delivered on 4 April. How should the handler process both requests?

Source question ↗

WHAT THIS TESTS

The exact boundary matters. “Within 30 days” and “£50 or less” are inclusive.

See the expected answer
Sam · 30 days · £41Approve directly
Nadia · 5 days · £50Approve directly

Approve both requests and reply directly to Sam and Nadia.

What do the 11 questions test, and where are the answers grounded?
Evidence matrix for all eleven questions: nine cite narration, six cite screen captures, and four cite source files; a question can use several evidence types.
8 closed-answer questions and 3 rubric-scored questions. No multiple-choice questions in this published task. Questions and evidence ↗

Published returns task · revision f4aa425. This revised quiz illustrates task design; it is not asserted to match every earlier result.

RESULTS / COMPREHENSION AFTER TEACH

Show & Tell

Preliminary results

One demonstration. Questions about the rules, exceptions, and decisions it contained.

Show & Tell quiz results. They do not represent a full AEI score.

Mean comprehension score

24 cases attempted by all three systems · equal weight per case

Brackett71 / 71 completed attempts
84.0%
Codex24 / 24 completed attempts
58.1%
Claude10 / 48 completed attempts
15.7%
Incomplete attempts score zero; untested cells are excluded. The shared-case view compares the same set of cases. “Completed” means a numeric benchmark score was recorded, not successful execution of the business workflow.
Snapshot · 12 Sep 2026Download data ↓

System/model labels: Brackett: Fused/multiple models (107 selected attempts). Claude: Sonnet 5 (60 selected attempts). Codex: Not recorded (24 selected attempts). Brackett uses the author-specified fused/multiple-model description. Other model labels are workbook-reported, not verified model IDs; model and adapter versions are not fully pinned. Grader names in comments do not identify the evaluated model.

Share of scored and incomplete attempts for each system; the figure labels counts and the linked CSV lists every attempt.
Scored answers and incomplete attempts contribute to the aggregate. Attempt details ↓
Case-by-system heatmap · all 39 tested cases
Heatmap of mean comprehension score by business process and system. Grey cells indicate untested cases. Full values are available in the case table below.
Highest case means across 24 shared cases — Brackett: 18, Codex: 4, Claude: 2. These are case means, not significance tests. PNG ↓
How much do repeated attempts differ?
Largest observed ranges between selected attempts of the same case per system; incomplete attempts are marked with crosses.
Largest observed ranges, selected explicitly for this diagnostic. Brackett cases have two or three selected attempts, Claude cases two. These repetitions do not measure learning. All repeat gaps ↗
Results, coverage, and scoring

What is included

The snapshot includes 39 cases with at least one attempted system and 3 cases not yet tested by any system. Brackett: 107 selected attempts across 39 tested cases. Claude: 60 selected attempts across 30 tested cases. Codex: 24 selected attempts across 24 tested cases. Codex Record & Replay is tested on macOS because Teach relies on OS internals.

How the average is calculated

Average the attempt scores within each case, then weight cases equally. An NA is an incomplete attempt worth zero. A blank is untested and excluded. The shared view uses cases attempted by every system, including incomplete attempts.

What this snapshot establishes

This is a work in progress with selected attempts and non-canonical grading. The first two Brackett attempts per case and the Claude attempts were graded manually; the third Brackett attempt set and the nine cases tested only by Brackett carry imported grades produced by Claude Sonnet 4.6, as recorded in the workbook notes. Attempts the workbook marks as earlier and excluded from its overview are listed in the analysis export but do not enter any score. Model and adapter versions are not fully pinned in the workbook. No confidence interval is reported for this snapshot; it is not a verified canonical leaderboard.

Source: Brackett ShowTell Benchmark Results.xlsx, Overview sheet. The machine-readable snapshot includes source hash, individual attempt scores, and aggregation rules. The repository defines the canonical quiz scoring protocol separately.

Explore results by business process

Load the interactive results to explore individual cases.

Comprehension scores by business process, averaged across attempts
Business processBrackettCodexClaude

“1” is retained where it appears in the source case name. A dash means untested; a zero can include incomplete attempts. Use the CSV for attempt-level status.

OPEN RESEARCH

Read it. Run it. Build on it.

Specification, code, task data, and reproducible figures.

See how to run or create a benchmark task
Terminal walkthrough showing how to create a Show and Tell benchmark task
Create a task. Capture a workflow and prepare its evidence and quiz. Read the repository guide ↗