The Task Exposure Indexv2026.Q3

Pre-registered predictions, release v2026.Q3

Registered: 15 September 2026 Resolves against: releases v2026.Q4 and v2027.Q1 Status: open

An index that never says anything checkable is a branding exercise. This page states, in advance, what we expect the next two quarters to show. Each prediction has a number, a deadline and a resolution rule written before the result is known. When the next release publishes, every one of these is scored here, including the ones that were wrong, and the scoring stays up permanently.

The rules are fixed now so they cannot be reinterpreted later:

it resolves wrong and the drafting error is recorded next to it.

Where the index stands today

These are the figures the predictions below move against.

Measurev2026.Q3
Activities rated2,087
Activities at capability 3 or 4816
Activities at capability 2322
Tasks banded exposed8,693 of 18,838
Occupations banded exposed413 of 923
Median occupation exposed share27.9%
Mean friction, context1.62 of 3
Mean friction, embodiment1.52 of 3

The predictions

1. Capability moves, friction does not

By v2027.Q1, the mean capability rating across all 2,087 activities rises by at least 0.10 of a point, while the mean of the five friction dimensions moves by less than 0.05 in either direction.

Reasoning: capability is a property of deployed systems and has moved every quarter for three years. Friction is a property of law, liability and physical reality, and those do not move on a quarterly cadence. If this is wrong, the two-factor structure of the index is measuring one thing rather than two, which would be the most important finding this project could produce about itself.

2. Context erodes first

By v2027.Q1, the mean context friction falls by at least 0.05 of a point, and falls by more than any other friction dimension.

Reasoning: context is the only friction that is an engineering problem rather than a social one. Retrieval, long horizons and tool access are all attacks on it, and all three are under active commercial development. Embodiment, presence and accountability are not.

3. The embodiment floor holds

In v2027.Q1, the 791 activities currently rated embodiment 3 have a mean capability rating below 0.60. It is 0.34 today.

Reasoning: general purpose robotics is not generally available to an ordinary organisation and will not be within two quarters. If this fails, it fails for a specific and visible reason, and that reason will be worth more than the prediction.

4. The draft-quality band empties upward

By v2027.Q1, at least 60 of the 322 activities currently rated capability 2 are rated 3 or higher, and fewer than 15 are rated lower than 2.

Reasoning: capability 2 means the output needs substantial rework. That is the band that moves, because it is the band where marginal model improvement changes the answer. Downgrades should be rare and should reflect rater correction rather than regression.

5. The assisted band grows

In v2027.Q1, more than 48 occupations fall in the assisted band.

Reasoning: assisted requires high capability and high friction together. Capability is rising into work that carries accountability and presence, so occupations should cross from untouched into assisted faster than they cross from assisted into exposed. The assisted band is the smallest today at 48 of 923, and we expect it to be the fastest growing in relative terms.

6. The pay relationship persists

In v2027.Q1, the rank correlation between median annual pay and exposed share stays positive and stays below +0.50.

Reasoning: capability has climbed the wage ladder, but the top of the ladder is defended by friction, which caps how far the correlation can run. A correlation that goes negative would mean exposure has rotated toward low paid work, which would be a real change in the character of the technology. One above +0.50 would mean friction at the top is failing.

7. No occupation reaches 80%

In v2027.Q1, no occupation has an exposed share above 80%. The maximum today is Telemarketers at 73.3%.

Reasoning: every occupation in the federal taxonomy contains some task that requires a body, a signature or a person present. An occupation above 80% would mean we have found one that does not, which would say something about the taxonomy as much as about AI.

8. Our own reliability holds

The v2027.Q1 re-rating reproduces exact agreement of at least 80% against the v2026.Q3 ratings on the 40-activity calibration set, excluding activities we deliberately re-scored.

Reasoning: if our raters cannot reproduce their own judgments a quarter later, the time series is noise and should be treated as such. This is the prediction we are least comfortable publishing, which is why it is here.

What we are deliberately not predicting

exposure score would be exactly the overreach the methodology page warns about.

measurement question, and guessing at it would add noise to a record that is meant to be about the measurement.

cannot be wrong, so it cannot be worth anything.

Scoring history

No release has been scored yet. This is the first set. The scoring table appears here when v2026.Q4 publishes, and every subsequent release adds a row rather than replacing one.

ReleasePredictionsCorrectWrongScore
v2026.Q38openopenpending v2026.Q4