The scale

H1 to H5: how much of a task stays human.

The Human Agency Scale is a five-level scale for how much of a task a person holds once AI is in it. H1 is AI alone. H5 is fully human. Higher always means more human agency, not less.

The scale is Stanford's, from the WORKBank research, which scored 844 tasks across 104 occupations. What is ours is the rubric that places a task on it and applying that rubric at task grain across four levels of an organization.

The five levels

Read the mark, not the code.

Fill shows the human share

H1 AI does it aloneH2 AI does it, person spot-checksH3 AI and person share itH4 Person leads, AI helpsH5 Fully human

The more AI carries at the low end, the more valuable the fully human work becomes. That is what makes the scale a design tool rather than a scoreboard: it shows you which work to move, and which work to invest in.

Placing a task

One question per level, asked in order.

Score the task, not the role, and score it as performed at the quality the work actually requires. Work upward from H1. The first "no" places the task.

LevelThe question
H1Could AI produce the complete output at required quality with no human touch at all — no review, no approval?
H2Does AI produce the primary output, with the person contributing review, approval or correction at defined checkpoints?
H3Do the person and AI each contribute substance, such that neither alone reaches required quality?
H4Does the person produce the output, with AI retrieving, drafting fragments or speeding things up — removable without changing who authors the work?
H5Does the task's value come specifically from human origin — accountability, trust, authority, relationship, presence?

Where scores drift

Three boundaries do most of the damage.

H1 / H2

Is there a real checkpoint?

One required approval makes it H2. "Someone glances at it out of habit" is not a checkpoint — ask what happens when the glance stops.

H2 / H3

Who holds the shape?

If AI drafts and the person's edits are corrective, it is H2. If the person supplies substance AI cannot, it is H3. The test: if their contribution stopped, would the output be wrong, or merely worse? Wrong is H3.

H4 / H5

Is the human origin the point?

Would the recipient accept the output knowing AI touched it? A client entrusting a relationship, a regulator holding an accountable officer, a negotiation — those are H5.

The two-axis check

When capability and accountability disagree, accountability wins.

A score carries two things at once: who can produce the output, and who has to own it. AI can draft a credit sanction. A person has to own it. That is H2 with a pinned approval checkpoint — and it cannot drift to H1 while the accountability requirement stands.

This is why the scale is a design tool. A step deliberately held at H5 is a decision with a reason behind it, recorded as such, not a gap someone forgot to close.

Keeping scores honest

A rubric with worked examples at every level.

Every score is scored twice — as the work runs today, and as it would run after redesign — and both carry the evidence they rest on. In a 164-task calibration study, independent raters agreed exactly on 89% of scores and within one level on 100%.

Common questions

What people ask about the scale.

Q01What is the Human Agency Scale?
A five-level scale, H1 to H5, for how much of a task a person holds once AI is in it. H1 means AI does the task alone. H5 means the task is fully human. It comes from Stanford’s WORKBank research, which scored 844 tasks across 104 occupations. Higher always means more human agency, not less.
Q02Does a higher H number mean more AI or more human?
More human. H1 is AI alone and H5 is fully human. The scale is never inverted. Reading it the other way round turns every downstream number upside down.
Q03How do you decide which level a task sits at?
One discriminating question per level, applied in order from H1 upward, and the first "no" places the task. Could AI produce the complete output with no human touch at all? Does AI produce the primary output with the human only reviewing at defined checkpoints? Do both contribute substance? Does the human author it with AI only accelerating? Does the value derive specifically from human origin?
Q04What is the difference between H2 and H3?
Who holds the shape of the output. If AI drafts and the human’s edits are corrective, it is H2. If the human supplies substance AI cannot, it is H3. The test: if the human’s contribution stopped, would the output be wrong, or merely worse? Wrong is H3. Merely worse is H2.
Q05What is the difference between H4 and H5?
Whether the human origin is constitutive. Ask whether the recipient would still accept the output knowing AI touched it. A client entrusting a relationship, a regulator holding an accountable officer, a negotiation — those are H5. If AI assistance is merely unmentioned but acceptable, it is H4.
Q06Can a task be scored H1 if a human is legally accountable for it?
No. A score encodes two things: who produces the output, and who must own it. When they diverge, accountability wins. AI can draft a credit sanction, but a human must own it — that is H2 with a pinned approval checkpoint, and it cannot drift to H1 while the accountability requirement stands.
Q07Is the Human Agency Scale Effectv’s?
No. The scale is Stanford’s, from the WORKBank research. What is ours is the scoring rubric with worked examples at each level, and applying it at task grain across four levels of an organization to drive decisions.

Score one of your own tasks against it.

Not a sales meeting. A conversation about whether this addresses a problem you recognize in your organization, and what you would need to see to hold the answer under scrutiny from your CFO and your board.

Book the thirty minutes

Pick a time on the calendar

Who you get

Rajeev, who runs the engagement. Not a qualifier.

Bring one workflow you already know is painful. We trace it one or two layers out loud with you, tell you what we would need to see to answer it properly, and say what we could not answer.