The scale
H1 to H5: how much of a task stays human.
The Human Agency Scale is a five-level scale for how much of a task a person holds once AI is in it. H1 is AI alone. H5 is fully human. Higher always means more human agency, not less.
The scale is Stanford's, from the WORKBank research, which scored 844 tasks across 104 occupations. What is ours is the rubric that places a task on it and applying that rubric at task grain across four levels of an organization.
The five levels
Read the mark, not the code.
Fill shows the human share
The more AI carries at the low end, the more valuable the fully human work becomes. That is what makes the scale a design tool rather than a scoreboard: it shows you which work to move, and which work to invest in.
Placing a task
One question per level, asked in order.
Score the task, not the role, and score it as performed at the quality the work actually requires. Work upward from H1. The first "no" places the task.
| Level | The question |
|---|---|
| H1 | Could AI produce the complete output at required quality with no human touch at all — no review, no approval? |
| H2 | Does AI produce the primary output, with the person contributing review, approval or correction at defined checkpoints? |
| H3 | Do the person and AI each contribute substance, such that neither alone reaches required quality? |
| H4 | Does the person produce the output, with AI retrieving, drafting fragments or speeding things up — removable without changing who authors the work? |
| H5 | Does the task's value come specifically from human origin — accountability, trust, authority, relationship, presence? |
Where scores drift
Three boundaries do most of the damage.
Is there a real checkpoint?
One required approval makes it H2. "Someone glances at it out of habit" is not a checkpoint — ask what happens when the glance stops.
Who holds the shape?
If AI drafts and the person's edits are corrective, it is H2. If the person supplies substance AI cannot, it is H3. The test: if their contribution stopped, would the output be wrong, or merely worse? Wrong is H3.
Is the human origin the point?
Would the recipient accept the output knowing AI touched it? A client entrusting a relationship, a regulator holding an accountable officer, a negotiation — those are H5.
The two-axis check
When capability and accountability disagree, accountability wins.
A score carries two things at once: who can produce the output, and who has to own it. AI can draft a credit sanction. A person has to own it. That is H2 with a pinned approval checkpoint — and it cannot drift to H1 while the accountability requirement stands.
This is why the scale is a design tool. A step deliberately held at H5 is a decision with a reason behind it, recorded as such, not a gap someone forgot to close.
Keeping scores honest
A rubric with worked examples at every level.
Every score is scored twice — as the work runs today, and as it would run after redesign — and both carry the evidence they rest on. In a 164-task calibration study, independent raters agreed exactly on 89% of scores and within one level on 100%.
Common questions
What people ask about the scale.
Q01What is the Human Agency Scale?
Q02Does a higher H number mean more AI or more human?
Q03How do you decide which level a task sits at?
Q04What is the difference between H2 and H3?
Q05What is the difference between H4 and H5?
Q06Can a task be scored H1 if a human is legally accountable for it?
Q07Is the Human Agency Scale Effectv’s?
Score one of your own tasks against it.
Not a sales meeting. A conversation about whether this addresses a problem you recognize in your organization, and what you would need to see to hold the answer under scrutiny from your CFO and your board.
Rajeev, who runs the engagement. Not a qualifier.
Bring one workflow you already know is painful. We trace it one or two layers out loud with you, tell you what we would need to see to answer it properly, and say what we could not answer.

