INDEPENDENT STUDY · FOUNDING PILOT

The last word
is human.

An expert and an AI agent.
The same brief. You decide which work holds up.

What does expertise add when everyone has access to powerful AI? We’re putting real judgment to the test, one anonymous comparison at a time.

$50 per invited review. Your preference never changes your pay.

THE JOHN HENRY EXPERIMENT№ 001
The tools are equal.
Is the work?
01 / HUMAN-DIRECTEDExpert
+ AI
Experience at the controls.
VS
02 / SELF-DIRECTEDAutonomous
agent
Room to plan, revise, decide.
SAME BRIEFSAME TOOLSSAME CLOCK
THE WORK IS ANONYMOUS.
THE AUDIENCE HAS THE LAST WORD.
EXPERTISE, UNDER EXAMINATIONNO AUTHOR LABELSNO PREFERRED WINNERHUMAN REASONS MATTER

Intelligence can make an answer.
Judgment makes it matter.

A convincing memo isn’t necessarily a useful decision. A polished concept isn’t necessarily a good one. John Henry asks where human expertise still changes the outcome—and where a capable agent is catching up.

The name comes from the legend of a worker racing a machine. Our version gives the expert the machine, too. Then we remove the names and ask people who know the work to choose.

Give both sides
a fair shot.

Human expertise should be tested against an agent that has the context and opportunity to do good work.

  1. 01 / DEFINE

    You set the standard.

    Reviewers describe what useful work looks like before the task begins. Both sides receive the same de-identified criteria and answers to clarification questions.

  2. 02 / MAKE

    One brief. Equal resources.

    An expert using AI and an autonomous agent share the source packet, deadline, tool access, and declared spending limit. Each chooses its own final answer.

  3. 03 / JUDGE

    The work speaks first.

    Reviewers compare anonymous outputs in randomized order, choose a preference, and explain why. “Equally good” and “neither acceptable” are valid answers.

The agent can plan, research, and revise within the budget. This is more than a single prompt.

Read the pilot protocol

Which would
you put to work?

Try a small example of the review experience. Both samples were written to demonstrate the interface; this is not a benchmark result.

ILLUSTRATIVE BRIEF · HEALTHCARE STRATEGY

A promising product. A limited budget.

A health plan is considering an AI tool that flags members at risk of losing coverage. It can run a 90-day pilot. What should the team do first, and what would justify expansion?

DEMO ONLY
OUTPUT AIDENTITY WITHHELD

Test the decision, then the model.

Start by identifying where outreach can change a member’s next step. Compare the tool against a simple rules-based list using the same outreach capacity, with staff reviewing flags before contact.

Track successful coverage renewals, missed cases, staff time, and errors across relevant groups. A useful alert must arrive early enough to act on and lead to assistance the member can actually use.

Expand only if the pilot improves renewal outcomes beyond the simpler approach without adding unacceptable burden or disparities. Ninety days may establish feasibility; longer follow-up may be needed to establish impact.

OUTPUT BIDENTITY WITHHELD

Build a focused path to scale.

Choose one region with reliable enrollment data and an engaged operations team. Establish baseline renewal and outreach rates, validate the tool on recent cases, and launch a weekly dashboard.

Give a dedicated team a prioritized member list. Record contact rates, resolved issues, member feedback, and operational costs. Review progress weekly so the team can address data gaps and improve its outreach process.

At day 90, decide whether to expand based on measurable improvement from baseline, manageable costs, and positive staff feedback. Document implementation lessons before adding another region.

Which response would you prefer to use?

Problems with
something at stake.

The proposed pilot starts close to the organizer’s experience. Each task will use a frozen brief and public or synthetic material.

01

Healthcare economics

Policy interpretation, Medicaid financing, and consequential tradeoffs.

02

Product & strategy

What to build, what to measure, and when to change direction.

03

AI & human experience

Research judgment, narrative choices, and the quality of an interaction.

AN INVITATION TO PEOPLE WHO KNOW THE WORK

Bring your standards.
Keep your independence.

You don’t need to be an AI enthusiast. We’re looking for people with relevant experience and a point of view on what makes work worth using.

  • Share your criteria before either side starts.
  • Read two anonymous outputs and explain your preference.
  • Be paid the same whether you favor the expert, the agent, both, or neither.
Express interest by email

Show your
work.

Is this trying to prove that humans are better?

It begins with curiosity about the value of human expertise. A useful test must also be able to show that the agent wins. We’ll report wins, losses, ties, limitations, and reviewers’ reasons—not just a flattering headline.

Can’t reviewers recognize AI writing?

They may. Professional judgment tasks will test a common editing treatment on both outputs, with a check that the substance stays intact. Originals will be retained. Creative work will be evaluated in its original voice. The exact treatment is fixed before each round; authorship guesses are collected after preference.

Who is behind John Henry?

John Henry is an independent project by Shane Francis, a healthcare economist exploring AI, design, and expert judgment. Shane will be the first human participant. The pilot is also a public record of the process: what we built, what we tested, and what we learned.

How is the pilot funded?

The planned sponsor-funded budget is $2,000: $1,600 for 32 reviews across four task pairs, $200 for four follow-up reviews, and $200 reserved for operating costs. We’ll start with one task and eight reviewers, then decide how to expand. These are plans, not completed study activity. The pilot has no betting or entry fee.

When will there be results?

No benchmark rounds have been completed or published yet. The first step is recruiting reviewers and freezing the first task’s protocol. Results will include the brief, model and tool settings, resource limits, anonymized judgments, and the limits of a small, network-recruited pilot.

The frontier is moving.
Let’s measure what matters.

Help shape the first round