← Back to the experiment

Fair conditions.
Open questions.

PROPOSED PILOT PROTOCOL · VERSION 0.1 · OCTOBER 9, 2026

We’re testing the contribution of an expert directing AI, compared with an agent working autonomously. Human reviewers decide which final work they would prefer to use.

This is the pilot’s design, not a report of completed experiments. Each round’s exact task, models, resource limits, editing treatment, and scoring rules will be frozen and published before work begins.

1. Define useful work before generating it.

Invite reviewers with relevant professional experience. Ask what decision the output needs to support, what makes it useful, which plausible mistakes would make it unacceptable, and which tradeoffs matter. De-identify the responses and share the same packet with both contestants.

Each contestant gets the same allowance for clarification questions. Pool the answers into a shared addendum and freeze it before the clock starts. Reviewers do not coach either contestant or provide draft feedback.

2. Give both contestants the same opportunity.

The main comparison is an expert using AI versus an autonomous agent. Both receive the same task, source material, reviewer criteria, tools, model access, wall-clock limit, and declared tool/model spending cap. We report actual time, tool use, and spending as well as the permitted limits. The human’s prior experience is the contribution under examination.

The agent may plan, search, generate multiple drafts, check its work, and revise within those limits. It must select its own final submission through a frozen process. The organizer does not cherry-pick its best answer. An optional one-shot baseline would be labeled separately and would not replace the primary comparison.

3. Reduce author cues without silently improving content.

Remove names and identifying metadata, normalize layout, and assign A/B order randomly for each reviewer. For professional judgment tasks, we plan to evaluate a common editing treatment on both submissions before using it. Freeze the editor, instructions, and settings, retain originals, and check that facts, citations, caveats, errors, and recommendations are preserved. If that cannot be achieved, report the limitation and use the original text with normalized formatting.

Creative tasks are judged in their original voice because style may be the skill at issue. An edited comparison, if included, is a separate condition. No humanizer has yet been validated or selected.

4. Ask for a preference and its reasons.

Reviewers read both outputs without author labels and record one of four outcomes:

Collect the reason, confidence, and material errors. Lock preference before asking which output the reviewer thinks came from the expert, with “unsure” available. This gives us an indication of blinding quality; it is not proof that authorship was concealed.

Once assigned a paid review, a reviewer receives the agreed $50 for completing it regardless of preference or criticism. Payment arrangements and the meaning of a completed review are confirmed before participation.

5. Keep the first pilot small and legible.

Begin with one task and eight paid reviews. If the process works, expand to four task pairs with eight reviews each. The planned $2,000 allocation is $1,600 for those 32 reviews, $200 for four follow-up reviews, and $200 for operating costs. Domain and hosting costs come from the operating reserve; any expansion depends on the remaining budget.

Shane Francis is the organizer and first expert participant. Reviewers may be recruited through existing professional relationships. Disclose relevant relationships, keep judgments independent, and treat this as an exploratory, network-recruited pilot. Multiple reviews of the same task do not constitute independent task replications.

6. Publish what happened, including the limits.

For each completed round, report the brief, shared context, dated model versions, agent instructions, allowed and actual resources, final outputs, any editing, preference counts, reasons, and failures. Report ties and “neither acceptable” separately. Publish anonymized reviewer material only with the permission agreed before participation.

Small samples, task selection, familiarity with the organizer, and writing cues limit what we can infer. A result describes these tasks and conditions; it does not establish universal human or AI superiority. The point is to build a comparable record as the frontier moves.

What the website does today

This release explains the pilot, demonstrates an anonymous ballot using illustrative samples, and lets prospective reviewers express interest by email. The demo does not collect or store votes. Real study ballots, participant consent, and result publication will be added for the first live round. The pilot does not accept wagers, collect entry fees, or hold participant funds.

Bring your judgment to the first round →