# 90-Day Human-Indexed Work Field Pilot

## Purpose

Determine whether personally operated AI stacks create durable, transferable capability beyond generic access to AI, and whether HIWP evidence improves employer decisions.

## Recommended first role slice

Choose one high-frequency, low-to-moderate-risk knowledge-work outcome with objective review. Good candidates include project-evidence synthesis, internal support triage, structured proposal production, software issue diagnosis, or catalog-content operations. Avoid hiring, clinical, legal, credit, safety, and other high-impact decisions in the first pilot.

## Participants

- 8–20 experienced practitioners;
- 2–4 participating employers or one employer with multiple teams;
- independent task authors and evaluators;
- one protocol steward and one security/privacy reviewer.

## Conditions

Each participant completes balanced tasks under:

1. conventional tools;
2. employer-provided generic AI;
3. their personal operator stack;
4. a substituted model/tool condition;
5. a clean-room target context.

## Primary measures

- blinded outcome acceptance;
- accepted outcomes per human hour;
- median and 95th-percentile human attention;
- serious-error and policy-event rate;
- exception detection and successful recovery;
- total cost per accepted outcome;
- model-substitution quality delta;
- clean-room adaptation time and leakage findings.

## Secondary measures

- operator workload and perceived control;
- employer willingness to interview, hire, or pay for the verified capability;
- evaluator agreement;
- maintenance time across 30 days;
- degree to which gains are operator-specific versus tool-access-specific;
- team knowledge-sharing effects.

## Timeline

### Days 1–15: role and protocol freeze

- Select the role slice and task population.
- Define blinded acceptance rubric and severe-error taxonomy.
- Inventory legal, security, privacy, and IP constraints.
- Freeze manifests, baselines, attention accounting, and analysis plan.
- Create synthetic transfer context and hidden test pool.

### Days 16–30: onboarding and dry run

- Verify participants and declare stacks.
- Separate operator, employer, and provider domains.
- Instrument human attention and receipts.
- Run non-scored calibration tasks.
- Fix protocol defects, then freeze v0.1-pilot.

### Days 31–60: controlled assessment

- Execute balanced comparison conditions.
- Run variation, novel task, exception, and adversarial trials.
- Record every attempted run and exclusion.
- Blind output scoring where possible.

### Days 61–75: substitution and transfer

- Replace one material dependency.
- Revoke original workplace access.
- Audit retained artifacts.
- Rebuild against clean data and rerun a test subset.

### Days 76–90: analysis and employer validation

- Compute preregistered results and uncertainty.
- Issue experimental Passports.
- Test whether employers interpret and use them correctly.
- Publish protocol, aggregate results, negative findings, and next-version decisions.

## Decision gates

Proceed to a larger pilot only if:

- the HCU condition materially improves accepted output per human hour over generic AI access;
- quality and serious-error guardrails hold;
- attention accounting is reliable;
- at least a useful subset passes clean transfer;
- employers find Passport evidence more decision-useful than a résumé and generic certificate;
- workers can preserve capability evidence without retaining protected employer content.

Pause or redesign if gains are primarily benchmark leakage, uncounted review labor, one vendor’s temporary advantage, or employer-specific data that cannot be separated.

## Minimum deliverables

- preregistration;
- task and rubric specification;
- signed manifests and evidence receipts;
- anonymized analysis dataset;
- incident and exclusion log;
- sample Capability Passports;
- clean-room audit reports;
- public findings memo, including null and negative results;
- HIWP v0.2 change log.
