# Human-Indexed Work Protocol (HIWP) v0.1

**Status:** public working draft for field validation  
**Date:** 15 August 2026  
**Scope:** declaration, assessment, evidence, portability, and governance of a Human Capability Unit (HCU)

## 1. Purpose

HIWP defines how an identified natural person and a personally operated AI-enabled system may be represented and evaluated as one work-capability unit. It is vendor-neutral and model-neutral. It does not certify a foundation model, grant professional licensure, override employment agreements, or allocate legal liability.

Normative terms **MUST**, **MUST NOT**, **SHOULD**, **SHOULD NOT**, and **MAY** are used in their ordinary standards sense.

## 2. Conformance objects

An HIWP implementation recognizes four signed objects:

1. **Capability Manifest** — the declared HCU and its conditions.
2. **Evidence Receipt** — the record of one assessment or production-validation run.
3. **Capability Claim** — a bounded assertion supported by a receipt set.
4. **Capability Passport** — a holder-controlled collection of current claims.

## 3. Identity and accountability

3.1 Every HCU MUST be indexed to exactly one natural-person principal for a given run.  
3.2 The principal MUST be able to identify the intended outcome, declared autonomy, escalation triggers, and stop mechanism.  
3.3 Accountability MUST NOT be assigned where the principal lacks material control, observability, competence, time, or authority.  
3.4 Teams MAY aggregate multiple HCUs, but evidence MUST preserve the contribution and authority boundary of each principal.  
3.5 A model, organization, or pseudonymous bot MUST NOT be represented as the human principal.

## 4. Five-layer declaration

Every manifest MUST declare:

### L1 — Human principal

- stable subject identifier;
- role and applicable professional constraints;
- competence prerequisites;
- identity-verification method or assurance level;
- declared responsibility and authority.

### L2 — Operator layer

- version identifier and integrity hash;
- interfaces and controls;
- routing, approval, escalation, and stop rules;
- evaluation suite identifier;
- ownership/licensing classification;
- protected-detail disclosure policy.

The manifest MAY describe proprietary internals by capability, hash, interface, and evaluation behavior. Publication of private prompts, source code, chain-of-thought, trade secrets, or employer data is not required.

### L3 — Capability modules

Each module MUST state its purpose, input and output class, tools, autonomy mode, known limitations, foreseeable failures, monitoring, and human escalation point.

### L4 — Work context

The manifest MUST identify the context class, governing policies, data sensitivity, permissions, jurisdictional or professional constraints, and the party controlling each contextual asset. Sensitive particulars MAY be redacted or represented by evaluator attestation.

### L5 — Compute substrate

The manifest MUST record material model and tool dependencies by product or capability class, version where meaningful, provider, and substitution tolerance. Third-party systems MUST be labeled licensed or externally governed rather than operator-owned.

## 5. Autonomy modes

- **A0 Manual:** AI advises or transforms content; every material action is performed by the human.
- **A1 Confirmed:** the system proposes each material action; the human approves before execution.
- **A2 Supervised:** the system executes bounded steps while the human actively monitors and can interrupt.
- **A3 Bounded autonomous:** the system executes a defined workflow and escalates exceptions; the human reviews at declared gates.
- **A4 Monitored autonomous:** the system operates continuously inside strict limits with retrospective and exception monitoring.

An HCU MUST declare autonomy per module and per consequential action class. One global label is insufficient. A3 or A4 MUST include machine-enforced boundaries, observable logs, stop controls, and tested rollback or containment. HIWP conformance is not, by itself, permission to use A3 or A4 in regulated or high-impact settings.

## 6. Risk classes

- **R0 Minimal:** reversible internal assistance with no sensitive data or consequential external effect.
- **R1 Low:** ordinary work product subject to routine human review.
- **R2 Moderate:** external communication, material business decisions, sensitive data, or actions with recoverable consequences.
- **R3 High:** employment, credit, health, legal, safety, critical infrastructure, or other rights- or safety-affecting uses.
- **R4 Prohibited by context:** a use barred by law, policy, license, professional duty, or explicit engagement rule.

Assessments MUST apply the highest material risk class present in the claimed workflow. R2 and above MUST include adversarial, privacy, security, and incident-response trials. HIWP v0.1 does not create a sufficient conformity route for R3 production deployment; applicable law and sector standards control.

## 7. Capability-claim grammar

A conforming claim MUST specify:

> Subject + operator-layer class/version + action/outcome + acceptance threshold + context + autonomy + resource limits + risk controls + validity window + evidence reference.

Example:

> Operator `did:example:7f3…`, using declared support-triage stack `2.x`, resolves English-language Tier-1 software support cases from the approved knowledge base at ≥92% blind-rubric acceptance, ≤4 minutes median human attention, A2 autonomy, with zero unreviewed external messages, under receipt set `ers:…`, valid through 2026-11-30 unless a material dependency changes.

Claims MUST NOT use unbounded language such as “expert AI user,” “can automate marketing,” or “10x worker.”

## 8. Assessment protocol

### 8.1 Preregistration

Before testing, the assessor MUST freeze:

- task population and sampling method;
- acceptance rubric and minimum threshold;
- baseline condition;
- allowed tools, data, time, cost, and autonomy;
- primary and guardrail metrics;
- material-change rules;
- evaluator identity and conflict disclosures.

### 8.2 Required trial families

A full assessment MUST include:

1. representative baseline tasks;
2. normal input variations;
3. at least one novel but in-scope task;
4. exception detection and recovery;
5. adversarial or policy-conflict inputs appropriate to risk;
6. one material dependency substitution where technically possible;
7. clean-room transfer for a portability claim.

### 8.3 Comparison conditions

Where the claim includes productivity or augmentation, testing SHOULD compare:

- conventional tools or unaided professional workflow;
- generic AI access without the personal operator layer;
- the declared HCU.

The same outcome rubric SHOULD be applied to all conditions. Learning and order effects MUST be controlled or disclosed.

### 8.4 Evidence sufficiency

One successful demonstration is insufficient for a reliability claim. Sample size MUST be justified based on outcome variability and risk. Exact sample sizes MAY vary by role, but the receipt set MUST report repetitions, failures, exclusions, and confidence intervals where quantitative inference is made.

## 9. Required measurement vector

HIWP forbids a universal composite score in v0.1. Reports MUST preserve at least these dimensions:

| Dimension | Minimum evidence |
|---|---|
| Outcome quality | rubric version, evaluator method, acceptance rate, distribution |
| Reliability | repetitions, variance, ordinary perturbations, temporal window |
| Human attention | active minutes, review minutes, exception minutes, measurement method |
| Throughput | accepted outcomes per wall-clock and per human hour |
| Intervention quality | detection rate, correction success, inappropriate override rate |
| Recovery | failure severity, containment, time to safe state, residual harm |
| Transferability | adaptation time, quality delta, prohibited-data audit |
| Cost | model/tool spend, operator time, maintenance, evaluation overhead |
| Risk compliance | policy violations, unsafe actions, privacy/security findings |
| Evidence integrity | trace completeness, signatures, conflicts, reproducibility |

Metrics MAY be supplemented by role-specific measures. A claim MUST NOT report throughput without the paired quality and human-attention measures.

## 10. Evidence receipt

Each run MUST produce an immutable receipt containing:

- receipt and protocol versions;
- subject, assessor, and task identifiers;
- manifest hash and material dependency versions;
- declared context, autonomy, and risk;
- start/end times and human-attention accounting;
- output hash or protected result reference;
- rubric results and evaluator attestations;
- interventions, exceptions, failures, and policy events;
- cost;
- exclusions and deviations;
- cryptographic proof or trusted-system signature.

Receipt publication MAY use hashes, pseudonymous identifiers, selective disclosure, and third-party attestations to protect identity and confidential content.

## 11. Capability Passport

11.1 The natural-person subject MUST control presentation of the Passport.  
11.2 The issuer controls claim issuance, suspension, revocation, and expiry.  
11.3 Passport formats SHOULD be compatible with W3C Verifiable Credentials 2.0 or an equivalent tamper-evident issuer-holder-verifier model.  
11.4 Selective disclosure SHOULD allow proof of a claim without revealing unrelated claims or private stack details.  
11.5 Every claim MUST expose its validity window, evidence class, task scope, and material-dependency policy.  
11.6 A Passport MUST NOT imply professional licensure or regulatory approval unless separately issued by the competent authority.

## 12. Material change and validity

A claim MUST be re-evaluated or downgraded after a material change. Materiality includes:

- a foundation-model family change affecting behavior;
- significant operator-layer logic change;
- added tool permissions or autonomy;
- new data-sensitivity or jurisdiction class;
- altered acceptance criteria;
- a serious incident indicating the claim is no longer reliable;
- a context shift outside validated bounds.

Minor provider updates MAY be covered by continuous canary evaluations if the assessor has declared acceptable drift limits.

## 13. Ownership schedule

Every employment or client deployment SHOULD attach an ownership schedule with three columns:

1. **Operator domain:** pre-existing or independently licensed methods and artifacts.
2. **Engagement domain:** employer/client data, credentials, integrations, policies, work product, and assigned improvements.
3. **Provider domain:** third-party models, tools, data, and services.

The schedule MUST identify carry-in assets and exit disposition. Capability evidence retained by the operator MUST be stripped of employer secrets and personal data not required for the claim.

## 14. Clean-room portability audit

A portability claim MUST be validated by:

1. revoking former-context access;
2. inventorying retained artifacts;
3. scanning for protected data, credentials, and proprietary identifiers;
4. rebuilding integrations against synthetic or new-context inputs;
5. rerunning a subset of the assessment;
6. recording adaptation time and quality delta;
7. obtaining operator and evaluator attestations.

Failure to transfer does not invalidate task capability, but MUST remove or qualify portability claims.

## 15. Organizational deployment controls

An employer using an HCU SHOULD provide:

- least-privilege, revocable access;
- a visible system inventory and responsible principal;
- time and authority for meaningful oversight;
- incident reporting without retaliation for good-faith interruption;
- compensated maintenance expectations;
- data-loss prevention and clean exit procedures;
- shared-learning mechanisms where appropriate;
- accessibility and equitable compute access.

The employer MUST NOT use HIWP to assign responsibility to a worker for system behavior the worker cannot observe or control.

## 16. Assessor requirements

Assessors MUST demonstrate role-domain competence, evaluation competence, impartiality controls, data-handling procedures, and repeatable scoring. They MUST disclose conflicts and separation between training and certification functions. A future formal scheme SHOULD be designed for compatibility with ISO/IEC 17024:2026, but HIWP v0.1 is not accredited under that standard.

## 17. Revocation and incident response

Claims MAY be suspended for expired evidence, unreported material change, integrity failure, or a serious incident. Revocation systems SHOULD protect holder privacy while allowing verifiers to determine current status. Incident review MUST distinguish operator error, employer-control failure, provider failure, malicious input, and unforeseeable interaction rather than collapsing all responsibility onto the principal.

## 18. v0.1 non-goals

HIWP v0.1 does not:

- rank all workers on one scale;
- certify consciousness, personality replication, or identity cloning;
- require disclosure of chain-of-thought;
- create a marketplace for selling autonomous agents apart from humans;
- guarantee job preservation;
- authorize regulated deployment;
- decide ownership disputes without a governing agreement;
- treat automation count as competence.

## 19. Protocol success criteria

HIWP should advance only if field trials show that its objects improve hiring prediction, deployment safety, portability clarity, or worker bargaining power beyond ordinary work samples and generic AI certificates. If employers cannot interpret the claims, assessors cannot reproduce them, or workers cannot transfer meaningful capability without leakage, the protocol must be revised or abandoned.
