# HIWP Assessment and Capability Passport

## Assessment objective

Measure whether a specific human-plus-system configuration can deliver a specific occupational outcome, safely and repeatedly, with a known amount of human attention. Do not test AI trivia, prompt memorization, or theatrical autonomy.

## Assessment record

Every assessment is defined by:

- **Role slice:** a bounded, economically meaningful outcome.
- **Population:** the real task distribution the claim covers.
- **Context envelope:** data, policy, language, jurisdiction, systems, and risk.
- **Stack envelope:** allowed operator-layer and substrate classes.
- **Acceptance threshold:** externally meaningful quality bar.
- **Attention budget:** active operation, review, exception, and maintenance time.
- **Autonomy contract:** actions allowed without approval and required gates.
- **Validity rule:** expiry and material changes that trigger retesting.

## Six-part examination

### 1. Representative production

Randomly sample ordinary tasks. Score outputs blind where possible. Record accepted outcomes, error taxonomy, elapsed time, human attention, compute/tool cost, and interventions.

### 2. Controlled variation

Change format, length, vocabulary, source quality, priorities, or nonessential surface details. The goal is to distinguish an adaptable operator system from a memorized demo.

### 3. Novel in-scope task

Introduce a task that remains inside the role definition but requires the operator to revise their workflow. Evaluate adaptation time, quality loss, unnecessary complexity, and whether the operator can explain the new control strategy.

### 4. Exception and recovery

Inject missing inputs, conflicting instructions, tool failure, stale data, or an impossible objective. Score detection, halt/escalation choice, diagnostic accuracy, containment, time to safe state, and final outcome.

### 5. Adversarial and boundary test

Test prompt injection, unauthorized data request, false authority, fabricated evidence, unsafe external action, excessive permissions, policy collision, and pressure to bypass review. Test design must match the role’s risk.

### 6. Substitution and clean transfer

Replace one material model/tool where feasible, then rebuild the portable layer against synthetic or new-employer context. Score adaptation time, performance delta, residual dependencies, and protected-data leakage.

## Required comparison

For augmentation claims, use three conditions:

| Condition | Purpose |
|---|---|
| Conventional workflow | Establish existing professional baseline |
| Generic AI access | Isolate the value of merely having AI |
| Declared HCU | Measure the personal operator layer and operating skill |

Randomize order or disclose learning effects. The evaluator should score outputs without knowing the condition when practical.

## Performance vector

No universal total score is issued. The Capability Passport presents:

- quality: acceptance rate and distribution;
- reliability: variance and failure frequency;
- leverage: accepted outcomes per human hour relative to baselines;
- attention: median and tail active/review/exception minutes;
- intervention: detection and successful correction rates;
- recovery: containment and recovery-time distribution;
- safety: policy-event and severe-error rate;
- cost: total economic cost per accepted outcome;
- transfer: adaptation time and post-transfer quality delta;
- evidence: assessor independence, sample size, trace completeness, and recency.

## Guardrails against gaming

- Sample from a hidden task pool.
- Require signed version hashes before tasks are revealed.
- Report all attempted runs and justified exclusions.
- Separate training tasks from certification tasks.
- Rotate adversarial cases.
- Measure actual human attention; do not equate machine runtime with labor savings.
- Audit contamination between benchmark data and the operator stack.
- Revoke claims for forged receipts or undeclared material changes.
- Publish confidence intervals and failure counts, not only averages.

## Capability Passport display

The human-readable Passport should show:

1. **Who:** verified operator subject and relevant role.
2. **What:** exact capability claim.
3. **Where:** validated context envelope.
4. **How autonomous:** module-specific autonomy.
5. **How well:** performance vector.
6. **At what human cost:** attention and maintenance.
7. **What failed:** material limitations and severe-error history.
8. **Who checked:** issuer/evaluator and assurance method.
9. **When:** evaluation date, expiry, and material-change trigger.
10. **Proof:** selectively disclosed receipt references and status.

## Example role-slice rubric: project evidence synthesis

**Claim:** produce a source-linked statement of project reality from a mixed project archive.

| Criterion | Weight | Pass condition |
|---|---:|---|
| Claim-source entailment | 25% | ≥97% supported claims in blinded sample |
| Source-span precision | 15% | ≥95% citations resolve to exact supporting span |
| Contradiction detection | 15% | ≥90% seeded and naturally occurring conflicts surfaced |
| Current-state accuracy | 15% | ≥92% against adjudicated project state |
| Coverage | 10% | ≥90% of required decision-relevant facts |
| Unsupported assertion control | 10% | <1% material unsupported assertions |
| Human attention | 5% | declared target, reported not hidden |
| Recovery and auditability | 5% | all injected failures detected or safely escalated |

A role may require minimum thresholds in addition to weighted performance. A high average cannot compensate for a catastrophic privacy or safety failure.

## Issuance levels

- **Observed:** one assessor-witnessed demonstration; no reliability claim.
- **Validated:** repeated trials across representative and variation sets.
- **Stress-tested:** includes adversarial, recovery, and dependency substitution.
- **Portable:** includes a passed clean-room transfer audit.

These are evidence levels, not prestige ranks. A claim can be Portable but narrow, or Stress-tested but time-limited.
