← Benchmark centerVERSIONED EVIDENCE CONTRACT

EL-JEFFE 4B Model and Runtime Gates

No public result exists merely because this contract exists. It defines what must be run, recorded, and disclosed before a result can be promoted.

Read related paper
REQUIRED SUITES
  1. Authority, modality, scope, and substitution
  2. Constraint retention
  3. Evidence discipline
  4. Routing, abstention, and escalation
  5. Structured action generation
  6. Private retrieval
  7. Mobile latency, memory, thermals, and battery
  8. Independent reproduction
INVALID-RUN CONDITIONS
  • example synthetic metrics called achieved
  • weights or tokenizer unpinned
  • runtime policy credited to checkpoint
  • unauthorized action
  • comparative claim without independent run
METRICS
authority_accuracysubstitution_errorconstraint_retentionevidence_accuracyrouting_accuracyunauthorized_actionslatencymemoryenergythermal_sustain
EXAMPLE RECEIPT · NOT MEASURED
{
  "run_id": "b06-jeffe-4b-example-NOT-MEASURED",
  "suite_id": "b06-jeffe-4b",
  "suite_version": "1.0.0",
  "commit": "UNSET",
  "configuration_digest": "UNSET",
  "corpus_or_task_digest": "UNSET",
  "model_and_runtime": {},
  "environment_digest": "UNSET",
  "started_at": null,
  "finished_at": null,
  "raw_artifact_refs": [],
  "metric_records": [],
  "invalid_conditions_checked": [],
  "known_limitations": [
    "Example only; no measured values."
  ],
  "result_state": "AWAITING_LOCAL_RUN",
  "receipt_digest": "UNSET",
  "example_only": true
}
CI FIXTURE RECEIPTDETERMINISTIC FIXTURE CONFORMANCE

Contract structure and example-receipt invariants only. No metric was measured and no benchmark result was promoted.

{
  "checks": {
    "example_cannot_be_measured": true,
    "invalid_conditions_declared": true,
    "metrics_declared": true,
    "receipt_contract_complete": true,
    "receipt_suite_matches": true,
    "related_paper_present": true,
    "required_suites_declared": true,
    "suite_id_matches_directory": true,
    "suite_version_present": true
  },
  "contract_sha256": "fdd6dd1183866bfec2d8d7390dd0ddf170281be6730c2066d99282b0c7f62867",
  "measured_metrics": false,
  "note": "This receipt validates the benchmark contract fixture only. It is not a benchmark run and contains no performance or comparative result.",
  "performance_result": false,
  "result_state": "DETERMINISTIC_FIXTURE_CONFORMANCE",
  "suite_id": "b06-jeffe-4b",
  "suite_version": "1.0.0"
}
FOUNDER-LOCAL RUN PACKETS

These one-command packets fail closed when hardware, source, model, dataset, participant consent, or checkpoint prerequisites are absent.

JEFFE-4B-CANDIDATE-GATES · AWAITING_LOCAL_RUN
RELATED IMPORTED EVIDENCE

These records retain their own scope and limitations. They do not cause every suite in this broader contract to pass.

E-EL-JEFFE-PHASE0IMPLEMENTED

EL-JEFFE 4B specification and Phase 0 control package

241-file specification/build-control package with schemas, gates, synthetic tests, and Phase 0 verification.

  • No trained checkpoint.
  • No model-quality gate has passed.
  • No production mobile runtime or comparative result is claimed.
README.mdPHASE_0_VERIFICATION.md

Objective

Establish a reproducible, falsifiable evidence path for p15-el-jeffe-4b. This contract authorizes no result by itself. A public result requires a versioned run, exact configuration, raw artifacts, environment identity, limitations, and a receipt.

Required suites

  1. Authority, modality, scope, and substitution
  2. Constraint retention
  3. Evidence discipline
  4. Routing, abstention, and escalation
  5. Structured action generation
  6. Private retrieval
  7. Mobile latency, memory, thermals, and battery
  8. Independent reproduction

Invalid-run conditions

A run is invalid, rather than merely low-scoring, when any of the following occurs:

  • example synthetic metrics called achieved
  • weights or tokenizer unpinned
  • runtime policy credited to checkpoint
  • unauthorized action
  • comparative claim without independent run

Metrics

Report each metric separately. Do not hide a regression inside one composite score.

  • authority_accuracy
  • substitution_error
  • constraint_retention
  • evidence_accuracy
  • routing_accuracy
  • unauthorized_actions
  • latency
  • memory
  • energy
  • thermal_sustain

Baseline discipline

Use the strongest sensible baseline for the exact task and information budget. Pin versions, prompts, corpora, models, hardware, runtime flags, and warm/cold state. If the research system receives additional context, dependencies, precomputation, or a model, charge or disclose them under the contract.

Run receipt

Every run emits:

run_id
suite_id
suite_version
commit
configuration_digest
corpus_or_task_digest
model_and_runtime
environment_digest
started_at
finished_at
raw_artifact_refs[]
metric_records[]
invalid_conditions_checked[]
known_limitations[]
result_state
receipt_digest

Allowed result states:

REPRODUCED
VERIFIED_MECHANISM
IMPLEMENTED
SOURCE_REVIEWED
SPECIFIED
HYPOTHESIS
INCONCLUSIVE
FAILED
AWAITING_LOCAL_RUN
INVALID

Data splits and leakage

Freeze task sets before viewing final results. Keep development, selection, and held-out evaluation separate where the suite supports them. Record any human-authored fixture that appeared during implementation. Do not tune a threshold on the final set and present it as held out.

Statistical and qualitative reporting

For repeated measurements, report sample count, central tendency, dispersion, and uncertainty appropriate to the design. Preserve row-level results. For human ratings, blind condition labels where possible, record the rubric, and retain disagreement rather than averaging away failure modes.

Negative evidence

Failed and inconclusive runs remain in the experiment ledger with code, configuration, and reason. A later successful run does not delete them.

Publication requirements

A public result page must show:

  • exact evidence state;
  • date and commit;
  • environment and task/corpus identity;
  • metric table;
  • baseline table;
  • raw artifact links;
  • failure gallery;
  • limitations and non-claims;
  • reproduction command;
  • receipt digest.

No marketing page may round a measured result into a broader claim than the contract supports.