← Field notes & protocolsVERIFIED MECHANISM

Benchmark Receipts

Reproducible evidence, dependency-adjusted accounting, and honest non-results

A benchmark number without its execution context is a marketing fragment.

A Glyphd benchmark receipt records:

run_id
family_id
commit
artifact/model version
task or corpus digest
split and sampling
configuration
environment
dependencies
start/end
raw result refs
metrics
failures
limitations
acceptance
receipt hash

Dependency-adjusted accounting

Compression and cache experiments report:

  1. object-specific bytes;
  2. amortized shared dependencies;
  3. cold-start recoverability.

Model tests report weights, tokenizer, runtime, quantization, hardware, context, prompts, and scorer.

Evidence states

A result may be:

  • reproduced;
  • verified mechanism;
  • source-reviewed;
  • implemented;
  • inconclusive;
  • failed;
  • invalid;
  • awaiting local run.

Mock rows cannot become live results. Source inspection cannot become runtime PASS. A deployment cannot become a model benchmark.

Raw artifacts

Every aggregate links to raw rows. A reader can reconstruct the metric and identify excluded or failed cells.

Pre-registration

Important thresholds, baselines, sampling, and kill criteria are frozen before results are inspected. A failed result remains in the ledger.

No PASS badge exists without a receipt.

Source register

  • N09-S01 — GlyphDrive Codec Benchmark Contract. GLYPHDRIVE-CODEC-BENCHMARK-CONTRACT-v0.1.md
  • N09-S02 — SchemaStack Benchmark Plan. V1_BENCHMARK_PLAN.md
  • N09-S03 — EL-JEFFE gate harness. PHASE_0_VERIFICATION.md