Benchmark Receipts
Reproducible evidence, dependency-adjusted accounting, and honest non-results
A benchmark number without its execution context is a marketing fragment.
A Glyphd benchmark receipt records:
run_id
family_id
commit
artifact/model version
task or corpus digest
split and sampling
configuration
environment
dependencies
start/end
raw result refs
metrics
failures
limitations
acceptance
receipt hash
Dependency-adjusted accounting
Compression and cache experiments report:
- object-specific bytes;
- amortized shared dependencies;
- cold-start recoverability.
Model tests report weights, tokenizer, runtime, quantization, hardware, context, prompts, and scorer.
Evidence states
A result may be:
- reproduced;
- verified mechanism;
- source-reviewed;
- implemented;
- inconclusive;
- failed;
- invalid;
- awaiting local run.
Mock rows cannot become live results. Source inspection cannot become runtime PASS. A deployment cannot become a model benchmark.
Raw artifacts
Every aggregate links to raw rows. A reader can reconstruct the metric and identify excluded or failed cells.
Pre-registration
Important thresholds, baselines, sampling, and kill criteria are frozen before results are inspected. A failed result remains in the ledger.
No PASS badge exists without a receipt.
Source register
- N09-S01 — GlyphDrive Codec Benchmark Contract.
GLYPHDRIVE-CODEC-BENCHMARK-CONTRACT-v0.1.md - N09-S02 — SchemaStack Benchmark Plan.
V1_BENCHMARK_PLAN.md - N09-S03 — EL-JEFFE gate harness.
PHASE_0_VERIFICATION.md