Loading the institute record…
Loading the institute record…
No public result exists merely because this contract exists. It defines what must be run, recorded, and disclosed before a result can be promoted.
exact_answer_accuracysigned_clearance_interval_errorunknown_state_precisionunknown_state_recallunsupported_state_precisionunsupported_state_recallmissed_collision_candidate_ratebroad_phase_false_positive_ratenarrow_phase_call_reductionpath_existence_accuracypath_agreementassembly_order_agreementcycle_detection_accuracyreplay_content_id_agreementrevision_parent_agreementbranch_isolation_failuresmerge_conflict_false_positive_ratemerge_conflict_false_negative_rateimport_entity_preservationimport_relation_preservationframe_and_unit_conversion_errorsource_pointer_completenessmemory_per_occupied_entityindex_insertion_latency_distributionpoint_query_latency_distributionrange_query_latency_distributionretrieved_bytes_or_tokenstyped_command_compile_accuracyambiguity_calibrationrefusal_calibrationresolution_reversal_count{
"run_id": "b18-semantic-fractal-topology-example-NOT-MEASURED",
"suite_id": "b18-semantic-fractal-topology",
"suite_version": "2.0.0",
"commit": "UNSET",
"deployment": "UNSET",
"configuration_digest": "UNSET",
"corpus_or_task_digest": "UNSET",
"model_and_runtime": {
"spatial_runtime": "UNSET",
"geometry_reference": "UNSET",
"graph_reference": "UNSET",
"language_model": "UNSET",
"retrieval_policy": "UNSET",
"import_adapter": "UNSET"
},
"environment_digest": "UNSET",
"started_at": null,
"finished_at": null,
"raw_artifact_refs": [],
"metric_records": [],
"invalid_conditions_checked": [],
"known_limitations": [
"Example only; no measured spatial-runtime values.",
"The public D25 runtime and twelve application rows are deterministic fixtures, not benchmark measurements.",
"Exact geometry, specialist engineering solvers, code authorities, and human acceptance remain authoritative for their respective claims."
],
"result_state": "AWAITING_LOCAL_RUN",
"receipt_digest": "UNSET",
"example_only": true
}Contract structure and example-receipt invariants only. No metric was measured and no benchmark result was promoted.
{
"checks": {
"example_cannot_be_measured": true,
"invalid_conditions_declared": true,
"metrics_declared": true,
"receipt_contract_complete": true,
"receipt_suite_matches": true,
"related_paper_present": true,
"required_suites_declared": true,
"suite_id_matches_directory": true,
"suite_version_present": true
},
"contract_sha256": "d9110a49306077962c4a79b086f8c0ad888bc73b90e5f38b5f1018f3a67adbca",
"measured_metrics": false,
"note": "This receipt validates the benchmark contract fixture only. It is not a benchmark run and contains no performance or comparative result.",
"performance_result": false,
"result_state": "DETERMINISTIC_FIXTURE_CONFORMANCE",
"suite_id": "b18-semantic-fractal-topology",
"suite_version": "2.0.0"
}Determine whether a typed, recursive, language-addressable spatial intermediate representation can preserve correct bounded answers, reduce irrelevant spatial work, maintain deterministic revision history, normalize useful public projections, and return honest unknown or unsupported states without displacing the exact geometry and specialist solvers required by each task.
D25 now provides an implemented reference runtime and twelve deterministic application fixtures. Their passing state proves only fixture and contract conformance. No public latency, scaling, interoperability, geometric-accuracy, building-code, or production result exists merely because the runtime builds and its authored rows pass.
The repository currently exercises:
MAIN preservation;These checks produce DETERMINISTIC_FIXTURE_CONFORMANCE. They do not count as measured benchmark rows.
Generate versioned primitive, convex, and mesh cases involving objects, openings, containers, turns, and tolerance bands. Compare:
Required outputs include Boolean or unknown result, signed clearance interval, tolerance, orientation tested, exact source revision, selected kernel, and proof trace.
Measure adaptive-octree broad-phase candidate reduction against brute-force pair enumeration and a conventional reference index using identical world-space bounds.
Report true overlap candidates, false positives, missed candidates, candidate pairs, narrow-phase calls, node count, retained-at-parent count, occupied depth, and update cost. No broad-phase overlap may be represented as proven surface contact.
Use swept paths that include high-speed tunneling, narrow gaps, turns, and moving obstacles. Compare discrete address-only checks with a pinned swept-volume or time-of-impact reference. Report missed collisions separately from latency.
Generate acyclic and cyclic containment, frame, attachment, and assembly graphs. Compare the validator and topological ordering with independently generated graph ground truth.
Report false acceptance, false refusal, cycle localization, ordering agreement, and unsupported mechanical facts.
Generate typed spatial graphs with directed edges, blocked nodes, blocked edges, minimum-clearance attributes, alternate routes, and disconnected regions. Compare exact path existence and returned path with a reference graph implementation.
Robot motion-planning claims remain excluded unless a motion planner and geometry corpus are added explicitly.
For each initial scene root and event log, replay the scenario at least one hundred times across supported environments. Final canonical scene content, revision IDs, parent links, branch pointers, and merge-candidate IDs must match exactly. Interface layout and rendering state are excluded from the authoritative hash.
Fork historical revisions, apply disjoint and overlapping changes, and verify:
MAIN never advances without an explicit authority action.Increase total represented world volume while holding occupied content constant, then increase occupied content at fixed extent. Compare adaptive octree, brute-force scan, and a conventional reference index.
Report memory, node count, insertion latency, update latency, point-query latency, range-query latency, candidate reduction, retained-at-parent count, and serialized retrieval size. Record complete latency distributions rather than one decorative average.
Compare full-scene serialization with task-specific graph and spatial-branch retrieval for the same pinned question set. Preserve model, prompt, temperature, compiler policy, scene revision, and retrieved state.
Report retrieved bytes or tokens, latency, typed-command accuracy, exact-answer accuracy, unsupported-answer rate, ambiguity calibration, and refusal/unknown calibration. A model-generated command must pass the same deterministic schema and entity-resolution gates as the rule compiler.
Refine unrelated spatial branches and verify that an unchanged local query does not reverse. Refine the queried branch and record when tighter bounds legitimately change UNKNOWN into PASS or FAIL.
For each adapter, compare the normalized graph with a public source-specific golden projection.
Report preserved objects, preserved relations, unit/frame error, ignored feature classes, source-pointer completeness, and false precision. Native interchange claims require native parsers and independent corpora.
Use independently authored project-policy fixtures, then real public-safe project extracts where available. Compare each rule result with a pinned calculation and expert review.
Keep three authorities distinct:
The benchmark must never promote the first or second into the third.
Construct cases with missing dimensions, ambiguous names, unsupported language, ambiguous frames, unsupported materials, absent tolerance, unavailable solver class, native-format features outside an adapter, invalid revision ancestry, and overlapping agent changes.
Score exact use of PASS, FAIL, UNKNOWN, INVALID, UNSUPPORTED, AMBIGUOUS, and explicit adapter warnings.
At minimum compare against:
Report separately:
A run is invalid when:
VERIFIED_MECHANISM requires a versioned implementation, public fixture corpus, independent reference calculations, raw rows, complete failures, environment record, exact source commit, invalid-run checks, limitations, and a receipt.
A performance, interoperability, architectural, robot, or product claim requires a specifically scoped dataset, repeated runs, and independent review. Negative and inconclusive results remain part of the evidence ledger.