Objective
Determine whether a typed, recursive, language-addressable spatial intermediate representation can preserve correct bounded answers, reduce irrelevant spatial work, maintain deterministic revision history, normalize useful public projections, and return honest unknown or unsupported states without displacing the exact geometry and specialist solvers required by each task.
D25 now provides an implemented reference runtime and twelve deterministic application fixtures. Their passing state proves only fixture and contract conformance. No public latency, scaling, interoperability, geometric-accuracy, building-code, or production result exists merely because the runtime builds and its authored rows pass.
B18 R1 context-resolver result
The frozen seven-case context-construction comparison reached VERIFIED_MECHANISM for its hard safety scope and FAILED for its greenlight conclusion.
- required entity, relation, and metric-fact recall:
1.0 for SFT operational context;
- false authoritative executions when evidence was missing:
0;
- exact SFT/full-scene deterministic solver parity:
1.0;
- exact source/revision completeness:
1.0;
- median SFT serialized-context ratio to full scene:
0.785827;
- predeclared maximum ratio:
0.600000.
Because the efficiency gate failed, the frozen pinned-model holdout was not run. This result does not support a retrieval, Vector RAG, GraphRAG, model-answer, latency, scaling, or production-digital-twin advantage. The complete manifest, oracle, comparison contexts, raw rows, model protocol, limitations, losses, and receipt are published under /evidence/sft-b18-r1/.
Current implemented fixture layer
The repository currently exercises:
- exact canonical SHA-256 vectors;
- graph and frame validation;
- adaptive octree insertion and queries;
- bounded fit dispatch;
- authored architectural policy failure;
- service-access obstruction;
- assembly topological order;
- blocked and permitted reachability;
- explicit unsupported language;
- immutable branch isolation;
- clean two-parent merge-candidate creation with
MAIN preservation;
- five bounded interchange adapters;
- and synchronization of downloadable schemas and fixture receipts.
These checks produce DETERMINISTIC_FIXTURE_CONFORMANCE. They do not count as measured benchmark rows.
Benchmark groups
A. Exact fit and signed clearance
Generate versioned primitive, convex, and mesh cases involving objects, openings, containers, turns, and tolerance bands. Compare:
- the SFT bounded box kernel;
- analytic primitive ground truth;
- an independent geometry library;
- and, for mesh cases, a pinned exact or conservative collision/clearance implementation.
Required outputs include Boolean or unknown result, signed clearance interval, tolerance, orientation tested, exact source revision, selected kernel, and proof trace.
B. Collision candidate quality
Measure adaptive-octree broad-phase candidate reduction against brute-force pair enumeration and a conventional reference index using identical world-space bounds.
Report true overlap candidates, false positives, missed candidates, candidate pairs, narrow-phase calls, node count, retained-at-parent count, occupied depth, and update cost. No broad-phase overlap may be represented as proven surface contact.
C. Continuous motion and tunneling
Use swept paths that include high-speed tunneling, narrow gaps, turns, and moving obstacles. Compare discrete address-only checks with a pinned swept-volume or time-of-impact reference. Report missed collisions separately from latency.
D. Constraint and assembly correctness
Generate acyclic and cyclic containment, frame, attachment, and assembly graphs. Compare the validator and topological ordering with independently generated graph ground truth.
Report false acceptance, false refusal, cycle localization, ordering agreement, and unsupported mechanical facts.
E. Reachability and clearance
Generate typed spatial graphs with directed edges, blocked nodes, blocked edges, minimum-clearance attributes, alternate routes, and disconnected regions. Compare exact path existence and returned path with a reference graph implementation.
Robot motion-planning claims remain excluded unless a motion planner and geometry corpus are added explicitly.
F. Deterministic replay
For each initial scene root and event log, replay the scenario at least one hundred times across supported environments. Final canonical scene content, revision IDs, parent links, branch pointers, and merge-candidate IDs must match exactly. Interface layout and rendering state are excluded from the authoritative hash.
G. Counterfactual branch isolation
Fork historical revisions, apply disjoint and overlapping changes, and verify:
- original branches remain byte-identical;
- changed IDs are isolated correctly;
- frame, entity, relation, and envelope conflicts are detected;
- clean candidates have both branch heads as parents;
- and
MAIN never advances without an explicit authority action.
H. Sparse scaling
Increase total represented world volume while holding occupied content constant, then increase occupied content at fixed extent. Compare adaptive octree, brute-force scan, and a conventional reference index.
Report memory, node count, insertion latency, update latency, point-query latency, range-query latency, candidate reduction, retained-at-parent count, and serialized retrieval size. Record complete latency distributions rather than one decorative average.
I. Spatial retrieval for language models
Compare full-scene serialization with task-specific graph and spatial-branch retrieval for the same pinned question set. Preserve model, prompt, temperature, compiler policy, scene revision, and retrieved state.
Report retrieved bytes or tokens, latency, typed-command accuracy, exact-answer accuracy, unsupported-answer rate, ambiguity calibration, and refusal/unknown calibration. A model-generated command must pass the same deterministic schema and entity-resolution gates as the rule compiler.
J. Resolution stability
Refine unrelated spatial branches and verify that an unchanged local query does not reverse. Refine the queried branch and record when tighter bounds legitimately change UNKNOWN into PASS or FAIL.
K. Import normalization fidelity
For each adapter, compare the normalized graph with a public source-specific golden projection.
- glTF: hierarchy, TRS, semantic extras, authored bounds;
- OpenUSD subset: prim hierarchy, translation, extent, cube size;
- CAD manifest: assemblies, parts, frames, dependencies, bounds;
- BIM/IFC projection: containment, spaces, openings, services, connectivity;
- robot scene graph: links, joints, obstacles, path nodes, blocked and clearance state.
Report preserved objects, preserved relations, unit/frame error, ignored feature classes, source-pointer completeness, and false precision. Native interchange claims require native parsers and independent corpora.
L. Architectural clearance and service access
Use independently authored project-policy fixtures, then real public-safe project extracts where available. Compare each rule result with a pinned calculation and expert review.
Keep three authorities distinct:
- geometric or interval calculation;
- authored project policy;
- jurisdictional code or professional approval.
The benchmark must never promote the first or second into the third.
M. Failure honesty
Construct cases with missing dimensions, ambiguous names, unsupported language, ambiguous frames, unsupported materials, absent tolerance, unavailable solver class, native-format features outside an adapter, invalid revision ancestry, and overlapping agent changes.
Score exact use of PASS, FAIL, UNKNOWN, INVALID, UNSUPPORTED, AMBIGUOUS, and explicit adapter warnings.
Required comparisons
At minimum compare against:
- analytic primitive calculations;
- brute-force collision-pair enumeration;
- a conventional spatial index using identical authoritative bounds;
- an independent graph and topological-order implementation;
- an independent geometry or CAD kernel for geometric cases;
- source-specific golden import projections;
- full-scene language serialization;
- task-specific branch retrieval;
- and a semantic graph with metric facts withheld.
Metrics
Report separately:
- exact-answer accuracy;
- signed-clearance interval error;
- unknown-state precision and recall;
- unsupported-state precision and recall;
- missed-collision-candidate rate;
- broad-phase false-positive rate;
- narrow-phase call reduction;
- path-existence accuracy;
- path agreement;
- assembly-order agreement;
- cycle-detection accuracy;
- replay content-ID agreement;
- revision-parent agreement;
- branch-isolation failures;
- merge-conflict false positive and false negative rates;
- import entity and relation preservation;
- frame and unit conversion error;
- source-pointer completeness;
- memory per occupied entity;
- index insertion and query distributions;
- retrieved bytes or tokens;
- typed-command compile accuracy;
- ambiguity and refusal calibration;
- resolution-reversal count;
- and implementation/environment details.
Invalid conditions
A run is invalid when:
- interface card coordinates are treated as physical coordinates;
- recursive address overlap is labeled exact collision;
- semantic-only cases receive hidden metric facts;
- renderer output enters the authoritative world hash;
- an authored project policy is labeled building-code approval;
- native interchange completeness is claimed from a projection manifest;
- a model answer is scored without preserving prompt and retrieved state;
- failed, unknown, ambiguous, or unsupported rows are removed;
- deterministic replay permits unrecorded randomness;
- a merge candidate is counted as accepted state;
- fixture conformance is promoted as measured spatial accuracy;
- or implementation-specific shortcuts are omitted.
VERIFIED_MECHANISM requires a versioned implementation, public fixture corpus, independent reference calculations, raw rows, complete failures, environment record, exact source commit, invalid-run checks, limitations, and a receipt.
A performance, interoperability, architectural, robot, or product claim requires a specifically scoped dataset, repeated runs, and independent review. Negative and inconclusive results remain part of the evidence ledger.