Deterministic compiler and fail-closed local read boundary
The checked-in conformance run admitted 40/40 valid controls, refused 400/400 scope-fault calls before scoring, and refused 9/9 injected mutants without returning a payload.
This confirms local conformance for the tested cut. It does not establish retrieval quality, cryptographic authenticity, exhaustive security, or production durability.Static additive-j retrieval improved a closed 300-row ladder
Across three checked-in MuSiQue/2Wiki runs, HSWM measured +0.0364 support recall@3, +0.0259 nDCG@10, and +0.0729 downstream answer F1 over cosine.
HSWM used 100 offline LLM judgments per run. The ladder omitted strong late-interaction and production graph retrievers, while cosine still led hit@3 and MRR; no state-of-the-art claim follows.General cognitive uplift over direct LLM reranking did not replicate
The preregistered cross-dataset criterion failed: pooled HSWM minus direct-LLM answer F1 was -0.1489; the positive 2Wiki delta was not significant at the stored paired-bootstrap threshold.
The supported statement is retrieval improvement over the listed lightweight baselines on this ladder, not a smarter reasoner.The tested query-time graph traversal path is deployment-OFF
MuSiQue and 2Wiki certificates selected mu=0, and none of nine tested traversal settings beat the static field on hop-drop.
The implementation falls back to the same snapshot's static field. This is not evidence of successful graph reasoning.Stale suppression works pointwise; the broader architectural novelty did not
A full-dose non-destructive supersede write removed the maximally confusable stale fact from top-10 in the two checked-in datasets while preserving it for audit.
An external graded revision arm was bit-exact, and a wrong write reduced current recall by 12.69 points on MuSiQue and 31.0 points on 2Wiki. Durable correction and replay are still required.Durable multi-writer memory remains an open systems problem
Crash-safe event replay, compensation, concurrent publication, signatures, and external trust distribution are not present-tense capabilities.
The public implementation is a compiler boundary and measurement prototype, not a production memory service.Real book-scale advantage is unmeasured
A synthetic experiment establishes mechanism sufficiency when judge-readable aboutness survives vector dilution.
No real NoCha, QASPER, NarrativeQA, or book-scale result has landed, so the website makes no long-document efficacy claim.