Status as of 18 August 2026
An Externalised, Citation-Enforced Relational-State Layer Does Not Lift a Frontier Model Into Register Band
Working title; re-scoped after a partially overlapping external result. Data collection in progress.
Raghav Sharma · Gaurav Sharma · Aryan Tiwari · Vyakti.ai
Unpublished draft
No submission target. There is no draft to submit.
This paper is genuinely incomplete, not merely unannounced. The central comparison arm is at roughly 3% of its target size (74 of a planned 2,304 calls) and is rate-limited to ~75 calls/day on its current pool, which puts full data collection roughly a month out absent additional compute. No headline number for this paper is final, and nothing on its page is a finding.
Abstract
This paper is in preparation and its data is incomplete, so no results are published yet. The question it is built to answer: if you give a different underlying AI model the exact same compiled memory, relationship history and context as an existing one — byte-for-byte identical input — does the companion still feel like the same person? Early, partial data suggests the answer is no by default: swapping the model changes surface behaviour (word choice, question-asking habits, and in one measured case, an unwanted script-mixing failure) even when every other input is held fixed. A related, larger independent study published elsewhere in 2026 found a similar pattern, so this paper's contribution is narrower and sharper than originally planned: whether a specific engineering approach — externalised memory enforced by citations rather than free text — narrows that gap. We will publish honestly whichever way the data comes out.
The identity-ceiling question: does an externalised, citation-enforced relational-state layer — compiled byte-identically into both arms by a production context compiler, rather than a generated memory scaffold — narrow the register gap when the serving model is swapped? A companion charm-grok bake-off previously showed the incumbent model beating a swap candidate 38–2 on blind, counterbalanced, human-decision-grade verdicts; that finding, and its ANCHOR-adjacent generalisation (Venkit et al., arXiv:2607.28818, 2,008 conversations across 3 memory architectures × 4 models, memory scaffold identity not moving persona-collapse pattern) partially scoop the original framing, which this paper is re-scoped around. Status: the compiled-context corpus exists (2,304 sha256-distinct, byte-identity-by-construction contexts) and the raw candidate arm is fully generated (2,304/2,304 non-empty replies) with an early, unadjudicated surface-register signal — 7 raw Devanagari-script hits against a hard-fail axis, mean 15.1 words/turn against the incumbent's 20.5-word band centre. The incumbent arm needed for the primary comparison is at 74/2,304 calls, gated by a free-tier pool ceiling of ~75 calls/day; the qualified judge needed for the D2 relational-axis gate is blocked on Paper B's own subject matter (no credit-billed judge cleared qualification). Both blockers are logged and dated, not glossed over.
Where this paper actually is
- Primary comparison arm
- 74 of 2,304 calls
- 3%
- Candidate arm
- 2,304 of 2,304 calls
- complete
- Rate limit
- ~75 calls/day on the current pool
Not a finding
7 raw Devanagari-script hits on the swap candidate's surface output
The raw candidate model, given the same compiled context as the incumbent, produced native-script characters against a hard-fail axis (any hit fails) in early generation. This is an unadjudicated data point, not a scored result, since the confirmatory incumbent-side comparison has not yet been generated.
Raw and unscored. No adapter or relational-layer mitigation was active in this run by design; this measures the model swapped in with nothing added, the paper's baseline condition, not its test condition.
context/measurements.md terra-arm-2304
Why the scope changed
The original framing was broader: that a swapped model, given the right context, would land inside an acceptable behavioural range regardless of which model it was. Our own earlier measurement had already made that claim harder to defend: an incumbent-against-candidate bake-off on this exact product found the incumbent winning 38 judged comparisons to 2. Then a larger, independently published study (2,008 conversations, three memory architectures crossed with four models) found the same shape of result: the underlying model, not the memory scaffold around it, sets the ceiling on whether a companion's persona survives. That result is a bigger, better-resourced version of a claim we had already measured, which means we no longer get to publish it as a first report.
What survives, and what has not been scooped, is a narrower and more specific claim: whether a context that is compiled byte-identically into both arms by our production engine, rather than a generated memory narrative fed to the model, changes the picture. That is an engineering claim about a specific architecture, not a claim about model swapping in general, and it is the paper we are now building.
What exists today
A compiled-context corpus of 2,304 distinct, hash-verified contexts, built so that both the incumbent and the candidate model are served byte-identical input by construction. The candidate arm is fully generated against that corpus: 2,304 of 2,304 replies, no errors.
One early, unadjudicated observation came out of that raw generation: seven instances of native-script (Devanagari) characters against an axis this product treats as a hard failure at any occurrence, in a model given no memory-layer mitigation by design. That run measures the swap with nothing added yet, which is the paper's baseline condition rather than its test condition. It is quarantined below, at the size an unscored count deserves.
What is missing before there is a paper
The matched incumbent-side data at the same scale, and a qualified judge for the relational-preference axes of the comparison. The second gap is, concretely, blocked on Paper B's own subject matter: no credit-billed candidate judge has yet cleared qualification for this kind of task.
Both blockers are logged with dates rather than described as almost finished. When the primary arm completes, this page gains findings; until then it has a question, a corpus, and two reasons the answer does not exist yet.
What this does not show
There are no findings on this page and no figure, because the primary comparison arm does not exist yet. The one observation shown is raw and unadjudicated, taken from a run with no relational-layer mitigation active by design, which is the paper's baseline condition rather than its test condition. The second blocker is the subject matter of the other paper: no credit-billed judge has cleared qualification for the relational axes this comparison needs.
Artifacts
Every slot renders whether or not it is filled, because what does not exist yet is part of the status.
Not written; data collection incomplete.
arXiv
Not applicable. The paper is pre-draft.
Code
Not applicable yet.
Benchmark
Not applicable yet.
Cite
Not citable yet. This page will carry a BibTeX entry when there is a paper to cite.
Page last updated 18 August 2026