Skip to content
In preparation

Status as of 18 August 2026

An Externalised, Citation-Enforced Relational-State Layer Does Not Lift a Frontier Model Into Register Band

Working title; re-scoped after a partially overlapping external result. Data collection in progress.

Raghav Sharma · Gaurav Sharma · Aryan Tiwari · Vyakti.ai

Unpublished draft

No submission target. There is no draft to submit.

This paper is genuinely incomplete, not merely unannounced. The central comparison arm is at roughly 3% of its target size (74 of a planned 2,304 calls) and is rate-limited to ~75 calls/day on its current pool, which puts full data collection roughly a month out absent additional compute. No headline number for this paper is final, and nothing on its page is a finding.

Abstract

This paper is in preparation and its data is incomplete, so no results are published yet. The question it is built to answer: if you give a different underlying AI model the exact same compiled memory, relationship history and context as an existing one — byte-for-byte identical input — does the companion still feel like the same person? Early, partial data suggests the answer is no by default: swapping the model changes surface behaviour (word choice, question-asking habits, and in one measured case, an unwanted script-mixing failure) even when every other input is held fixed. A related, larger independent study published elsewhere in 2026 found a similar pattern, so this paper's contribution is narrower and sharper than originally planned: whether a specific engineering approach — externalised memory enforced by citations rather than free text — narrows that gap. We will publish honestly whichever way the data comes out.

Where this paper actually is

Primary comparison arm
74 of 2,304 calls
3%
Candidate arm
2,304 of 2,304 calls
complete
Rate limit
~75 calls/day on the current pool

Not a finding

7 raw Devanagari-script hits on the swap candidate's surface output

The raw candidate model, given the same compiled context as the incumbent, produced native-script characters against a hard-fail axis (any hit fails) in early generation. This is an unadjudicated data point, not a scored result, since the confirmatory incumbent-side comparison has not yet been generated.

Raw and unscored. No adapter or relational-layer mitigation was active in this run by design; this measures the model swapped in with nothing added, the paper's baseline condition, not its test condition.

context/measurements.md terra-arm-2304

Why the scope changed

The original framing was broader: that a swapped model, given the right context, would land inside an acceptable behavioural range regardless of which model it was. Our own earlier measurement had already made that claim harder to defend: an incumbent-against-candidate bake-off on this exact product found the incumbent winning 38 judged comparisons to 2. Then a larger, independently published study (2,008 conversations, three memory architectures crossed with four models) found the same shape of result: the underlying model, not the memory scaffold around it, sets the ceiling on whether a companion's persona survives. That result is a bigger, better-resourced version of a claim we had already measured, which means we no longer get to publish it as a first report.

What survives, and what has not been scooped, is a narrower and more specific claim: whether a context that is compiled byte-identically into both arms by our production engine, rather than a generated memory narrative fed to the model, changes the picture. That is an engineering claim about a specific architecture, not a claim about model swapping in general, and it is the paper we are now building.

What exists today

A compiled-context corpus of 2,304 distinct, hash-verified contexts, built so that both the incumbent and the candidate model are served byte-identical input by construction. The candidate arm is fully generated against that corpus: 2,304 of 2,304 replies, no errors.

One early, unadjudicated observation came out of that raw generation: seven instances of native-script (Devanagari) characters against an axis this product treats as a hard failure at any occurrence, in a model given no memory-layer mitigation by design. That run measures the swap with nothing added yet, which is the paper's baseline condition rather than its test condition. It is quarantined below, at the size an unscored count deserves.

What is missing before there is a paper

The matched incumbent-side data at the same scale, and a qualified judge for the relational-preference axes of the comparison. The second gap is, concretely, blocked on Paper B's own subject matter: no credit-billed candidate judge has yet cleared qualification for this kind of task.

Both blockers are logged with dates rather than described as almost finished. When the primary arm completes, this page gains findings; until then it has a question, a corpus, and two reasons the answer does not exist yet.

What this does not show

There are no findings on this page and no figure, because the primary comparison arm does not exist yet. The one observation shown is raw and unadjudicated, taken from a run with no relational-layer mitigation active by design, which is the paper's baseline condition rather than its test condition. The second blocker is the subject matter of the other paper: no credit-billed judge has cleared qualification for the relational axes this comparison needs.

Artifacts

Every slot renders whether or not it is filled, because what does not exist yet is part of the status.

PDF

Not written; data collection incomplete.

arXiv

Not applicable. The paper is pre-draft.

Code

Not applicable yet.

Benchmark

Not applicable yet.

Cite

Not citable yet. This page will carry a BibTeX entry when there is a paper to cite.

← Back to research

Page last updated 18 August 2026