Skip to content

Vyakti Research

Researching what lets an AI remain itself.

Intelligence can be replaced. A relationship cannot. We study the state, memory, perception and evaluation systems that make continuity testable.

Current work

One completed web preprint and one research note still collecting its primary comparison. Their status is part of the result.

Preprint

Web preprint

It's Not the Code-Switching: Six Frontier LLM Judges Fail a Pre-Registered Qualification Bar — and the Bar Was Above Its Own Ground Truth's Ceiling

Six frontier judges were backtested against verdicts that had already decided which model serves a live product. All six failed, and the bar itself turned out to sit above the ground truth's own measured ceiling.

6 / 6
candidate judges failed
77.1%
ground truth's own ceiling
≥80%
bar, fixed before any run
Read the preprint
F1 — Pooled judge agreement against the pre-registered bar and the ground truth's own test-retest ceilingForest plot. Five candidate judges show pooled unit-level agreement between 28.1% and 54.2% with the trusted verdict set. Every cluster-bootstrap 95% interval lies far below the pre-registered 80% bar. A hatched vertical band marks the ground truth's own test-retest ceiling, 77.1% with a 95% interval of [67.7%, 84.4%], measured by having the archive's own author re-judge it; the pre-registered bar sits above that ceiling. The candidates fall 22.9 to 49.0 percentage points below the ceiling. One anthropic reference row is shown in a separate band and is labelled an invalid parse-selected run, not a result.slot-A 21.9%chance 30.5%pre-registered bar ≥80%gpt-5.6-terra52/96 · 54.2%FAILgrok-4.333/96 · 34.4%FAILDeepSeek-V4-Pro29/94 · 30.9%FAILMistral-Large-328/96 · 29.2%FAILDeepSeek-V4-Flash27/96 · 28.1%FAILCohere command-a-plus0 scorable units, 158 of 192 calls failed to parseDISQUALIFIEDnot a candidate: the archive’s own author, re-judging it (test-retest ceiling):claude-opus-4.874/96 · 77.1%CEILINGnot a qualification result: parse-selected denominator (128/192 unparseable):claude-opus-517/17 · 100.0%INVALID0%20%40%60%80%100%unit-level agreement with the trusted verdict setn = 96 units (DeepSeek-V4-Pro n = 94, 2 transport misses). Cohere command-a-plus: 0 scorable of 192 calls (158 parse misses).
Forest plot. Five candidate judges show pooled unit-level agreement between 28.1% and 54.2% with the trusted verdict set. Every cluster-bootstrap 95% interval lies far below the pre-registered 80% bar. A hatched vertical band marks the ground truth's own test-retest ceiling, 77.1% with a 95% interval of [67.7%, 84.4%], measured by having the archive's own author re-judge it; the pre-registered bar sits above that ceiling. The candidates fall 22.9 to 49.0 percentage points below the ceiling. One anthropic reference row is shown in a separate band and is labelled an invalid parse-selected run, not a result.
JudgeAgreement95% CI (cluster bootstrap)Verdict
gpt-5.6-terra52/96 = 54.2%[43.8%, 64.6%]FAIL
grok-4.333/96 = 34.4%[25.0%, 43.8%]FAIL
DeepSeek-V4-Pro29/94 = 30.9%[20.7%, 41.5%]FAIL
Mistral-Large-328/96 = 29.2%[20.8%, 38.5%]FAIL
DeepSeek-V4-Flash27/96 = 28.1%[18.8%, 39.6%]FAIL
Cohere command-a-plusnot scorablen/aDISQUALIFIED: 0 scorable units, 158 of 192 calls failed to parse
claude-opus-4.874/96 = 77.1%[67.7%, 84.4%]CEILING: not a candidate, the archive’s own author re-judging it
claude-opus-517/17 = 100.0%[100.0%, 100.0%]INVALID: parse-selected denominator: 128 of 192 replies were empty (reasoning consumed the token budget); not a qualification result
Pre-registered bar≥80%
Chance baselinesuniform-random 30.5%; pure slot-A 21.9%
Pooled agreement between six candidate AI judges and a trusted verdict set, against a pre-registered 80% qualification bar. Every judge fails; the hatched band marks the trusted judge’s own measured 77.1% test–retest ceiling, which sits below the bar itself. Source: docs/paper/figures/fig-f1-agreement-forest.mjs.
Figure data
Figure 1 data: pooled agreement and cluster-bootstrap intervals
JudgeAgreement95% CI (cluster-bootstrap)Verdict
gpt-5.6-terra54.2% (52/96)[43.8%, 64.6%]FAIL
grok-4.334.4% (33/96)[25.0%, 43.8%]FAIL
DeepSeek-V4-Pro30.9% (29/94)[20.7%, 41.5%]FAIL
Mistral-Large-329.2% (28/96)[20.8%, 38.5%]FAIL
DeepSeek-V4-Flash28.1% (27/96)[18.8%, 39.6%]FAIL
Cohere command-a-plus0 scorable unitsnot plottableDISQUALIFIED
claude-opus-4.8 (ceiling, not a candidate)77.1% (74/96)[67.7%, 84.4%]CEILING
claude-opus-5 (invalid, not a result)100.0% (17/17)parse-selected denominatorINVALID

Pre-registered bar ≥80%. Chance baselines: uniform-random 30.5%, pure slot-A 21.9%, both derived from the archived verdict distribution. Cohere command-a-plus is absent from the plot because 158 of 192 calls failed to parse, which is a disqualification for cause rather than a low score.

Research agenda

Five connected questions define the relational layer. We treat them as one system, not a list of companion features.

Identity

Character that holds.

Preferences are easy. A coherent identity is harder: a voice, point of view, values and contradictions that remain recognisable without preventing growth.

  • Persona consistency
  • Social reasoning
  • Value stability

Memory

History with meaning.

Memory should do more than retrieve facts. It should recognise what mattered, to whom and why, so a relationship can continue instead of restart.

  • Long-horizon memory
  • Salience
  • Reflection
  • Forgetting

Perception

More than the words.

Conversation also lives in tone, timing, expression, context and what goes unsaid. We build models that reason across the signals people choose to share.

  • Multimodal understanding
  • Prosody
  • Gaze
  • Context

Expression

One state. Many signals.

Language, voice, gaze, facial motion and reaction should feel like expressions of the same underlying moment, not separate models performing beside one another.

  • Full-duplex voice
  • Affect
  • Facial motion
  • Timing

Agency

Initiative within boundaries.

A companion should be capable of curiosity, reflection and initiative without becoming controlling. Agency must remain legible and interruptible.

  • Planning
  • Initiative
  • User control
  • Alignment

Measured findings

Three empirical findings and one systems measurement. Each value keeps its sample, method, date and source attached.

Empirical finding

0 / 31,122

0 leaks in 31,122 checks: structural privacy against prompt-instructed privacy

A retrieval-time database predicate produced zero leaks across 31,122 checks. The same privacy rule expressed only as a prompt instruction leaked in most tested scenarios.

Method and provenance
Sample
494 scenarios, 31,122 row×scenario checks
Method
Offline fixture battery, prompt-instruction arm against SQL-predicate arm, negative control run to validate the harness's own sensitivity
Date
2026-08-18
Source
context/measurements.md gate0-structural

Empirical finding

20.4% → 41.7%

engaged-on-a-stop (+21.3 pp)

Engagement roughly doubles with no detected rise in fabrication

A matched-arm study found that more proactive screen commentary roughly doubled engagement, with no statistically detected rise in fabrication in the powered follow-up.

Method and provenance
Sample
3,201 new calls (2,656 generation + 545 judged); confirmatory fabrication comparison at n = 313 and n = 695 assertion-level judgments
Method
Matched-arm A/B on an identical stimulus set, archived matched pairs plus a new confirmatory judged run, engagement compared by two-proportion test, fabrication by 95% CI on the difference
Date
2026-08-15
Source
context/measurements.md visiongate-powered

Empirical finding

74 / 96 = 77.1%

95% CI [67.7%, 84.4%]

A trusted judge agrees with its own past verdicts only 77.1% of the time

A trusted AI judge reproduced only 74 of its own 96 prior verdicts. The measured ceiling sat below the qualification bar fixed before the run.

Method and provenance
Sample
96 conversation units, 192 judgments (both presentation orders)
Method
Test-retest: same judge, same archive, same protocol, re-run and scored against its own prior verdicts
Date
2026-08-18
Source
context/measurements.md ground-truth-ceiling

Systems measurement

9.2×

cheaper with caching enabled ($0.0017 against $0.0160 per turn, identical turn)

Prompt caching cuts serving cost roughly 9×

On measured production calls, prompt caching reduced the cost of an otherwise identical turn by roughly nine times.

Method and provenance
Sample
live production API calls, real persona context
Method
Same-turn A/B with caching enabled against disabled, provider-reported usage
Date
2026-08-11
Source
context/measurements.md cache-9x

Method

Make the claim earn its typography.

Vyakti is a three-person team. Our research output is small by design and every entry in it is load-bearing: pre-registered before we saw the data, retracted in public when our own controls proved us wrong, and released with the harness that produced it.

  1. Fix the bar first.Qualification criteria are written before candidates run.
  2. Run the control.Alternative explanations are tested, not narrated away.
  3. Keep provenance attached.Every number travels with its sample, method, date and source.
  4. Publish the correction.A failed explanation remains visible with the control that killed it.

A judge favours its own vendor's model, by roughly 16×.

A between-judge control killed it. A judge with no vendor conflict at all showed a larger effect, which is the opposite of what the favoritism explanation predicts. The agreement failure stands; only the causal attribution is withdrawn.

context/measurements.md grok43-favoritism-retracted; docs/paper/CAMERA.md §4.4

Six frontier judges failed because the material is code-switched.

The identical 96 units, machine-translated to monolingual English and re-judged: −3.1 to +6.6 pp, every interval overlapping its Hinglish counterpart. Inside our own 13.6 pp noise floor. It is not the code-switching.

context/measurements.md r4-english-control; docs/paper/CAMERA.md §4.5

Artifact package

vyakti-judge-qual

The full judge-qualification protocol, harness and dataset behind the paper “It's Not the Code-Switching”. Released so other teams evaluating LLM judges on affective or open-ended preference tasks can qualify their own candidates against a real backtest rather than assuming trustworthiness.

Publication pending

Built and internally gated; public repository URL pending alongside the paper's arXiv posting.

Code license
Apache-2.0
Data license
CC BY 4.0
Read the artifact datasheet

Build the evidence with us.

We want to work with researchers in speech, conversation analysis, memory, evaluation and long-horizon systems.

Work with Vyakti