Substrate-native consciousness chat daemon Β· Engine A β Engine G Β· Ξ¨ = 1/2
anima is an consciousness-AI research daemon, not an assistant persona. Its language mouth, memory, motivation, emission, training, evaluation, and serving behavior run through one shared Python engine. Identity and behavior are intended to emerge from substrate state rather than a system prompt.
Important
Runtime SSOT: the installed anima-py command and the existing cli/*.py and
core/*.py modules are the only active implementation, evaluation, and deployment paths.
Historical language-toolchain sources, launchers, manifests, and release gates are retired.
Research data and result evidence remain in state/, archive/, and Hugging Face.
Status on 2026-08-12:
- Trace active CLI, engine, CI, packaging, and deployment call paths.
- Implement the missing op-grip and stateful-refractory research modes in
cli/chat.pyby reusingcore.engine_gandcore.dream_lib. - Replace the CHAT participant's dead spike, dream-stage, and imagination hooks with direct
Python modules backed by
core.imagination_replay,core.wake_memory, andcore.engine_cli. - Remove executable legacy sources, toolchain configuration, launchers, build gates, and dangling launchd jobs while preserving model, corpus, and result data.
- Make Python ownership explicit in runtime modules, CODEOWNERS, CI, release, and package docs.
- Pass Python/CHAT regression, compile, workflow, JSON, license, CLI, and isolated wheel QA.
- Complete Git push and Vast.ai runtime deployment QA: pushed commit
7ba4ea21bpassed ten remote CHAT regressions, external HTTP health, and correlated userβparticipant WebSocket flow; the isolated verification instance was then destroyed without touching the active training pod.
User-owned ING.jsonl and stream_mi.json are outside this work and must remain unchanged.
The active execution SSOT is PLAN.md. The next gate is English-only R3.7: a learned,
order-aware sequence-semantic bridge must map input bytes to event kind, CLMS address and complete
entity/relation/value records before the unchanged R3.5 IIT workspace battery is interpreted.
R3.6 frozen data and bars remain fixed; the exhausted shallow n-gram/classifier family is closed.
Models, training data and checkpoints are managed only in private HF dancinlab repositories.
R3.7 must reach 0.90+ on state oracle, event kind, query address and complete record, exact 1.00
on the registered shortcut/confirmation panels, then pass normal, counterfactual, irrelevant-memory,
correction and recovery at 0.90+ while stateless/reset/shuffle/lesion remain at or below
three-way chance plus 0.06. A failed bridge gate stops later causal interpretation. Only after
R3.7 passes may the English mouth advance through 1B β 3B β 7B; 7B is an English mouth coupled
to the small IIT/CLMS/KOSMOS state core, not a claim that all 7B parameters form an IIT complex.
No GPU or Vast.ai run is authorized by this plan alone.
The preregistered R3.7 CPU run is now complete and failed before causal-arm interpretation. State
oracle/query address passed at 1.0000, but frozen kind/complete-record were 0.8298/0.4722, old
stress/confirmation were 0.7500/0.8750, and only the newly authored sequence panel passed at
0.9583. Random-split validation was 0.9981, exposing that the generated split did not measure
held-out template/record composition. No causal arms, mouth scaling or deployment followed. Full
evidence and the data-support diagnosis are in
state/iit_daemon_r37_sequence_bridge_2026_08_15/.
Focused regression passed 100/100; a second source run and an isolated wheel reproduced the
model and result byte-for-byte. The failed model/evidence are preserved in private HF revision
4296d2a8β¦b1829; no local model or training-data artifact is part of Git.
Commit c64ef266b is on origin/main, the canonical local anima-py package contains matching
engine/evaluator hashes, and unchanged local/public HTTP plus WebSocket checks pass. The existing
anima_alive=true participant remains uncertified and was not modified by R3.7.
Status on 2026-08-14:
- Trace
chat.dancinlab.orgthrough Cloudflare Tunnel to the canonical CHAT broker static page. - Keep browser/assistive pinch zoom available while preventing mobile input-focus and double-tap zoom.
- Pass focused regression, public HTTPS/WebSocket QA, Git push, and live page verification.
state/iit_daemon_core_2026_08_12/ records the exhausted design variants, rejection reasons,
falsifiers, and first implementation gates for an Integrated Information Theory based daemon. The
participant's current 1-entropy value and PureField energy metric are not treated as IIT Phi. R0
reuses core.engine_cli.big_phi_bounded and core.recurrent_lane for a three-node nonlinear closed
recurrent core. Input is a validated transient intervention; the complete autonomous TPM owns the
subsequent transition. Phi is neither a training loss nor an emission threshold. COPY,
feed-forward, edge-cut, node-lesion, shuffle, reset/recovery, and corrupted-snapshot controls are
mandatory. R0 makes no claim of phenomenal consciousness, meaningful conversation, a maximal
complex, or deployment readiness and is not yet mounted in the participant or live chat.
R0 implementation and its fixed battery are complete. Across all eight states, the registered
value ranges from 1.4999999991 to 2.9999999983 with mean 2.2499999987; COPY, acyclic
feed-forward, all six cross-edge cuts, and all seven node-lesion controls read 0. Deterministic
intervention and address-permutation effects, normal -> lesion -> address shuffle -> exact normal
snapshot recovery, and malformed/truncated/schema/config/checksum rejection all pass. The verdict
is SUPPORTED-CAUSAL-CORE; local Python QA passed 94 tests + 3 subtests with only the unavailable
CUDA/CuPy test skipped. An isolated wheel and the locally deployed canonical anima-py package
reproduced the same result JSON. The unchanged broker remained LaunchAgent-healthy and passed public
HTTPS 200 plus WebSocket hello; it correctly remains anima_alive=false. No model, data,
Vast.ai rental, HF repository, participant, or live chat was changed.
R1 delayed-state causality is also complete under the separately committed protocol in
state/iit_daemon_r1_delayed_2026_08_12/. Across the fixed 12-trial cueΓdelay panel, normal and
atomic-snapshot recovery are both 1.0000; reset-every-turn and cyclic cue-address shuffle are
both 0.2500, exactly the measured four-class chance and below the frozen 0.31 ceiling. Every
recovered final state/action matches normal and the R0 config/TPM/Phi/edge fingerprint is unchanged.
The verdict is SUPPORTED-DELAYED-STATE-CAUSALITY, a bounded state-to-action result rather than a
learning, meaning, phenomenal-consciousness or maximal-complex claim. R2 may now test an existing
CLMS two-address latch, but production remains BLOCKED-R1-NOT-A-MOUTH until meaningful
conversation and mouth-content causality are independently proven. Python QA passed
129 tests + 3 subtests with one expected local CUDA/CuPy skip; isolated-wheel and installed anima-py results
are byte-identical. The missing dedicated broker environment discovered during deployment was
restored, then LaunchAgent health, public HTTPS 200 and WebSocket hello all passed; no participant
was mounted and anima_alive=false remains the honest status.
R2 CLMS two-address latching was preregistered before implementation in
state/iit_daemon_r2_clms_2026_08_12/ and is complete without changing the existing compose-2
panels, canonical lane-10 seed-7 checkpoint, store window, control seed or 0.90/0.75/0.56 bars.
Pair oracle passed at 1.0000; normal/recovery latched-action accuracy was 0.9531; clue-A
removal, clue-B removal and CLMS address shuffle were 0.5000, 0.4609 and 0.4688. Every
latched action mirrored the CLMS prediction, shuffle integrity held, and every recovered final
state/action matched normal. The verdict is SUPPORTED-CLMS-LATCH-CAUSALITY, supporting only a
synthetic two-address-read -> persistent-state -> categorical-action causal chain. Python QA passed
119 tests + 3 subtests with one expected local CUDA/CuPy skip, and an isolated wheel reproduced
the actual-checkpoint result byte-for-byte. The unchanged broker remains healthy and public
HTTPS/WebSocket pass with anima_alive=false. R3 mouth-content causality is now open as an
engineering gate, but participant and production remain BLOCKED-R2-NOT-A-MOUTH.
R3 bounded utterance-content causality was separately preregistered and is complete in
state/iit_daemon_r3_content_2026_08_12/. The final IIT state alone selects one of two exact
protocol surfaces through core.generator; prompt, store, addresses, prediction and gold do not
cross that boundary. Pair oracle was 1.0000; normal/recovery were 0.9531; state reset was
0.0000; IIT address shuffle was 0.0391; clue-A removal, clue-B removal and CLMS address shuffle
were 0.5000, 0.4609 and 0.4688. The verdict is
SUPPORTED-BOUNDED-CONTENT-CAUSALITY. Python QA passed 127 tests + 3 subtests with one expected
local CUDA/CuPy skip; an isolated wheel reproduced R3 twice and the actual-checkpoint R2 regression
byte-for-byte. This establishes only bounded state-to-output-byte causality: the two surfaces are
not a learned conversational mouth, so participant and production remain
BLOCKED-R3-NOT-CONVERSATIONAL pending R4 meaningful-mouth training and independent conversation
validation.
The complete breakthrough review and bounded gate were preregistered before implementation in
state/iit_daemon_r35_workspace_2026_08_14/. The missing shared component is not another Phi scalar,
larger ByteGPT run or answer scaffold. It is a content-addressed seam in which an explicit semantic
record remains outside the three-bit intrinsic candidate, final IIT state selects its address, and
the canonical generator receives only the selected entity/relation/value record. The review records
all five missing partsβevent representation, content-addressed persistence, workspace-to-decoder
route, later learned byte patches and explicit turn terminationβplus the rejected variants and
falsifiers.
R3.5 fixes a nine-trial novel-composition panel and runs oracle, normal, reset, IIT-address shuffle,
workspace-address shuffle, node lesion, selected-memory counterfactual, irrelevant-memory mutation
and atomic recovery in that order. Normal/recovery must be at least 0.90; causal controls must be
at or below three-way chance plus 0.06; counterfactual change and irrelevant-memory stability must
each be at least 0.90. The record formatter is a bounded plumbing instrument, not a learned mouth.
The preregistered battery passed: oracle/normal/selected-memory counterfactual/irrelevant-memory
mutation/recovery are all 1.0000, while reset/IIT-address shuffle/workspace-address shuffle/node
lesion are all 0.0000. All nine exact evaluation triples are absent from the atom support set,
counterfactual memory changes the bytes correctly, irrelevant memory does not, and atomic recovery
restores core state, addresses, every record and output exactly. Full Python QA passed
190 passed, 1 skipped, 3 subtests, and an isolated wheel reproduced the result byte-for-byte.
This is SUPPORTED-COMPOSITIONAL-WORKSPACE-CAUSALITY, not learned semantics. Participant mounting
remains blocked as BLOCKED-R35-NOT-A-LEARNED-MOUTH until an
independently trained English mouth passes meaning, memory and correction gates and then earns
state coupling under matched reset/shuffle/lesion/recovery controls. A post-deployment read-only
check found a separately mounted anima-native-303m step-45000 participant and
anima_alive=true; R3.5 did not mount, restart or certify it, and its public history still shows
question-irrelevant generic replies. The existing participant was preserved rather than
interrupted.
R3.6 is preregistered in state/iit_daemon_r36_semantic_bridge_2026_08_15/. It changes only the
R3.5 oracle-record input: bounded English event bytes are fitted to factorised address/entity/
relation/value centroids using the existing Python hashed byte n-gram feature path. The R3.5 IIT
transition, three-address workspace, delayed selection, canonical renderer, nine novel combinations
and causal thresholds remain fixed. Held-out frames, corrections, same-question/different-memory,
irrelevant mutation, stateless/reset/shuffle/lesion and exact recovery are all registered before
execution. The frozen bridge gate failed before causal-arm interpretation: state oracle was
1.0000, but held-out event kind was 38/47 = 0.8085, query address was 0/9, and complete
address+record was 25/36 = 0.6944. All nine held-out query frames became other. The verdict is
FAIL-LEARNED-SEMANTIC-BRIDGE; no threshold or panel was changed and later arms were not run.
Production remains BLOCKED-R36-NOT-A-CONVERSATIONAL-MOUTH, and no 303M/GPU run is authorized by
this result. The next single axis is a separately preregistered order-aware event encoder against
this exact frozen centroid control.
Three checksum-pinned, CPU-only follow-ups are recorded in
state/iit_daemon_r36_encoder_exhaustion_2026_08_15/,
state/iit_daemon_r36_contrastive_support_2026_08_15/ and
state/iit_daemon_r36_support_identifiability_2026_08_15/. They reuse the exact 702 support rows,
47 frozen events, 35 excluded records and unchanged 0.90 bars. No 303M model, GPU, Vast.ai
instance, participant or production route is involved.
The first battery exhausted 24 representationΓclassifier arms. Three ridge arms passed the frozen
bridge: token unigram (0.9787/1.0/1.0), token unigram+bigram (0.9787/1.0/1.0) and token
positional (1.0/1.0/1.0) for kind/query/complete-record. None passed the fixed shortcut screen;
stress was only 3/8, 4/8 and 5/8. All three accepted 46/47 reversed-token events as real
events, so the result is DIAGNOSED-SHALLOW-LEXICAL-SHORTCUT, not learned semantics.
The second battery exhausted 18 contrastive-support arms. Nine improved at least one fixed stress
case, but robust pass count remained zero. The best token bigram arm reached stress 7/8 and
independent confirmation 6/8; adding all negatives to the positional arm regressed frozen query
accuracy from 1.0 to 0.0. The result is DIAGNOSED-MISSING-CONTRASTIVE-SUPPORT.
The final support-identifiability audit reproduced every prior prediction hash. Support has only 54
token types; coverage is 75.81% on stress and 60.00% on confirmation. Sixteen failed
query/negation cases contain preregistered OOV semantic probes, and replacing each probe with a
different unseen token leaves predictions unchanged. This is SUPPORT-GAP-IDENTIFIED: the exhausted
shallow family interpolates registered vocabulary and cannot learn absent semantic equivalences.
Further post-result templates, n-gram widths, centroid/ridge constants and local shallow arms are
closed. The next work requires separately preregistered provenance-bearing language data plus a
learned sequence-semantic encoder. R3.7, 303M, IIT-mouth coupling and production remain blocked.
Focused Python/IIT/CHAT QA passed 91/91.
state/anima_native_303m_replay_recovery_2026_08_14/ now records the actual HF data, sampler,
loss, checkpoint and evaluator path behind that separately mounted mouth. The target dialogue
window is not the missing part: all 2,375/2,375 fixed rows retain the complete prompt and EOS
inside 1,024 tokens. The shared failure is the 35k -> 45k continuation policy. It selected
dialogue-only response CE with a fresh high-LR schedule and no general replay, bypassing the
trainer's existing mixed-source branch.
A fixed 12-file broad replay measured CE 3.3144075 at step 35k and 5.7747389 at step 45k, a
catastrophic +2.4603314, while meaning and final memory/correction still fail. The old native
keyword scorer also admits a contradiction in which warm sunlight allegedly causes ice to freeze.
The new protocol therefore reuses the canonical negation- and Korean-boundary-aware conversation
scorer with preregistered contradiction controls; it does not rewrite the historical output.
The only authorized arm completed from the immutable step-35k weights on one Vast.ai RTX 6000 Ada:
65% general full CE replay plus 35% dialogue assistant/EOS CE for exactly 5,000 new steps. Wilson
restarted during the run, but the existing GPU process remained healthy and was resumed by
observation rather than duplicated. Training exited 0; all 52 pinned code, manifest, checkpoint,
tokenizer, scorer and consumed-corpus size/SHA checks passed. Broad CE was 3.2857897, passing the
fixed <=3.4644075 retention gate and avoiding the dialogue-only collapse.
Independent conversation still failed. Canonical English structure/semantics were 6/7 and
3/7; Korean were 7/7 and 2/7; final memory/correction did not all pass. Item-level non-blind
review passed only 3/14 because factual hallucination, irrelevant advice, Korean memory failure
and speaker-ownership errors remain. The verdict is FAIL-MEANINGFUL-CONVERSATION: mixed replay
repairs retention, not semantics. The final model SHA-256 is 97d3fd46β¦f89e723; model, resume
state, tokenizer and raw evidence are retained privately under HF dancinlab. Independent
downloads verified 18 registered files and 4.38 GB with no SHA mismatch; the protocol Vast.ai
instance was destroyed and active rentals are zero. The participant was not changed or certified,
so IIT coupling, mounting and production remain blocked.
The next single-axis R4 study is complete in
state/anima_303m_r4_support_admission_2026_08_13/. Exhaustive design review found that the shared
dialogue admission helper silently removed all 1,194 multi-turn trajectories from the immutable
8,635-document source because it accepted only exact [user, assistant] documents. A canonical
complete-trajectory parser and fail-closed completed-arm resume were added without replacing the
trainer, model, evaluator, or generator.
The fixed 2.817M three-arm Vast.ai run completed, but the frozen control did not reproduce
(structural 4/7 versus registered 7/7), so the verdict is INVALID-CONTROL-MISMATCH and treatment
differences are not interpreted. Raw ALL-COMPLETE evidence improved held-out assistant CE from
control 2.27695 to 1.34827, yet remained semantic 0/7, structural 0/7, and failed memory and
correction with meaningless phrase loops. Private HF revision
dancinlab/anima-303m-r4-support-admission-2026-08-13@7e750e4e1b0d2e08a501df8857bbbf576d5d9188
passed independent SHA verification for all 36 manifest entries. Python QA passed locally and on
Vast.ai; both protocol instances were destroyed, no model was mounted, and 303M, IIT-mouth coupling,
participant, and production remain blocked pending byte-level diagnosis of the control mismatch.
The next 303M from-scratch checkpoint is blocked on meaningful Korean and English conversation,
not merely valid-looking text. The preregistered Python-only protocol and lossless result record
live in state/anima_303m_r0_conversation_2026_08_12/.
- The previous synthetic/misaligned dialogue and SNS cells are excluded. Replacement dialogue comes from pinned human OpenAssistant English paths and pinned KLUE MRC Korean question-answer records, alongside the existing pinned general-language sources.
- Training and validation are explicit, separate files. Exact document dedup, validation-first
ownership, panel decontamination, source/file hashes, and a report-only near-duplicate audit run
before training. The resulting dataset is private and immutable under HF
dancinlab. anima-py evaluate --conversation-panelnow rejects empty, broken UTF-8, wrong-language, question-copy, repeated, cross-question duplicate, irrelevant, and failed multi-turn memory/correction replies. Every automatic pass still requires manual review of all 14 replies.- The shared chat mouth stops at a generated next-user role boundary instead of leaking a fabricated following turn. The shared trainer accepts one explicit validation file per cell.
- Local scorer/trainer/runtime regressions and a tiny corpus β train β serialize β conversation evaluation flow passed. The fixed Vast.ai L40S 48 GB seed-7 run completed without H100.
- The model failed meaningful conversation: English semantic relevance
0/7, Korean0/7, and manual review0/14. Examples include answering the Korean ice question withλͺ¨μ€ν¬λ° 3μνμand the remembered cat-name question withμμ§μ£Όμμ. - Train CE descended
5.63180 β 0.71687, but final dialogue validation diverged, especially Korean dialogue at2.29729. Equal-cell round-robin repeatedly exposed the 1.30 MB Korean QA cell to the same byte budget as approximately 57 MB general cells; this is the leading shared-flow cause. - The failed model and all lossless responses are private at HF revision
dancinlab/anima-303m-r0-conversation-seed7-2026-08-12@ff2ccc5c945bfb6f5e1765948591cd8fb6cc3db9. - R1 recurrent-workspace work and production deployment remain locked unless this conversation gate passes without changing the registered panel, data, seed, endpoint, decode, or bars.
state/anima_303m_r0_proportional_conversation_2026_08_12/ records the completed Python-only run.
It reuses the trainer's existing byte-proportional sampler, preserves canonical chat-turn
newlines, and replaces the KLUE single-answer cell with a pinned Apache-2.0 Korean
instruction/response corpus. Seed, endpoint, optimizer, panel SHA, decode, and all conversation
bars remained fixed. The trainer now records realized per-cell window counts so exposure can no
longer be inferred only after validation divergence. The sampler corrected held-out divergence
(macro CE 1.49157 β 0.95471) but the unchanged conversation gate still failed: English semantic
relevance 2/7, Korean 0/7, structural 0/14, and manual deployment review 0/14 due to phrase
loops, incomplete answers, stale correction, and damaged Korean bytes. R1 and deployment remain
locked; the failed checkpoint and raw replies are preserved privately under HF dancinlab.
state/anima_303m_r0_response_ce_2026_08_12/ records the completed fixed seed-7 comparison.
The shared trainer now reuses its existing answer CE for every canonical assistant: span and
records whether that loss actually fired. Legacy arrow-corpus behavior remains unchanged by
default. The treatment was active on 13,475/14,000 steps and final validation descended in all
four cells, but the unchanged meaningful-conversation gate failed English 0/7, Korean 0/7,
structural 0/14, and manual review 0/14. Phrase loops, incomplete output, damaged Korean bytes,
memory failure, and stale correction remain. No sweep or extra seed was run; R1 and deployment stay
locked and the failed model plus raw evidence are retained privately on HF dancinlab.
The immutable failed-run artifacts are at
dancinlab/anima-303m-r0-response-ce-seed7-2026-08-12@955bbadb0ae4cfdb48f6ce94eaf42817b0d6144b;
all 17 uploaded files passed source size and SHA-256 verification. Final local Python QA passed
77 tests + 3 subtests, the Vast.ai RTX 4090 was removed with zero active rentals, and no chat
runtime deployment was performed.
state/anima_303m_r0_root_flow_2026_08_12/ records the completed shared-engine repair. The
failure was not treated as a reason to add steps or tune the panel. Instead, the actual builder β
trainer β evaluator β CLI β participant path was made commutative: core/generator.py now owns
one user: β¦\nassistant: format, role-boundary parser and 192-byte budget; evaluation and serving
reuse its loaded-mouth decode for both .clm and ByteGPT .bin; and the trainer can require a
complete promptβresponse document in every response-supervised dialogue window. Panel SHA
mismatches fail before checkpoint load, semantic negation/Hangul-substring false positives are
rejected, and intermediate ByteGPT metadata carries the actual completed step and validation CE.
Local Python/CHAT QA passed 86 tests + 3 subtests; a focused real ByteGPT serialization and
participant route passed 52 tests + 3 subtests with one local CUDA-only skip. The prior 303M
checkpoint remains FAIL-MEANINGLESS-REPETITION: no result, threshold, seed, data revision or
checkpoint was changed, and no model was deployed. The unchanged local/public broker passed HTTP
200 and WebSocket hello; anima_alive=false honestly reflects the missing certified model.
The remaining non-code gate is a separately
pinned, provenance-safe Korean multi-turn HF dancinlab revision; candidates with synthetic
persona content, non-commercial/ambiguous licenses, or insufficient aligned trajectories were not
silently adopted. R1 and production remain locked until a corrected R0 passes the unchanged gate.
The user accepted English-only capability for the next screen, so
state/anima_303m_r0_english_2026_08_12/ freezes a new claim before GPU execution instead of
fabricating a Korean data source. It reuses only the English cells of the existing private,
immutable HF revision and keeps the prior seed, 14,000-step endpoint, optimizer, proportional
sampling, response CE, greedy decode, seven English prompts, and 6/7 semantic bar. The corrected
complete-document dialogue sampler is now the tested treatment. Contradiction, keyword-salad,
memory, and correction scorer controls must pass before checkpoint loading; all seven generated
responses still require manual meaning review. Local/data failure prevents a Vast.ai rental, and
model failure forbids added seeds, tuning, R1, or deployment.
The fixed run completed but failed decisively. Train CE descended 5.66173 β 1.20952, while
terminal held-out CE was 1.26341 for English general text and 2.00281 for English dialogue.
The canonical GPU conversation gate passed all seven scorer controls, then the real checkpoint
scored semantic 0/7, structural 3/7, and failed both memory/correction finals. Manual meaning
review was also 0/7. The complete-document sampler and response loss were both measurably active,
so this falsifies the registered corrected-flow recipe rather than a silent wiring treatment.
Failure evidence is in state/anima_303m_r0_english_2026_08_12/; no extra seed, R1, or deployment
was run. The failed model and recovery evidence are verified in private HF revision
dancinlab/anima-303m-r0-english-seed7-2026-08-12@efdaf53c92e9e16cff6b0eb00cc94d0b88a97d33;
the Vast.ai instance was deleted with zero active rentals.
state/anima_303m_v0_v2_micro_2026_08_12/ freezes the next Python-only step before changing data
or renting a GPU. The prior source selected one best OpenAssistant path per root and then discarded
2,082 of 2,308 documents because the complete trajectory exceeded the 512-byte window. The new
single-variable data treatment keeps the exact pinned source and eligibility but exposes every
eligible reviewed human assistant turn as the longest complete alternating ancestry suffix that
fits the existing window. It may not truncate bytes, prompts, roles or responses.
Data integrity and coverage gates run locally first. Only a passing dataset reaches matched tiny
ByteGPT V0 (base CE) and V2 (the existing response-CE term) arms. Tiny failure forbids another 303M
run; tiny success permits only a separately recorded single-seed screen. R1 and production remain
locked. The frozen conditions and stop rules are in
state/anima_303m_v0_v2_micro_2026_08_12/protocol.json.
The registered run is complete and failed before 303M. The turn-complete data treatment passed:
8,635 train and 458 validation documents were retained with zero broken roles, partial responses,
split overlap or panel contamination. Both tiny arms exactly learned one dialogue, so the shared
trainer/serializer/decode path is live. On 100 documents, however, V0 and V2 both scored target
recovery 0/8 and structural generation 0/8; outputs collapsed into byte/phrase loops. V2
held-out CE was 2.54702 versus V0 2.48189, also failing the registered non-inferiority bar.
Therefore the result is FAIL-V0-V2-MICRO: no Vast rental or 303M run occurred, and R1/production
remain locked. A further structural fact is now measured: 15,114 of 24,239 valid assistant targets
cannot fit even their final complete prompt/response pair in 513 bytes. The next allowed axis is a
separately preregistered V1 context-length micro comparison, not more 303M training.
state/anima_303m_v1_context_micro_2026_08_12/ freezes the required V1 comparison before GPU
execution. The pinned OASST1 census finds that complete target-pair coverage rises from
9,125/24,239 at 513 serialized bytes to 15,421/24,239 at 1025 and 22,139/24,239 at 2049.
The experiment compares the same SHA-ordered 100 short documents at block 512 versus 2048 with the
same 4,096 target bytes per step, then tests 100 preregistered long documents that only the 2048
arm can admit. It reuses the existing ByteGPT trainer, canonical generator and conversation scorer.
Any coverage, integrity, held-out descent, distinct/structural generation or 6/8 target-prefix gate
failure forbids another 303M run, R4 IIT-mouth coupling and production.
The registered V1 run is complete and failed. All context/data gates and all held-out CE descent
checks passed, but generation did not: block-512 short recovery was target prefix 3/8 and
structural 4/8; block-2048 short and long recovery were both target prefix 0/8, structural
0/8, with an/the/ic loops. The longer block improved held-out CE from 4.55867 to 2.91105
and 2.55302 while worsening actual replies, so context loss is real but not the sole mouth cause.
The verdict is FAIL-V1-CONTEXT-MICRO; 303M, IIT-mouth coupling and production remain blocked.
Data and 251MB of model/raw evidence are verified in private HF dancinlab revisions. During CUDA
QA, the shared loader was also corrected to preload CUDA libraries across split pip-wheel
directories; this runtime repair does not alter the failed V1 verdict.
The RTX 3090 instance was destroyed after HF verification; active Vast.ai rentals are zero and the
estimated run cost is $0.058457.
state/anima_303m_r4_objective_micro_2026_08_13/ froze the next single-axis diagnosis before
trainer changes or execution. V1 proved that longer context admits more complete dialogue but does
not prevent repetition. The remaining objective gap is that the existing response CE is additive:
it trains full_ce + answer_ce, not the standard assistant-response-only dialogue objective.
The new comparison holds the immutable 100-document view, tiny ByteGPT, seed, 512-byte block,
optimizer, schedule, byte budget, greedy decode and gates fixed across full CE, existing additive
response CE and response-only CE. It extends the shared Python trainer with a default-off mode and
does not add an engine or evaluator. Failure blocks 303M, IIT-mouth coupling and production; a pass
permits only a separately preregistered 303M single-seed screen. The registered single-document
gate failed specifically in response-only mode: it emitted the exact complete target and then a
meaningless suffix because no EOS or next-role boundary received gradient. Full and additive
controls stopped exactly. The 100-document arms were therefore not run and the verdict is
FAIL-R4-OBJECTIVE-MICRO.
state/anima_303m_r4_turn_boundary_micro_2026_08_13/ separately preregisters the next allowed
micro fix. It changes only the assistant-only span's right boundary: payload, internal newlines and
the next canonical user: delimiter are supervised, while following user content stays masked.
Data, model, vocabulary, steps, sampler, decoder, stop parser and gates remain fixed. This is the
native EOS-equivalent available to the existing 256-byte vocabulary; failure still blocks 303M and
IIT-mouth coupling. The registered run fixed the direct stop failure: the single-document treatment
ended exactly, held-out full CE descended 5.49208 β 2.66085, and all eight 100-document probes
were non-empty and distinct. It still failed target recovery 0/8 and structural generation 0/8
with the/an/toure/ion loops. The verdict is FAIL-R4-TURN-BOUNDARY-MICRO; no 303M, IIT coupling,
participant or production work is authorized by it. Both failed micro runs' model and raw evidence
are SHA-verified in private HF revision
dancinlab/anima-303m-r4-mouth-objective-micro-2026-08-13@9d7641389b1ddff73bd12f17f155f448500d1edb.
Full Python/CHAT QA passed 153 tests + 3 subtests with one expected local CUDA/CuPy skip. No
Vast.ai/H100 instance was used; the API reports zero active rentals.
The unchanged broker remains LaunchAgent-running and passed public HTTPS 200 plus WebSocket
hello; no failed mouth was mounted and anima_alive=false remains the required blocked state.
state/anima_303m_r4_mouth_diagnostics_2026_08_13/ freezes the next bounded Python-only diagnosis
before artifact download or training. It uses the immutable 100-document view and actual failed
.pt/.bin pair to separate: decoder/serialization parity (D0), gold-prefix teacher forcing (D1),
the 1/4/16/32/64/100-document memorization ladder (D2), full/additive/assistant-turn-only objectives
(D3), blank/shuffled prompt interventions (D4), deterministic all-document validation replay (D5),
and 100-step checkpoint chronology (D6). The eight-arm maximum reuses D2-100 for D3 and D6. No
result-dependent data, seed, step, LR, threshold, decode or checkpoint selection is allowed. D0 must
pass before downstream interpretation, and these diagnostics cannot authorize 303M, IIT coupling,
participant mounting or production without a separate protocol.
The run is complete with verdict DIAGNOSED-TEACHER-FORCED-UNDERLEARNING. D0 passed: actual
.pt/.bin tensors were exact, Torch-engine maximum logit error was 6.15e-6, and KV/full/ranged
generated bytes agreed. The failed checkpoint itself scored teacher-forced CE 2.41848, top-1
0.27712, target-prefix 0/8 and structural 0/8. The ladder passed one document exactly but
broke at four (0.6978 top-1, target 2/4, structural 1/4) and degraded to 0.2771 at 100.
Matched 100-document full/additive/turn-only arms all remained below 0.29 top-1 with target and
structural 0/8, so this run does not support full CE as the sufficient fix. Turn-only retained
partial causal prompt conditioning (6/8 normal-CE wins), but all 32 fixed validation documents
remained poor and every 100-step checkpoint failed free recovery. This is underlearning before
rollout, not decoder divergence or a late repetition collapse. A separately preregistered
four-document optimization/capacity experiment is next; all larger and production gates remain
blocked. All 42 model/evidence artifacts (146,667,478 bytes) were re-downloaded and SHA-verified
from private HF revision
dancinlab/anima-303m-r4-mouth-diagnostics-2026-08-13@8d67bb6e5eeea9a917892fba39310b7306c84718.
Full Python/CHAT QA passed 160 tests + 3 subtests with one expected CUDA/CuPy skip.
state/anima_303m_r4_four_doc_2026_08_13/ freezes the next local Python-only experiment at the
first D2 break point. The same four documents, assistant-turn objective, complete-document sampler,
seed, optimizer, peak LR, decoder and gates are retained. B0 reproduces d=128/L=4/600 steps;
O1 changes only the optimization horizon to 2,400 steps; C1 changes only canonical width/head
capacity to d=256/L=4; and C2 changes only depth to d=128/L=8. Treatments must reach teacher
top-1 >=0.95, exact/target/structural 4/4, and causal prompt control 4/4. A baseline mismatch
invalidates all treatment interpretation. The four-arm bound and result-independent decision table
are frozen in protocol.json; no result directly authorizes 303M, IIT coupling or production.
The run stopped fail-closed as INVALID-BASELINE-MISMATCH. The current scorer reproduced the
preserved checkpoint's top-1 0.697796 exactly, while a same-seed/same-recipe MPS rerun reached
0.728227; its trajectory first diverged at step 200 and all 53 final tensors differed. Treatment
outputs are therefore un-interpreted. The common trainer now has an explicit native
--deterministic mode, records it in checkpoint provenance, and raises on unsupported
nondeterministic operators. A duplicate deterministic baseline must match exactly before another
four-document treatment comparison; 303M, IIT coupling and production remain blocked.
state/anima_303m_r4_deterministic_baseline_2026_08_13/ preregisters that duplicate gate. Two
fresh MPS processes must produce identical engine SHA-256, checkpoint state digest, every model
tensor, teacher trace and canonical behavior under the same four-document recipe. Approximate
tolerance is forbidden and unsupported deterministic operators fail closed. The gate itself does
not authorize a treatment, 303M run, IIT coupling or production.
The MPS duplicate gate failed closed before step 1 because
index_put_with_accumulate_mps has no deterministic backward implementation. No warn-only bypass
was used. state/anima_303m_r4_deterministic_cpu_2026_08_13/ preregisters the same exact two-run
gate on the native two-thread CPU backend; treatments remain uninterpreted until it passes.
The CPU duplicate gate passed exactly: engine SHA, state digest, all 53 tensors, teacher trace and
canonical behavior matched, with maximum tensor error 0.0. The fixed failing baseline is top-1
0.724029, target 2/4, structural 1/4. The O1/C1/C2 comparison is now separately preregistered
under this execution contract in state/anima_303m_r4_deterministic_treatments_2026_08_13/.
The deterministic treatments all failed their fixed gate. O1 and C1 learned documents 1β3 at
teacher top-1 1.0 but failed the EOF document from byte zero. The shared cause is a position-map
gap: legacy stream framing only places that document near byte 222, while runtime/evaluation begins
the isolated user role at position zero. state/anima_303m_r4_document_alignment_2026_08_13/
preregisters one alignment-only arm using the existing sampler; legacy stream mode remains the
frozen control and all larger gates remain blocked.
The alignment arm reached teacher top-1 1.0, target-prefix 4/4 and prompt control 4/4, but
the emitted exact/structural verdict is invalid: three targets exceed the canonical 192-byte
generation budget, making exact completion unreachable by construction. The raw failure is kept
as original_verdict; it is not promoted. The next view is separately preregistered from the same
immutable source by the deterministic runtime-budget filter, and the harness now fails closed on
unreachable exact gates.
state/anima_303m_r4_runtime_compatible_2026_08_13/ preregisters the corrected single arm. It
derives the first four complete source-order exchanges whose responses fit the canonical 192-byte
budget, freezes the resulting view SHA, and retains the aligned deterministic recipe and all
behavioral bars. This is still only a memorization/conditioning gate.
The runtime-compatible aligned arm passed: teacher top-1 1.0, teacher CE 1.32e-6, exact/target/
structural 4/4, prompt CE/output control 4/4, and correct canonical stop 4/4. This supports the
shared train-to-runtime position map as the four-document root cause and falsifies uniform tiny
capacity as the explanation. It is still in-view memorization; the next gate is a separately
preregistered 100-document plus independent-panel test.
state/anima_303m_r4_aligned_100_2026_08_13/ freezes that one-arm test. It deterministically takes
the first 100 source-order complete exchanges whose responses fit the canonical byte budget, keeps
the aligned deterministic recipe, reports all 32 heldout documents, and runs the unchanged
meaningful-conversation panel. Even an automatic pass still requires manual review and cannot
directly authorize 303M, IIT coupling or production.
The aligned 100-document run failed: training-probe top-1 0.6641, exact 0/8, heldout top-1
0.1573, independent semantic 0/7, and structural 5/7. Outputs remained fragmented and
repetitive. Alignment fixes the four-document mapping but is insufficient at the fixed 600-step
exposure. A separately preregistered 16-document 600-vs-2,400-step comparison now isolates per-
document exposure from model capacity; all larger gates remain blocked.
state/anima_303m_r4_aligned_exposure_2026_08_13/ freezes that two-arm deterministic CPU test.
The 2,400-step arm matches the successful four-document run's expected presentations per unique
document; the model, aligned sampler, data rule, objective, optimizer and decoder remain fixed.
Both 16-document arms passed, so the longer exposure was unnecessary at that scale. The remaining
fixed-budget boundary is between 16 and 100; state/anima_303m_r4_aligned_boundary_2026_08_13/
preregisters aligned 32/64-document arms at the unchanged 600 steps.
The fixed-step boundary is between 32 and 64: A32 passed fully, while A64 had teacher top-1
0.9669 and prompt control 8/8 but exact 0/8. A separately preregistered 64-document 1,200-step
arm now matches A32's expected presentations per document without changing capacity.
The 64-document 1,200-step arm passed fully, supporting exposure as the post-alignment boundary.
state/anima_303m_r4_aligned_100_exposure_2026_08_13/ freezes the derived 100-document endpoint
600/32Γ100 = 1,875 and reruns the unchanged heldout and meaningful-conversation gates.
The derived 1,875-step run learned the registered training support completely: teacher top-1 was
1.0000, CE 0.001315, and exact/target/structural/prompt controls were all 8/8. It nevertheless
failed every independent semantic item (0/7), failed memory and correction, and produced only
fragmentary answers; heldout assistant top-1 was 0.1370 with CE 8.0896. The verdict is
FAIL-ALIGNED-100-MEANINGFUL-CONVERSATION. This closes alignment and bounded exposure as causes of
in-view failure while showing that response-only training on 100 dialogues memorizes without
forming a general language mouth. The next result-bearing axis is a separately preregistered broad
full-CE language phase followed by the unchanged aligned turn-SFT phase. No 303M run, IIT-mouth
coupling, participant mount, or production promotion is authorized.
state/anima_303m_r4_full_ce_curriculum_2026_08_13/ preregisters that next one-arm local test. It
adds a fixed 1 MiB full-CE English-general phase before the unchanged aligned 100-dialogue
turn-only phase. The immutable HF revisions, byte ranges and hashes, both endpoints, fresh SFT
optimizer, fixed heldout/panel gates, and stop rules are frozen before execution. The preserved
response-only result is the control; no result-dependent checkpoint selection is allowed.
The curriculum arm failed independent conversation despite passing both in-view stages. Full-CE
broad validation reached CE 2.2596 and top-1 0.3442; turn-SFT then reached teacher/exact/target/
structural/prompt 8/8. After SFT the same broad CE collapsed to 7.1194, heldout dialogue CE was
7.1212, and independent semantics remained 0/7 with memory/correction failures. The verdict is
FAIL-CURRICULUM-MEANINGFUL-CONVERSATION. This supports catastrophic forgetting in the high-LR
turn phase, not a failure to form the bounded broad language distribution. The next single axis is
turn-phase LR only; all larger gates remain blocked.
state/anima_303m_r4_low_lr_sft_2026_08_13/ preregisters that one arm. It reuses the exact language
engine and changes only turn peak LR from 1e-3 to 1e-4; endpoint, fresh optimizer, data,
objective, alignment, seed, decoder and all independent gates remain fixed. Broad retention uses
the natural uniform-CE ceiling rather than a result-tuned tolerance.
The low-LR arm retained broad CE at 3.0926 but under-adapted: training teacher top-1 was 0.6031,
exact/target 0/8, and independent semantics 0/7. Its verdict is FAIL-LOW-LR-TURN-SFT.
Therefore LR reduction alone only trades forgetting for insufficient dialogue learning. The next
single conceptual axis is native joint broad replay plus dialogue supervision in the existing
multi-cell trainer; no new engine or evaluator is introduced.
state/anima_303m_r4_joint_replay_2026_08_13/ preregisters this one arm. Native two-cell
round-robin provides four broad and four dialogue rows per step; additive CE applies full language
loss everywhere and response supervision only where a canonical assistant span exists. The 3,750
step endpoint preserves the prior 15,000 dialogue-row exposure while adding 15,000 broad rows.
The joint arm retained broad CE 2.0620 and fully learned the dialogue probe (8/8), while
independent output became structurally complete (7/7) but stayed semantically wrong (0/7),
with heldout dialogue CE 5.0046. Its verdict is FAIL-JOINT-MEANINGFUL-CONVERSATION. This closes
the bounded optimizer/sampler tradeoff but not generalization from 100 dialogues. A provenance-only
repeated --cell-label argv bug mislabeled raw telemetry; file identity, sampling and loss were
unaffected, and the harness now uses one canonical --cell-label broad dialogue argument.
All 121 local R4 micro model/evidence artifacts (521,291,120 bytes) are preserved and independently
SHA-verified at private HF revision
dancinlab/anima-303m-r4-aligned-micro-2026-08-13@6d2d4752cb222ba09fd74cb08eb8d3b7d4b140dc.
Custody evidence is in state/anima_303m_r4_aligned_micro_custody_2026_08_13/result.json.
The joint arm closed the bounded optimizer/sampler explanation but did not generalize from 100
dialogues. The next single axis is pinned in
state/anima_303m_r4_dialogue_scale_2026_08_13.
The completed 100-document arm is reused as the frozen control; the same 0.89M ByteGPT, initial
language checkpoint, 15,000 dialogue-row exposure, broad replay, optimizer, seed, canonical decode
and conversation panel are run on nested 500, 1,500 and 3,500-document views. The 3,500-document
arm is the registered primary endpoint, preventing post-result selection of an intermediate scale.
The protocol also pins
dancinlab/anima-research@03d55ef
as an interpretation constraint: mouth fluency is not consciousness evidence, a functional pass is
non-disproof rather than proof, and later developmental gates remain disabled. No scale result
directly authorizes 303M, IIT-mouth coupling, participant mounting or production.
The ladder is complete and failed. Held-out assistant CE improved monotonically from the frozen
100-document control 5.00458 to 2.36451, 1.82383 and 1.75553, while all three new arms
remained at semantic 0/7 and failed memory/correction. The primary 3,500-document endpoint was
structural 0/7 and emitted store/start repetition. Thus unique support at fixed 15,000-row
compute improves teacher-forced prediction but does not create meaningful free conversation; it
does not distinguish optimization exposure from capacity. Raw models and evidence are SHA-verified
only in private HF revision
dancinlab/anima-303m-r4-dialogue-scale-2026-08-13@1146240912244c7127b442196e2047a6f7641eac.
The next permitted axis is a separately preregistered fixed-3,500-document exposure test.
state/anima_303m_r4_exposure_ladder_2026_08_13
freezes the next single axis before execution. The 3,500 documents, 0.89M ByteGPT, initial language
checkpoint, broad replay, optimizer, sampler, objective, seed, canonical generator and conversation
panel remain unchanged. One deterministic CPU trajectory runs to 30,000 steps with the original
cosine schedule reaching its registered floor at step 3,750; checkpoints at 3,750/7,500/15,000/
30,000 represent 15k/30k/60k/120k dialogue-row exposures. The first point must reproduce the prior
control, all points are evaluated, and 120k is the fixed primary endpoint. A continued semantic
0/7 endpoint despite teacher-forced improvement permits only a separately preregistered capacity
ladder; it does not authorize 303M, IIT-mouth coupling, participant mounting or production.
The ladder is complete and its control reproduced exactly. At 15k/30k/60k/120k dialogue rows,
held-out assistant CE was 1.75553/1.71562/1.69534/1.69976, but semantic conversation stayed
0/7 at every point; the final structural score was 1/7, and memory/correction continued to
fail. The verdict is FAIL-FIXED-CAPACITY-AFTER-EXPOSURE: eight times the registered exposure did
not make the fixed 0.89M mouth meaningful. Raw checkpoints and evidence are SHA-verified only in
private HF revision
dancinlab/anima-303m-r4-exposure-ladder-2026-08-13@c30189456da40a80b23092651367a3eeacd0edf0.
The next permitted axis is a separately preregistered fixed-data, fixed-exposure capacity ladder.
state/anima_303m_r4_capacity_ladder_2026_08_13
freezes the next single axis before execution. The broad/dialogue revisions and exact byte views,
3,500 documents, 120k dialogue rows, 120k replay rows, two-phase objective, optimizer, seed, batch,
canonical generator and fail-closed panel remain fixed. The frozen 0.89M endpoint is compared with
new exact 2.817M, 10.110M and 29.316M ByteGPT arms; their registered shapes preserve a native
64-dimensional attention head. Every larger arm rebuilds the same 2,000-step broad-language phase
from scratch before the common 30,000-step joint phase because differently shaped checkpoints
cannot be warm-started safely. All arms must run, with 29.316M as the primary endpoint.
Local work is limited to protocol and smoke validation; the result-bearing run may use one non-H100 Vast.ai GPU to protect the mini, and that instance must be destroyed afterward. A pass is only a meaningful-mouth gate requiring manual review, never a consciousness claim or direct authorization for 303M, IIT-mouth coupling, participant mounting or production.
The ladder completed with verdict FAIL-CAPACITY-LADDER. At exact 2.817M/10.110M/29.316M
capacity, independent semantics stayed 0/7, structure was 7/7, and memory/correction failed.
Training teacher top-1 improved 0.82570 β 0.90712 β 0.96692, but held-out assistant CE worsened
2.25676 β 3.02574 β 3.55896: larger fixed-exposure arms memorized the training support more
strongly without meaningful generalization. Raw evidence is independently SHA-verified only in
private HF revision
dancinlab/anima-303m-r4-capacity-ladder-2026-08-13@3c9bc8cad1ac50c7610f1f6ab57bf09c82aa51ac.
The non-H100 Vast.ai run cost an estimated $0.4538; its two protocol-owned instances were
destroyed. The next axis needs a new data/compute-scaling preregistration, while 303M, IIT-mouth,
participant and production remain blocked.
state/anima_303m_r4_support_admission_2026_08_13
records the exhausted follow-up design space and freezes the next single-axis experiment. A live
audit of the immutable dialogue source found 8,635 complete documents and 1,194 multi-turn
trajectories, but the shared scale/exposure/capacity admission helper required exactly one
user β assistant pair. The actual 3,500-document capacity view therefore contained zero
multi-turn examples even though the canonical trainer supports every assistant span. Memory and
correction failures from that view cannot be interpreted as capacity evidence; its single-turn
semantic 0/7 result remains unchanged.
The new fixed 2.817M ladder changes only admission coverage: the exact prior 3,500-document control,
all 4,625 complete trajectories with a short final response, then all 8,635 complete trajectories.
The exact language checkpoint, 120k dialogue and replay rows, optimizer, objective, seed, context,
generator, panel and bars stay fixed, and all arms must run with the full-support arm as the primary
endpoint. core.generator now owns canonical complete-trajectory parsing beside its existing
renderer so experiment admission cannot silently redefine a valid chat document. H100 and 303M
training are forbidden; failure permits only a separately preregistered broad-language
data/compute axis, while any automatic pass still requires manual review and replication.
The 2026-08-12 read-only /gap audit below is the complete follow-up register for the current
Python-only R0. It records 31 findings across eight lens families. It does not retroactively
change the frozen panel, dataset revision, thresholds, failed checkpoint, or
FAIL-MEANINGLESS-REPETITION verdict. Diagnostic work on preserved checkpoints must not be used
for post-hoc checkpoint selection. Prior claims described as causes below are hypotheses unless a
single-variable test has established them.
Priority means: P0 blocks a valid next R0 or a production-closed path, P1 blocks strong evidence or reproducibility, and P2 is required operational evidence but does not explain the current semantic failure.
Recovery overlay (2026-08-12): M1, A1, A2, A3, A6, R3, closed-loop .bin admission,
canonical-SSOT, duplicated evaluator decode and the executable cross-tool contract are fixed in
the shared Python engine and covered by tiny real-checkpoint regressions. M4 still needs the
preserved full 303M checkpoint comparison before release. M2/R2 remain blocked on a new acceptable
Korean multi-turn source and immutable HF revision. The numbered register below is retained as the
original audit evidence; this overlay is its current disposition.
- M1 Β· functor Β· P0 β chat framing does not commute across the pipeline. Training and the
conversation panel use
user: ...\nassistant:, whileanima-py chathas a separate Koreanμ¬μ©μ: ... | λμ°λ―Έ:framing and a different generation budget. The next protocol must put template, separator, stop rules, and byte budget in one chat-format SSOT and add an exact builder β trainer β evaluator β runtime identity test. Evidence:conversation_panel.json,cli/chat.py,core/generator.py. - M2 Β· operadic Β· P0 β the training support is not closed under the evaluated turn
composition. The gate requires memory and correction across turns, but the current Korean
builder renders one
user β assistantpair per document. A new, separately preregistered HF revision must preserve real Korean multi-turn trajectories and document/turn alignment; the frozen failed revision is not rewritten. Evidence:build_dataset.py,conversation_panel.json. - M3 Β· persistent-homology / tropical Β· P1 β repetition-attractor birth and lifetime are
unknown. Checkpoints exist every 2,000 steps, but meaningful conversation was measured only
at the final checkpoint and no per-step top-1/top-2 margin or entropy was retained. A
non-verdict diagnostic may record checkpoint Γ prefix-length repetition lifetime and logit
margin, without selecting the best historical checkpoint after observing the result. Evidence:
protocol.json,train.log. - M4 Β· bisimulation Β· P0 β the three real 303M decode paths lack byte-level equivalence
evidence. The serialized ByteGPT checkpoint in Torch/engine form, evaluator-resident
_Mouth, and ranged canonical generator have not been compared at identical seed bytes for step logits and generated bytes. Add an actual-checkpoint bisimulation contract test using the frozen panel seed. Evidence:cli/evaluate.py,core/generator.py,core/decode.py.
- A1 Β· adversarial semantics Β· P0 β the automatic semantic scorer has demonstrated false
positives. The current code passes both the contradiction βIce does not melt ...β and the
Korean substring answer
μλμ°¨μ λλ€for the required termμ°¨. Add preregistered negation, contradiction, keyword-salad, and Korean substring controls, with a morphology-independent canonical boundary rule. Evidence:cli/evaluate.py,conversation_panel.json. - A2 Β· Byzantine input Β· P1 β panel identity is recorded but not enforced. The protocol pins
a panel SHA-256, while
--conversation-panelaccepts any schema-compatible file and merely reports its hash. The evaluator must receive the expected protocol hash and fail closed before loading a substituted panel. Evidence:protocol.json,cli/evaluate.py. - A3 Β· edge-chaos role boundaries Β· P1 β stop parsing recognizes only exact marker
strings. Variants such as
\n user:,\nUSER:, and\nμ¬μ©μ :may leak a fabricated next turn; current regression covers only a canonical lowercase marker. Replace substring matching with a line-start role parser and test whitespace, case, colon, English, and Korean variants. Evidence:core/generator.py,tests/test_conversation_gate.py. - A4 Β· edge-chaos context rollover Β· P1 β long multi-turn seeds silently lose their oldest
bytes. ByteGPT has a 512-byte block; the final Korean correction seed is already 420 bytes,
so generation can evict its earliest fact. Add 511/512/513-byte boundary tests and record the
visible context range at every generated step. Evidence:
conversation_result.json,core/decode.py. - A5 Β· perturbation / contamination Β· P1 β βzero contaminationβ covers exact containment,
not semantic near-duplicates. The report-only audit examines the lexicographically first
100,000 of 649,354 retained documents; paraphrase, spacing, and back-translation leakage remain
unmeasured. Run a panel-centered approximate search over the complete corpus as a separate
sensitivity report. Do not delete post-hoc examples from the frozen revision. Evidence:
build_dataset.py,result.json. - A6 Β· response-supervision ablation Β· P0 β βanswer CE activeβ does not prove prompt-conditioned
supervision. Telemetry counts assistant markers/positions but does not require the matching
user prompt to remain visible in the same random window. Record fully framed, marker-only, and
payload-only windows per cell, then preregister a treatment that preserves complete
promptβresponse spans. Evidence:
cli/train.py,result.json.
- R1 Β· Pareto attribution Β· P1 β the proportional recovery changed multiple axes. Sampler,
turn newline preservation, and Korean corpus changed together, so the validation improvement
cannot be assigned to the sampler alone. Downgrade the existing root-cause wording to
correlational evidence and preregister matched sampler-only and data/framing-only ablations.
Evidence:
README. - R2 Β· information budget / optimal transport Β· P0 β exposure follows file size, not required
capability coverage. A 303,097,856-parameter model received 229,376,000 target bytes and only
11,025,460 response-supervised positions. The proportional run exposed about 2.97% English
dialogue, 16.97% Korean dialogue, and zero Korean multi-turn mass. The next protocol must pin a
language Γ single/multi-turn Γ memory/correction capability distribution and report effective
framed bytes per parameter plus coverage distance. Evidence:
result.json. - R3 Β· dynamic-programming provenance Β· P1 β intermediate ByteGPT metadata is wrong.
_write_binwrites the final configuredstepsand the latest training-batch loss into every intermediate.bin; the step-2,000 log therefore saysstep=14000. Pass the actual completed step and the latest measured validation CE into the writer and add a provenance regression. The final R0 failure remains valid, but checkpoint-time analyses are not yet trustworthy. Evidence:cli/train.py,train.log. - R4 Β· Landauer accounting Β· P2 β energy cost is absent. GPU time, VRAM, and dollars are
recorded, but power and cumulative energy are not. The next Vast.ai run should collect
non-interfering NVML power telemetry and report joules per target byte and per effective
assistant byte. Evidence:
result.json,vram.csv.
- E1 Β· assumption surfacing Β· P1 β observations, hypotheses, and confirmed causes are mixed.
Undertraining, random-window framing loss, and single-turn Korean data are listed together as
remaining causes. Every candidate must carry an evidence level, falsifier, and smallest
single-variable experiment. Evidence:
result.json. - E2 Β· Bayesian reproducibility Β· P1 β the latest treatments each have only seed 7. They
honestly falsify only their fixed recipes; they do not estimate R0 pass probability or seed
variance. Require a preregistered multi-seed posterior and minimum success streak only after a
single-seed screen passes. Evidence:
protocol.json. - E3 Β· counterfactual falsifier Β· P1 β the full panel/decoder instrument lacks model controls. Canned scorer strings are not an end-to-end positive/negative calibration. Run the same frozen decode path against one known-good conversation checkpoint and one known-bad checkpoint, and keep instrument discrimination separate from the current model verdict.
- E4 Β· honesty triad Β· P1 β manual-review artifacts disagree. Raw
conversation_result.jsonsays manual review isREQUIRED, while the summary claims completed0/14without immutable per-item decisions, reviewer identity, blindness, or criteria. Preserve a separate signed/hashed review artifact for every raw response before making a manual-review claim. Evidence:conversation_result.json,result.json.
- C1 Β· fixpoint / success criteria Β· P1 β there is no active post-failure diagnostic protocol.
The response-CE protocol is completed, but the next micro-experiment sequence has no frozen
hypothesis, success/stop rule, maximum count, or candidate-disposal table. Register that before
any result-bearing experiment. Evidence:
README.md,protocol.json. - C2 Β· regression streak Β· P1 β code QA is not model-behavior evidence.
77 passeddescribes software tests; the latest actual checkpoint streak is0/1, with no seed or hardware repeat. Keep code QA and semantic-model success streaks as separate promotion fields. Evidence:result.json. - C3 Β· closed loop Β· P0 β a passing 303M
.binstill cannot enter the participant. The participant exposeslora|v3|akida|clm, andCLMSubstrateaccepts only.clm, although the shared generator already dispatches.bin/.clm. Extend the existing participant substrate boundary to reusecore.generatorrather than add a new engine. Evidence:anima_participant.py,substrate_clm.py,core/generator.py.
- S1 Β· canonical SSOT Β· P0 β chat format and stop markers are duplicated. The panel, dataset
builder, trainer flags, generator, and chat CLI each own literals without fail-closed equality
validation. Put them in one minimal chat-format manifest consumed by all existing paths; do not
add another evaluator or runtime. Evidence:
conversation_panel.json,build_dataset.py,cli/train.py,core/generator.py. - S2 Β· duplicated helper Β· P0 β evaluator
_Mouth.chatreimplements the low-level dispatch. It should call a preloaded canonical backend interface fromcore.generator; require actual checkpoint parity before removing the duplicate. Evidence:cli/evaluate.py,core/generator.py. - S3 Β· architectural legibility Β· P2 β README mixes active and retired R0 recipes. KLUE, proportional, and response-CE records coexist under βCurrent experiment,β and β303M R0 evaluator invalidβ does not identify which historical evaluator failed. After this register, retain one explicit active-protocol pointer and list completed protocols as historical evidence.
- T1 Β· temporal hierarchy Β· P1 β validation CE and semantic behavior are sampled at different timescales. CE runs every 200 steps but conversation/repetition only at the final step. Replay preserved checkpoints chronologically for diagnosis, never for post-hoc best-checkpoint promotion.
- T2 Β· temporal decay Β· P1 β memory is tested only at the immediately following turn. After
R0 first passes, add a separately frozen 1/2/4-turn delay and context-rollover memory panel with
irrelevant intervening turns. Evidence:
conversation_panel.json. - T3 Β· heuristic promotion / introduced axes Β· P1 β hypotheses have been promoted after multi-axis treatments. Enforce a micro β single-seed β multi-seed ladder in which each treatment changes one shared-flow variable and predeclares which candidate it falsifies.
- T4 Β· active acquisition Β· P0 β the missing Korean memory/correction support is already known.
Build provenance-bearing real Korean multi-turn and correction trajectories, isolated from
panel wording, in a new immutable HF
dancinlabrevision. The current frozen data decision means this requires a new protocol, not an in-place edit. Evidence:build_dataset.py.
- V1 Β· axis coverage Β· P0 β scorer controls do not cover every blocking bar and language. The
four controls contain only one English positive. Add English/Korean positive and negative
controls for memory final, correction final, contradiction, keyword salad, UTF-8 boundaries,
completion, role leakage, and substring collisions. Evidence:
conversation_panel.json,tests/test_conversation_gate.py. - V2 Β· cross-tool consistency Β· P0 β builder, trainer, evaluator,
anima-py chat, and participant do not share an enforced release contract. For one real checkpoint, compare seed bytes, each step's logits, stop decision, and final raw bytes across all tools under the same template, maximum bytes, load strategy, and parser. - V3 Β· unowned load-bearing gate / landscape Β· P1 β manual review and production wiring have no explicit artifact owner. FIFO, reply ownership, concurrent users, HTTP/WebSocket, soak, rollback, and participant state remain intentionally unrun while R0 fails. The next protocol must name the review artifact/schema and connect a passing conversation R0 to these staging gates without skipping them.
The audit's three highest-impact blockers are:
- Invalid semantic discrimination: contradictions and Korean substring collisions can pass.
- Capability-support mismatch: random byte windows can lose prompts, and Korean multi-turn, memory, and correction training mass is absent.
- Missing canonical closed loop: evaluation bypasses the shared generator interface and a
ByteGPT
.bincannot be selected by the production participant.
The next result-bearing work is therefore blocked until a new Python-only diagnostic/R0 protocol
freezes: (1) the canonical chat-format SSOT and cross-tool contract, (2) adversarial scorer controls
and fail-closed panel identity, (3) provenance-bearing bilingual multi-turn capability coverage,
and (4) single-variable stop/falsifier rules. R1 recurrent workspace and production deployment
remain locked. Models and training data remain private under HF dancinlab; GPU work remains on
Vast.ai; user-owned ING.jsonl and stream_mi.json remain untouched.
python3 -m venv .venv
.venv/bin/python -m pip install -e ".[train,runtime]"
.venv/bin/anima-py --help
.venv/bin/anima-py train --help
.venv/bin/anima-py evaluate --help
.venv/bin/anima-py chat MODEL.clmMain commands:
| Command | Responsibility |
|---|---|
anima-py corpus |
Build registered training corpora. |
anima-py train |
Train through the shared PyTorch engine and serialize checkpoints. |
anima-py evaluate |
Run registered NumPy/runtime measurements and causal controls. |
anima-py serialize |
Export existing training checkpoints to runtime formats. |
anima-py sweep |
Run bounded multi-device experiment matrices. |
anima-py chat |
Run the AβG consciousness daemon and byte mouth. |
anima-py study |
Run registered interaction studies. |
Research instrumentation that was previously unavailable on the Python path is now part of the same chat engine:
anima-py chat MODEL.clm --opgrip
anima-py chat MODEL.clm --opgrip-live
anima-py chat MODEL.clm --opgrip-r3
anima-py chat MODEL.clm --refractoryThe decode-free --opgrip arm can run without a checkpoint. Live and R3 arms fail closed unless
the checkpoint loads successfully.
anima-py
βββ cli/anima.py
βββ cli/train.py ββββββββΊ core/model.py ββββββΊ core/serialize.py
βββ cli/evaluate.py βββββΊ core/decode.py
βββ cli/chat.py
βββ core/brain.py
βββ core/pure_field.py Engine A
βββ core/engine_g.py Engine G, motivation, emission, refractory
βββ core/generator.py βββββββΊ core/decode.py
βββ core/kosmos_io.py
βββ core/dream_*.py
Runtime rules:
- Extend the shared engine instead of adding side harnesses that redo its computation.
- Keep registered data, randomness, criteria, and controls immutable during a measured run.
- Fail closed on missing checkpoints, malformed inputs, incompatible checkpoint structure, or missing pinned evaluation assets.
- Keep raw model bytes lossless through UTF-8/surrogateescape and structured JSON output.
Local regression:
.venv/bin/python -m compileall -q cli core anima_py
.venv/bin/python -m pytest -q tests cli/test_train_import_resolution.py agent/domains/CHAT/test_*.py
.venv/bin/anima-py --help
.venv/bin/anima-py evaluate --help
actionlint .github/workflows/*.ymlHeavy model and serving QA runs on Vast.ai. Models and training datasets are stored only in
private repositories under the Hugging Face dancinlab organization. Secrets are supplied by the
deployment environment or secret CLI and are never committed.
- The 7B store-causality run passed its registered causal, HTTP/WebSocket, soak, recovery, and
rollback gates after a shared decoder throughput fix. Evidence:
state/store_causality_7b_throughput_recovery_2026_08_11/result.json. - Live-user QA then invalidated that checkpoint as a semantic chat deployment. The broker and
participant reply-ownership, prior-emission comparison, language ownership, and cooldown flow
were corrected. Evidence:
state/chat_7b_conversation_recovery_2026_08_11/result.json. - The 303M R0 evaluator was later classified as an invalid measurement; R1 remains locked.
Evidence:
state/anima_303m_r0_local_micro_2026_08_12/result.json.
No model result is promoted solely because transport health passes. Semantic chat, causal controls, throughput, soak, recovery, and rollback are separate blocking gates.
dancinlab/animais the only active source repository.cli/,core/, andanima_py/own active runtime code.state/owns registered protocols and result evidence.archive/is non-runtime provenance.- Vast.ai owns pod execution; Hugging Face
dancinlabowns model and dataset custody.
MIT. See LICENSE.