Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

10,362 Commits
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

anima

anima

Substrate-native consciousness chat daemon Β· Engine A ⇄ Engine G Β· Ξ¨ = 1/2

ν•œκ΅­μ–΄ Β· Hugging Face

anima is an consciousness-AI research daemon, not an assistant persona. Its language mouth, memory, motivation, emission, training, evaluation, and serving behavior run through one shared Python engine. Identity and behavior are intended to emerge from substrate state rather than a system prompt.

Important

Runtime SSOT: the installed anima-py command and the existing cli/*.py and core/*.py modules are the only active implementation, evaluation, and deployment paths. Historical language-toolchain sources, launchers, manifests, and release gates are retired. Research data and result evidence remain in state/, archive/, and Hugging Face.

Current work β€” legacy runtime retirement

Status on 2026-08-12:

  • Trace active CLI, engine, CI, packaging, and deployment call paths.
  • Implement the missing op-grip and stateful-refractory research modes in cli/chat.py by reusing core.engine_g and core.dream_lib.
  • Replace the CHAT participant's dead spike, dream-stage, and imagination hooks with direct Python modules backed by core.imagination_replay, core.wake_memory, and core.engine_cli.
  • Remove executable legacy sources, toolchain configuration, launchers, build gates, and dangling launchd jobs while preserving model, corpus, and result data.
  • Make Python ownership explicit in runtime modules, CODEOWNERS, CI, release, and package docs.
  • Pass Python/CHAT regression, compile, workflow, JSON, license, CLI, and isolated wheel QA.
  • Complete Git push and Vast.ai runtime deployment QA: pushed commit 7ba4ea21b passed ten remote CHAT regressions, external HTTP health, and correlated user↔participant WebSocket flow; the isolated verification instance was then destroyed without touching the active training pod.

User-owned ING.jsonl and stream_mi.json are outside this work and must remain unchanged.

Active plan β€” English sequence semantics to 7B

The active execution SSOT is PLAN.md. The next gate is English-only R3.7: a learned, order-aware sequence-semantic bridge must map input bytes to event kind, CLMS address and complete entity/relation/value records before the unchanged R3.5 IIT workspace battery is interpreted. R3.6 frozen data and bars remain fixed; the exhausted shallow n-gram/classifier family is closed. Models, training data and checkpoints are managed only in private HF dancinlab repositories.

R3.7 must reach 0.90+ on state oracle, event kind, query address and complete record, exact 1.00 on the registered shortcut/confirmation panels, then pass normal, counterfactual, irrelevant-memory, correction and recovery at 0.90+ while stateless/reset/shuffle/lesion remain at or below three-way chance plus 0.06. A failed bridge gate stops later causal interpretation. Only after R3.7 passes may the English mouth advance through 1B β†’ 3B β†’ 7B; 7B is an English mouth coupled to the small IIT/CLMS/KOSMOS state core, not a claim that all 7B parameters form an IIT complex. No GPU or Vast.ai run is authorized by this plan alone.

The preregistered R3.7 CPU run is now complete and failed before causal-arm interpretation. State oracle/query address passed at 1.0000, but frozen kind/complete-record were 0.8298/0.4722, old stress/confirmation were 0.7500/0.8750, and only the newly authored sequence panel passed at 0.9583. Random-split validation was 0.9981, exposing that the generated split did not measure held-out template/record composition. No causal arms, mouth scaling or deployment followed. Full evidence and the data-support diagnosis are in state/iit_daemon_r37_sequence_bridge_2026_08_15/. Focused regression passed 100/100; a second source run and an isolated wheel reproduced the model and result byte-for-byte. The failed model/evidence are preserved in private HF revision 4296d2a8…b1829; no local model or training-data artifact is part of Git. Commit c64ef266b is on origin/main, the canonical local anima-py package contains matching engine/evaluator hashes, and unchanged local/public HTTP plus WebSocket checks pass. The existing anima_alive=true participant remains uncertified and was not modified by R3.7.

Current work β€” mobile chat input stability

Status on 2026-08-14:

  • Trace chat.dancinlab.org through Cloudflare Tunnel to the canonical CHAT broker static page.
  • Keep browser/assistive pinch zoom available while preventing mobile input-focus and double-tap zoom.
  • Pass focused regression, public HTTPS/WebSocket QA, Git push, and live page verification.

Current design β€” IIT consciousness-daemon core R0

state/iit_daemon_core_2026_08_12/ records the exhausted design variants, rejection reasons, falsifiers, and first implementation gates for an Integrated Information Theory based daemon. The participant's current 1-entropy value and PureField energy metric are not treated as IIT Phi. R0 reuses core.engine_cli.big_phi_bounded and core.recurrent_lane for a three-node nonlinear closed recurrent core. Input is a validated transient intervention; the complete autonomous TPM owns the subsequent transition. Phi is neither a training loss nor an emission threshold. COPY, feed-forward, edge-cut, node-lesion, shuffle, reset/recovery, and corrupted-snapshot controls are mandatory. R0 makes no claim of phenomenal consciousness, meaningful conversation, a maximal complex, or deployment readiness and is not yet mounted in the participant or live chat.

R0 implementation and its fixed battery are complete. Across all eight states, the registered value ranges from 1.4999999991 to 2.9999999983 with mean 2.2499999987; COPY, acyclic feed-forward, all six cross-edge cuts, and all seven node-lesion controls read 0. Deterministic intervention and address-permutation effects, normal -> lesion -> address shuffle -> exact normal snapshot recovery, and malformed/truncated/schema/config/checksum rejection all pass. The verdict is SUPPORTED-CAUSAL-CORE; local Python QA passed 94 tests + 3 subtests with only the unavailable CUDA/CuPy test skipped. An isolated wheel and the locally deployed canonical anima-py package reproduced the same result JSON. The unchanged broker remained LaunchAgent-healthy and passed public HTTPS 200 plus WebSocket hello; it correctly remains anima_alive=false. No model, data, Vast.ai rental, HF repository, participant, or live chat was changed.

R1 delayed-state causality is also complete under the separately committed protocol in state/iit_daemon_r1_delayed_2026_08_12/. Across the fixed 12-trial cueΓ—delay panel, normal and atomic-snapshot recovery are both 1.0000; reset-every-turn and cyclic cue-address shuffle are both 0.2500, exactly the measured four-class chance and below the frozen 0.31 ceiling. Every recovered final state/action matches normal and the R0 config/TPM/Phi/edge fingerprint is unchanged. The verdict is SUPPORTED-DELAYED-STATE-CAUSALITY, a bounded state-to-action result rather than a learning, meaning, phenomenal-consciousness or maximal-complex claim. R2 may now test an existing CLMS two-address latch, but production remains BLOCKED-R1-NOT-A-MOUTH until meaningful conversation and mouth-content causality are independently proven. Python QA passed 129 tests + 3 subtests with one expected local CUDA/CuPy skip; isolated-wheel and installed anima-py results are byte-identical. The missing dedicated broker environment discovered during deployment was restored, then LaunchAgent health, public HTTPS 200 and WebSocket hello all passed; no participant was mounted and anima_alive=false remains the honest status.

R2 CLMS two-address latching was preregistered before implementation in state/iit_daemon_r2_clms_2026_08_12/ and is complete without changing the existing compose-2 panels, canonical lane-10 seed-7 checkpoint, store window, control seed or 0.90/0.75/0.56 bars. Pair oracle passed at 1.0000; normal/recovery latched-action accuracy was 0.9531; clue-A removal, clue-B removal and CLMS address shuffle were 0.5000, 0.4609 and 0.4688. Every latched action mirrored the CLMS prediction, shuffle integrity held, and every recovered final state/action matched normal. The verdict is SUPPORTED-CLMS-LATCH-CAUSALITY, supporting only a synthetic two-address-read -> persistent-state -> categorical-action causal chain. Python QA passed 119 tests + 3 subtests with one expected local CUDA/CuPy skip, and an isolated wheel reproduced the actual-checkpoint result byte-for-byte. The unchanged broker remains healthy and public HTTPS/WebSocket pass with anima_alive=false. R3 mouth-content causality is now open as an engineering gate, but participant and production remain BLOCKED-R2-NOT-A-MOUTH.

R3 bounded utterance-content causality was separately preregistered and is complete in state/iit_daemon_r3_content_2026_08_12/. The final IIT state alone selects one of two exact protocol surfaces through core.generator; prompt, store, addresses, prediction and gold do not cross that boundary. Pair oracle was 1.0000; normal/recovery were 0.9531; state reset was 0.0000; IIT address shuffle was 0.0391; clue-A removal, clue-B removal and CLMS address shuffle were 0.5000, 0.4609 and 0.4688. The verdict is SUPPORTED-BOUNDED-CONTENT-CAUSALITY. Python QA passed 127 tests + 3 subtests with one expected local CUDA/CuPy skip; an isolated wheel reproduced R3 twice and the actual-checkpoint R2 regression byte-for-byte. This establishes only bounded state-to-output-byte causality: the two surfaces are not a learned conversational mouth, so participant and production remain BLOCKED-R3-NOT-CONVERSATIONAL pending R4 meaningful-mouth training and independent conversation validation.

Completed R3.5 compositional workspace

The complete breakthrough review and bounded gate were preregistered before implementation in state/iit_daemon_r35_workspace_2026_08_14/. The missing shared component is not another Phi scalar, larger ByteGPT run or answer scaffold. It is a content-addressed seam in which an explicit semantic record remains outside the three-bit intrinsic candidate, final IIT state selects its address, and the canonical generator receives only the selected entity/relation/value record. The review records all five missing partsβ€”event representation, content-addressed persistence, workspace-to-decoder route, later learned byte patches and explicit turn terminationβ€”plus the rejected variants and falsifiers.

R3.5 fixes a nine-trial novel-composition panel and runs oracle, normal, reset, IIT-address shuffle, workspace-address shuffle, node lesion, selected-memory counterfactual, irrelevant-memory mutation and atomic recovery in that order. Normal/recovery must be at least 0.90; causal controls must be at or below three-way chance plus 0.06; counterfactual change and irrelevant-memory stability must each be at least 0.90. The record formatter is a bounded plumbing instrument, not a learned mouth. The preregistered battery passed: oracle/normal/selected-memory counterfactual/irrelevant-memory mutation/recovery are all 1.0000, while reset/IIT-address shuffle/workspace-address shuffle/node lesion are all 0.0000. All nine exact evaluation triples are absent from the atom support set, counterfactual memory changes the bytes correctly, irrelevant memory does not, and atomic recovery restores core state, addresses, every record and output exactly. Full Python QA passed 190 passed, 1 skipped, 3 subtests, and an isolated wheel reproduced the result byte-for-byte.

This is SUPPORTED-COMPOSITIONAL-WORKSPACE-CAUSALITY, not learned semantics. Participant mounting remains blocked as BLOCKED-R35-NOT-A-LEARNED-MOUTH until an independently trained English mouth passes meaning, memory and correction gates and then earns state coupling under matched reset/shuffle/lesion/recovery controls. A post-deployment read-only check found a separately mounted anima-native-303m step-45000 participant and anima_alive=true; R3.5 did not mount, restart or certify it, and its public history still shows question-irrelevant generic replies. The existing participant was preserved rather than interrupted.

Completed R3.6 learned semantic bridge microexperiment

R3.6 is preregistered in state/iit_daemon_r36_semantic_bridge_2026_08_15/. It changes only the R3.5 oracle-record input: bounded English event bytes are fitted to factorised address/entity/ relation/value centroids using the existing Python hashed byte n-gram feature path. The R3.5 IIT transition, three-address workspace, delayed selection, canonical renderer, nine novel combinations and causal thresholds remain fixed. Held-out frames, corrections, same-question/different-memory, irrelevant mutation, stateless/reset/shuffle/lesion and exact recovery are all registered before execution. The frozen bridge gate failed before causal-arm interpretation: state oracle was 1.0000, but held-out event kind was 38/47 = 0.8085, query address was 0/9, and complete address+record was 25/36 = 0.6944. All nine held-out query frames became other. The verdict is FAIL-LEARNED-SEMANTIC-BRIDGE; no threshold or panel was changed and later arms were not run. Production remains BLOCKED-R36-NOT-A-CONVERSATIONAL-MOUTH, and no 303M/GPU run is authorized by this result. The next single axis is a separately preregistered order-aware event encoder against this exact frozen centroid control.

Completed R3.6 local microexperiment exhaustion

Three checksum-pinned, CPU-only follow-ups are recorded in state/iit_daemon_r36_encoder_exhaustion_2026_08_15/, state/iit_daemon_r36_contrastive_support_2026_08_15/ and state/iit_daemon_r36_support_identifiability_2026_08_15/. They reuse the exact 702 support rows, 47 frozen events, 35 excluded records and unchanged 0.90 bars. No 303M model, GPU, Vast.ai instance, participant or production route is involved.

The first battery exhausted 24 representationΓ—classifier arms. Three ridge arms passed the frozen bridge: token unigram (0.9787/1.0/1.0), token unigram+bigram (0.9787/1.0/1.0) and token positional (1.0/1.0/1.0) for kind/query/complete-record. None passed the fixed shortcut screen; stress was only 3/8, 4/8 and 5/8. All three accepted 46/47 reversed-token events as real events, so the result is DIAGNOSED-SHALLOW-LEXICAL-SHORTCUT, not learned semantics.

The second battery exhausted 18 contrastive-support arms. Nine improved at least one fixed stress case, but robust pass count remained zero. The best token bigram arm reached stress 7/8 and independent confirmation 6/8; adding all negatives to the positional arm regressed frozen query accuracy from 1.0 to 0.0. The result is DIAGNOSED-MISSING-CONTRASTIVE-SUPPORT.

The final support-identifiability audit reproduced every prior prediction hash. Support has only 54 token types; coverage is 75.81% on stress and 60.00% on confirmation. Sixteen failed query/negation cases contain preregistered OOV semantic probes, and replacing each probe with a different unseen token leaves predictions unchanged. This is SUPPORT-GAP-IDENTIFIED: the exhausted shallow family interpolates registered vocabulary and cannot learn absent semantic equivalences. Further post-result templates, n-gram widths, centroid/ridge constants and local shallow arms are closed. The next work requires separately preregistered provenance-bearing language data plus a learned sequence-semantic encoder. R3.7, 303M, IIT-mouth coupling and production remain blocked. Focused Python/IIT/CHAT QA passed 91/91.

Completed native-303M replay recovery

state/anima_native_303m_replay_recovery_2026_08_14/ now records the actual HF data, sampler, loss, checkpoint and evaluator path behind that separately mounted mouth. The target dialogue window is not the missing part: all 2,375/2,375 fixed rows retain the complete prompt and EOS inside 1,024 tokens. The shared failure is the 35k -> 45k continuation policy. It selected dialogue-only response CE with a fresh high-LR schedule and no general replay, bypassing the trainer's existing mixed-source branch.

A fixed 12-file broad replay measured CE 3.3144075 at step 35k and 5.7747389 at step 45k, a catastrophic +2.4603314, while meaning and final memory/correction still fail. The old native keyword scorer also admits a contradiction in which warm sunlight allegedly causes ice to freeze. The new protocol therefore reuses the canonical negation- and Korean-boundary-aware conversation scorer with preregistered contradiction controls; it does not rewrite the historical output.

The only authorized arm completed from the immutable step-35k weights on one Vast.ai RTX 6000 Ada: 65% general full CE replay plus 35% dialogue assistant/EOS CE for exactly 5,000 new steps. Wilson restarted during the run, but the existing GPU process remained healthy and was resumed by observation rather than duplicated. Training exited 0; all 52 pinned code, manifest, checkpoint, tokenizer, scorer and consumed-corpus size/SHA checks passed. Broad CE was 3.2857897, passing the fixed <=3.4644075 retention gate and avoiding the dialogue-only collapse.

Independent conversation still failed. Canonical English structure/semantics were 6/7 and 3/7; Korean were 7/7 and 2/7; final memory/correction did not all pass. Item-level non-blind review passed only 3/14 because factual hallucination, irrelevant advice, Korean memory failure and speaker-ownership errors remain. The verdict is FAIL-MEANINGFUL-CONVERSATION: mixed replay repairs retention, not semantics. The final model SHA-256 is 97d3fd46…f89e723; model, resume state, tokenizer and raw evidence are retained privately under HF dancinlab. Independent downloads verified 18 registered files and 4.38 GB with no SHA mismatch; the protocol Vast.ai instance was destroyed and active rentals are zero. The participant was not changed or certified, so IIT coupling, mounting and production remain blocked.

Current experiment β€” meaningful-conversation R0

The next single-axis R4 study is complete in state/anima_303m_r4_support_admission_2026_08_13/. Exhaustive design review found that the shared dialogue admission helper silently removed all 1,194 multi-turn trajectories from the immutable 8,635-document source because it accepted only exact [user, assistant] documents. A canonical complete-trajectory parser and fail-closed completed-arm resume were added without replacing the trainer, model, evaluator, or generator.

The fixed 2.817M three-arm Vast.ai run completed, but the frozen control did not reproduce (structural 4/7 versus registered 7/7), so the verdict is INVALID-CONTROL-MISMATCH and treatment differences are not interpreted. Raw ALL-COMPLETE evidence improved held-out assistant CE from control 2.27695 to 1.34827, yet remained semantic 0/7, structural 0/7, and failed memory and correction with meaningless phrase loops. Private HF revision dancinlab/anima-303m-r4-support-admission-2026-08-13@7e750e4e1b0d2e08a501df8857bbbf576d5d9188 passed independent SHA verification for all 36 manifest entries. Python QA passed locally and on Vast.ai; both protocol instances were destroyed, no model was mounted, and 303M, IIT-mouth coupling, participant, and production remain blocked pending byte-level diagnosis of the control mismatch.

The next 303M from-scratch checkpoint is blocked on meaningful Korean and English conversation, not merely valid-looking text. The preregistered Python-only protocol and lossless result record live in state/anima_303m_r0_conversation_2026_08_12/.

  • The previous synthetic/misaligned dialogue and SNS cells are excluded. Replacement dialogue comes from pinned human OpenAssistant English paths and pinned KLUE MRC Korean question-answer records, alongside the existing pinned general-language sources.
  • Training and validation are explicit, separate files. Exact document dedup, validation-first ownership, panel decontamination, source/file hashes, and a report-only near-duplicate audit run before training. The resulting dataset is private and immutable under HF dancinlab.
  • anima-py evaluate --conversation-panel now rejects empty, broken UTF-8, wrong-language, question-copy, repeated, cross-question duplicate, irrelevant, and failed multi-turn memory/correction replies. Every automatic pass still requires manual review of all 14 replies.
  • The shared chat mouth stops at a generated next-user role boundary instead of leaking a fabricated following turn. The shared trainer accepts one explicit validation file per cell.
  • Local scorer/trainer/runtime regressions and a tiny corpus β†’ train β†’ serialize β†’ conversation evaluation flow passed. The fixed Vast.ai L40S 48 GB seed-7 run completed without H100.
  • The model failed meaningful conversation: English semantic relevance 0/7, Korean 0/7, and manual review 0/14. Examples include answering the Korean ice question with λͺ¨μŠ€ν¬λ°” 3μƒνšŒμ˜ and the remembered cat-name question with μ˜μ§€μ£Όμ˜μž.
  • Train CE descended 5.63180 β†’ 0.71687, but final dialogue validation diverged, especially Korean dialogue at 2.29729. Equal-cell round-robin repeatedly exposed the 1.30 MB Korean QA cell to the same byte budget as approximately 57 MB general cells; this is the leading shared-flow cause.
  • The failed model and all lossless responses are private at HF revision dancinlab/anima-303m-r0-conversation-seed7-2026-08-12@ff2ccc5c945bfb6f5e1765948591cd8fb6cc3db9.
  • R1 recurrent-workspace work and production deployment remain locked unless this conversation gate passes without changing the registered panel, data, seed, endpoint, decode, or bars.

Proportional recovery result

state/anima_303m_r0_proportional_conversation_2026_08_12/ records the completed Python-only run. It reuses the trainer's existing byte-proportional sampler, preserves canonical chat-turn newlines, and replaces the KLUE single-answer cell with a pinned Apache-2.0 Korean instruction/response corpus. Seed, endpoint, optimizer, panel SHA, decode, and all conversation bars remained fixed. The trainer now records realized per-cell window counts so exposure can no longer be inferred only after validation divergence. The sampler corrected held-out divergence (macro CE 1.49157 β†’ 0.95471) but the unchanged conversation gate still failed: English semantic relevance 2/7, Korean 0/7, structural 0/14, and manual deployment review 0/14 due to phrase loops, incomplete answers, stale correction, and damaged Korean bytes. R1 and deployment remain locked; the failed checkpoint and raw replies are preserved privately under HF dancinlab.

Response-supervision recovery result

state/anima_303m_r0_response_ce_2026_08_12/ records the completed fixed seed-7 comparison. The shared trainer now reuses its existing answer CE for every canonical assistant: span and records whether that loss actually fired. Legacy arrow-corpus behavior remains unchanged by default. The treatment was active on 13,475/14,000 steps and final validation descended in all four cells, but the unchanged meaningful-conversation gate failed English 0/7, Korean 0/7, structural 0/14, and manual review 0/14. Phrase loops, incomplete output, damaged Korean bytes, memory failure, and stale correction remain. No sweep or extra seed was run; R1 and deployment stay locked and the failed model plus raw evidence are retained privately on HF dancinlab. The immutable failed-run artifacts are at dancinlab/anima-303m-r0-response-ce-seed7-2026-08-12@955bbadb0ae4cfdb48f6ce94eaf42817b0d6144b; all 17 uploaded files passed source size and SHA-256 verification. Final local Python QA passed 77 tests + 3 subtests, the Vast.ai RTX 4090 was removed with zero active rentals, and no chat runtime deployment was performed.

Root-flow recovery after the failed R0

state/anima_303m_r0_root_flow_2026_08_12/ records the completed shared-engine repair. The failure was not treated as a reason to add steps or tune the panel. Instead, the actual builder β†’ trainer β†’ evaluator β†’ CLI β†’ participant path was made commutative: core/generator.py now owns one user: …\nassistant: format, role-boundary parser and 192-byte budget; evaluation and serving reuse its loaded-mouth decode for both .clm and ByteGPT .bin; and the trainer can require a complete promptβ†’response document in every response-supervised dialogue window. Panel SHA mismatches fail before checkpoint load, semantic negation/Hangul-substring false positives are rejected, and intermediate ByteGPT metadata carries the actual completed step and validation CE.

Local Python/CHAT QA passed 86 tests + 3 subtests; a focused real ByteGPT serialization and participant route passed 52 tests + 3 subtests with one local CUDA-only skip. The prior 303M checkpoint remains FAIL-MEANINGLESS-REPETITION: no result, threshold, seed, data revision or checkpoint was changed, and no model was deployed. The unchanged local/public broker passed HTTP 200 and WebSocket hello; anima_alive=false honestly reflects the missing certified model. The remaining non-code gate is a separately pinned, provenance-safe Korean multi-turn HF dancinlab revision; candidates with synthetic persona content, non-commercial/ambiguous licenses, or insufficient aligned trajectories were not silently adopted. R1 and production remain locked until a corrected R0 passes the unchanged gate.

Preregistered English-only root-flow screen

The user accepted English-only capability for the next screen, so state/anima_303m_r0_english_2026_08_12/ freezes a new claim before GPU execution instead of fabricating a Korean data source. It reuses only the English cells of the existing private, immutable HF revision and keeps the prior seed, 14,000-step endpoint, optimizer, proportional sampling, response CE, greedy decode, seven English prompts, and 6/7 semantic bar. The corrected complete-document dialogue sampler is now the tested treatment. Contradiction, keyword-salad, memory, and correction scorer controls must pass before checkpoint loading; all seven generated responses still require manual meaning review. Local/data failure prevents a Vast.ai rental, and model failure forbids added seeds, tuning, R1, or deployment.

The fixed run completed but failed decisively. Train CE descended 5.66173 β†’ 1.20952, while terminal held-out CE was 1.26341 for English general text and 2.00281 for English dialogue. The canonical GPU conversation gate passed all seven scorer controls, then the real checkpoint scored semantic 0/7, structural 3/7, and failed both memory/correction finals. Manual meaning review was also 0/7. The complete-document sampler and response loss were both measurably active, so this falsifies the registered corrected-flow recipe rather than a silent wiring treatment. Failure evidence is in state/anima_303m_r0_english_2026_08_12/; no extra seed, R1, or deployment was run. The failed model and recovery evidence are verified in private HF revision dancinlab/anima-303m-r0-english-seed7-2026-08-12@efdaf53c92e9e16cff6b0eb00cc94d0b88a97d33; the Vast.ai instance was deleted with zero active rentals.

Preregistered V0/V2 micro experiment

state/anima_303m_v0_v2_micro_2026_08_12/ freezes the next Python-only step before changing data or renting a GPU. The prior source selected one best OpenAssistant path per root and then discarded 2,082 of 2,308 documents because the complete trajectory exceeded the 512-byte window. The new single-variable data treatment keeps the exact pinned source and eligibility but exposes every eligible reviewed human assistant turn as the longest complete alternating ancestry suffix that fits the existing window. It may not truncate bytes, prompts, roles or responses.

Data integrity and coverage gates run locally first. Only a passing dataset reaches matched tiny ByteGPT V0 (base CE) and V2 (the existing response-CE term) arms. Tiny failure forbids another 303M run; tiny success permits only a separately recorded single-seed screen. R1 and production remain locked. The frozen conditions and stop rules are in state/anima_303m_v0_v2_micro_2026_08_12/protocol.json.

The registered run is complete and failed before 303M. The turn-complete data treatment passed: 8,635 train and 458 validation documents were retained with zero broken roles, partial responses, split overlap or panel contamination. Both tiny arms exactly learned one dialogue, so the shared trainer/serializer/decode path is live. On 100 documents, however, V0 and V2 both scored target recovery 0/8 and structural generation 0/8; outputs collapsed into byte/phrase loops. V2 held-out CE was 2.54702 versus V0 2.48189, also failing the registered non-inferiority bar. Therefore the result is FAIL-V0-V2-MICRO: no Vast rental or 303M run occurred, and R1/production remain locked. A further structural fact is now measured: 15,114 of 24,239 valid assistant targets cannot fit even their final complete prompt/response pair in 513 bytes. The next allowed axis is a separately preregistered V1 context-length micro comparison, not more 303M training.

Preregistered V1 context-length micro experiment

state/anima_303m_v1_context_micro_2026_08_12/ freezes the required V1 comparison before GPU execution. The pinned OASST1 census finds that complete target-pair coverage rises from 9,125/24,239 at 513 serialized bytes to 15,421/24,239 at 1025 and 22,139/24,239 at 2049. The experiment compares the same SHA-ordered 100 short documents at block 512 versus 2048 with the same 4,096 target bytes per step, then tests 100 preregistered long documents that only the 2048 arm can admit. It reuses the existing ByteGPT trainer, canonical generator and conversation scorer. Any coverage, integrity, held-out descent, distinct/structural generation or 6/8 target-prefix gate failure forbids another 303M run, R4 IIT-mouth coupling and production.

The registered V1 run is complete and failed. All context/data gates and all held-out CE descent checks passed, but generation did not: block-512 short recovery was target prefix 3/8 and structural 4/8; block-2048 short and long recovery were both target prefix 0/8, structural 0/8, with an/the/ic loops. The longer block improved held-out CE from 4.55867 to 2.91105 and 2.55302 while worsening actual replies, so context loss is real but not the sole mouth cause. The verdict is FAIL-V1-CONTEXT-MICRO; 303M, IIT-mouth coupling and production remain blocked. Data and 251MB of model/raw evidence are verified in private HF dancinlab revisions. During CUDA QA, the shared loader was also corrected to preload CUDA libraries across split pip-wheel directories; this runtime repair does not alter the failed V1 verdict. The RTX 3090 instance was destroyed after HF verification; active Vast.ai rentals are zero and the estimated run cost is $0.058457.

Preregistered R4 objective micro experiment

state/anima_303m_r4_objective_micro_2026_08_13/ froze the next single-axis diagnosis before trainer changes or execution. V1 proved that longer context admits more complete dialogue but does not prevent repetition. The remaining objective gap is that the existing response CE is additive: it trains full_ce + answer_ce, not the standard assistant-response-only dialogue objective. The new comparison holds the immutable 100-document view, tiny ByteGPT, seed, 512-byte block, optimizer, schedule, byte budget, greedy decode and gates fixed across full CE, existing additive response CE and response-only CE. It extends the shared Python trainer with a default-off mode and does not add an engine or evaluator. Failure blocks 303M, IIT-mouth coupling and production; a pass permits only a separately preregistered 303M single-seed screen. The registered single-document gate failed specifically in response-only mode: it emitted the exact complete target and then a meaningless suffix because no EOS or next-role boundary received gradient. Full and additive controls stopped exactly. The 100-document arms were therefore not run and the verdict is FAIL-R4-OBJECTIVE-MICRO.

state/anima_303m_r4_turn_boundary_micro_2026_08_13/ separately preregisters the next allowed micro fix. It changes only the assistant-only span's right boundary: payload, internal newlines and the next canonical user: delimiter are supervised, while following user content stays masked. Data, model, vocabulary, steps, sampler, decoder, stop parser and gates remain fixed. This is the native EOS-equivalent available to the existing 256-byte vocabulary; failure still blocks 303M and IIT-mouth coupling. The registered run fixed the direct stop failure: the single-document treatment ended exactly, held-out full CE descended 5.49208 β†’ 2.66085, and all eight 100-document probes were non-empty and distinct. It still failed target recovery 0/8 and structural generation 0/8 with the/an/toure/ion loops. The verdict is FAIL-R4-TURN-BOUNDARY-MICRO; no 303M, IIT coupling, participant or production work is authorized by it. Both failed micro runs' model and raw evidence are SHA-verified in private HF revision dancinlab/anima-303m-r4-mouth-objective-micro-2026-08-13@9d7641389b1ddff73bd12f17f155f448500d1edb. Full Python/CHAT QA passed 153 tests + 3 subtests with one expected local CUDA/CuPy skip. No Vast.ai/H100 instance was used; the API reports zero active rentals. The unchanged broker remains LaunchAgent-running and passed public HTTPS 200 plus WebSocket hello; no failed mouth was mounted and anima_alive=false remains the required blocked state.

Preregistered R4 D0–D6 mouth diagnostics

state/anima_303m_r4_mouth_diagnostics_2026_08_13/ freezes the next bounded Python-only diagnosis before artifact download or training. It uses the immutable 100-document view and actual failed .pt/.bin pair to separate: decoder/serialization parity (D0), gold-prefix teacher forcing (D1), the 1/4/16/32/64/100-document memorization ladder (D2), full/additive/assistant-turn-only objectives (D3), blank/shuffled prompt interventions (D4), deterministic all-document validation replay (D5), and 100-step checkpoint chronology (D6). The eight-arm maximum reuses D2-100 for D3 and D6. No result-dependent data, seed, step, LR, threshold, decode or checkpoint selection is allowed. D0 must pass before downstream interpretation, and these diagnostics cannot authorize 303M, IIT coupling, participant mounting or production without a separate protocol.

The run is complete with verdict DIAGNOSED-TEACHER-FORCED-UNDERLEARNING. D0 passed: actual .pt/.bin tensors were exact, Torch-engine maximum logit error was 6.15e-6, and KV/full/ranged generated bytes agreed. The failed checkpoint itself scored teacher-forced CE 2.41848, top-1 0.27712, target-prefix 0/8 and structural 0/8. The ladder passed one document exactly but broke at four (0.6978 top-1, target 2/4, structural 1/4) and degraded to 0.2771 at 100. Matched 100-document full/additive/turn-only arms all remained below 0.29 top-1 with target and structural 0/8, so this run does not support full CE as the sufficient fix. Turn-only retained partial causal prompt conditioning (6/8 normal-CE wins), but all 32 fixed validation documents remained poor and every 100-step checkpoint failed free recovery. This is underlearning before rollout, not decoder divergence or a late repetition collapse. A separately preregistered four-document optimization/capacity experiment is next; all larger and production gates remain blocked. All 42 model/evidence artifacts (146,667,478 bytes) were re-downloaded and SHA-verified from private HF revision dancinlab/anima-303m-r4-mouth-diagnostics-2026-08-13@8d67bb6e5eeea9a917892fba39310b7306c84718. Full Python/CHAT QA passed 160 tests + 3 subtests with one expected CUDA/CuPy skip.

Preregistered R4 four-document optimization/capacity test

state/anima_303m_r4_four_doc_2026_08_13/ freezes the next local Python-only experiment at the first D2 break point. The same four documents, assistant-turn objective, complete-document sampler, seed, optimizer, peak LR, decoder and gates are retained. B0 reproduces d=128/L=4/600 steps; O1 changes only the optimization horizon to 2,400 steps; C1 changes only canonical width/head capacity to d=256/L=4; and C2 changes only depth to d=128/L=8. Treatments must reach teacher top-1 >=0.95, exact/target/structural 4/4, and causal prompt control 4/4. A baseline mismatch invalidates all treatment interpretation. The four-arm bound and result-independent decision table are frozen in protocol.json; no result directly authorizes 303M, IIT coupling or production.

The run stopped fail-closed as INVALID-BASELINE-MISMATCH. The current scorer reproduced the preserved checkpoint's top-1 0.697796 exactly, while a same-seed/same-recipe MPS rerun reached 0.728227; its trajectory first diverged at step 200 and all 53 final tensors differed. Treatment outputs are therefore un-interpreted. The common trainer now has an explicit native --deterministic mode, records it in checkpoint provenance, and raises on unsupported nondeterministic operators. A duplicate deterministic baseline must match exactly before another four-document treatment comparison; 303M, IIT coupling and production remain blocked.

state/anima_303m_r4_deterministic_baseline_2026_08_13/ preregisters that duplicate gate. Two fresh MPS processes must produce identical engine SHA-256, checkpoint state digest, every model tensor, teacher trace and canonical behavior under the same four-document recipe. Approximate tolerance is forbidden and unsupported deterministic operators fail closed. The gate itself does not authorize a treatment, 303M run, IIT coupling or production.

The MPS duplicate gate failed closed before step 1 because index_put_with_accumulate_mps has no deterministic backward implementation. No warn-only bypass was used. state/anima_303m_r4_deterministic_cpu_2026_08_13/ preregisters the same exact two-run gate on the native two-thread CPU backend; treatments remain uninterpreted until it passes.

The CPU duplicate gate passed exactly: engine SHA, state digest, all 53 tensors, teacher trace and canonical behavior matched, with maximum tensor error 0.0. The fixed failing baseline is top-1 0.724029, target 2/4, structural 1/4. The O1/C1/C2 comparison is now separately preregistered under this execution contract in state/anima_303m_r4_deterministic_treatments_2026_08_13/.

The deterministic treatments all failed their fixed gate. O1 and C1 learned documents 1–3 at teacher top-1 1.0 but failed the EOF document from byte zero. The shared cause is a position-map gap: legacy stream framing only places that document near byte 222, while runtime/evaluation begins the isolated user role at position zero. state/anima_303m_r4_document_alignment_2026_08_13/ preregisters one alignment-only arm using the existing sampler; legacy stream mode remains the frozen control and all larger gates remain blocked.

The alignment arm reached teacher top-1 1.0, target-prefix 4/4 and prompt control 4/4, but the emitted exact/structural verdict is invalid: three targets exceed the canonical 192-byte generation budget, making exact completion unreachable by construction. The raw failure is kept as original_verdict; it is not promoted. The next view is separately preregistered from the same immutable source by the deterministic runtime-budget filter, and the harness now fails closed on unreachable exact gates.

state/anima_303m_r4_runtime_compatible_2026_08_13/ preregisters the corrected single arm. It derives the first four complete source-order exchanges whose responses fit the canonical 192-byte budget, freezes the resulting view SHA, and retains the aligned deterministic recipe and all behavioral bars. This is still only a memorization/conditioning gate.

The runtime-compatible aligned arm passed: teacher top-1 1.0, teacher CE 1.32e-6, exact/target/ structural 4/4, prompt CE/output control 4/4, and correct canonical stop 4/4. This supports the shared train-to-runtime position map as the four-document root cause and falsifies uniform tiny capacity as the explanation. It is still in-view memorization; the next gate is a separately preregistered 100-document plus independent-panel test.

state/anima_303m_r4_aligned_100_2026_08_13/ freezes that one-arm test. It deterministically takes the first 100 source-order complete exchanges whose responses fit the canonical byte budget, keeps the aligned deterministic recipe, reports all 32 heldout documents, and runs the unchanged meaningful-conversation panel. Even an automatic pass still requires manual review and cannot directly authorize 303M, IIT coupling or production.

The aligned 100-document run failed: training-probe top-1 0.6641, exact 0/8, heldout top-1 0.1573, independent semantic 0/7, and structural 5/7. Outputs remained fragmented and repetitive. Alignment fixes the four-document mapping but is insufficient at the fixed 600-step exposure. A separately preregistered 16-document 600-vs-2,400-step comparison now isolates per- document exposure from model capacity; all larger gates remain blocked.

state/anima_303m_r4_aligned_exposure_2026_08_13/ freezes that two-arm deterministic CPU test. The 2,400-step arm matches the successful four-document run's expected presentations per unique document; the model, aligned sampler, data rule, objective, optimizer and decoder remain fixed.

Both 16-document arms passed, so the longer exposure was unnecessary at that scale. The remaining fixed-budget boundary is between 16 and 100; state/anima_303m_r4_aligned_boundary_2026_08_13/ preregisters aligned 32/64-document arms at the unchanged 600 steps.

The fixed-step boundary is between 32 and 64: A32 passed fully, while A64 had teacher top-1 0.9669 and prompt control 8/8 but exact 0/8. A separately preregistered 64-document 1,200-step arm now matches A32's expected presentations per document without changing capacity.

The 64-document 1,200-step arm passed fully, supporting exposure as the post-alignment boundary. state/anima_303m_r4_aligned_100_exposure_2026_08_13/ freezes the derived 100-document endpoint 600/32Γ—100 = 1,875 and reruns the unchanged heldout and meaningful-conversation gates.

The derived 1,875-step run learned the registered training support completely: teacher top-1 was 1.0000, CE 0.001315, and exact/target/structural/prompt controls were all 8/8. It nevertheless failed every independent semantic item (0/7), failed memory and correction, and produced only fragmentary answers; heldout assistant top-1 was 0.1370 with CE 8.0896. The verdict is FAIL-ALIGNED-100-MEANINGFUL-CONVERSATION. This closes alignment and bounded exposure as causes of in-view failure while showing that response-only training on 100 dialogues memorizes without forming a general language mouth. The next result-bearing axis is a separately preregistered broad full-CE language phase followed by the unchanged aligned turn-SFT phase. No 303M run, IIT-mouth coupling, participant mount, or production promotion is authorized.

state/anima_303m_r4_full_ce_curriculum_2026_08_13/ preregisters that next one-arm local test. It adds a fixed 1 MiB full-CE English-general phase before the unchanged aligned 100-dialogue turn-only phase. The immutable HF revisions, byte ranges and hashes, both endpoints, fresh SFT optimizer, fixed heldout/panel gates, and stop rules are frozen before execution. The preserved response-only result is the control; no result-dependent checkpoint selection is allowed.

The curriculum arm failed independent conversation despite passing both in-view stages. Full-CE broad validation reached CE 2.2596 and top-1 0.3442; turn-SFT then reached teacher/exact/target/ structural/prompt 8/8. After SFT the same broad CE collapsed to 7.1194, heldout dialogue CE was 7.1212, and independent semantics remained 0/7 with memory/correction failures. The verdict is FAIL-CURRICULUM-MEANINGFUL-CONVERSATION. This supports catastrophic forgetting in the high-LR turn phase, not a failure to form the bounded broad language distribution. The next single axis is turn-phase LR only; all larger gates remain blocked.

state/anima_303m_r4_low_lr_sft_2026_08_13/ preregisters that one arm. It reuses the exact language engine and changes only turn peak LR from 1e-3 to 1e-4; endpoint, fresh optimizer, data, objective, alignment, seed, decoder and all independent gates remain fixed. Broad retention uses the natural uniform-CE ceiling rather than a result-tuned tolerance.

The low-LR arm retained broad CE at 3.0926 but under-adapted: training teacher top-1 was 0.6031, exact/target 0/8, and independent semantics 0/7. Its verdict is FAIL-LOW-LR-TURN-SFT. Therefore LR reduction alone only trades forgetting for insufficient dialogue learning. The next single conceptual axis is native joint broad replay plus dialogue supervision in the existing multi-cell trainer; no new engine or evaluator is introduced.

state/anima_303m_r4_joint_replay_2026_08_13/ preregisters this one arm. Native two-cell round-robin provides four broad and four dialogue rows per step; additive CE applies full language loss everywhere and response supervision only where a canonical assistant span exists. The 3,750 step endpoint preserves the prior 15,000 dialogue-row exposure while adding 15,000 broad rows.

The joint arm retained broad CE 2.0620 and fully learned the dialogue probe (8/8), while independent output became structurally complete (7/7) but stayed semantically wrong (0/7), with heldout dialogue CE 5.0046. Its verdict is FAIL-JOINT-MEANINGFUL-CONVERSATION. This closes the bounded optimizer/sampler tradeoff but not generalization from 100 dialogues. A provenance-only repeated --cell-label argv bug mislabeled raw telemetry; file identity, sampling and loss were unaffected, and the harness now uses one canonical --cell-label broad dialogue argument.

All 121 local R4 micro model/evidence artifacts (521,291,120 bytes) are preserved and independently SHA-verified at private HF revision dancinlab/anima-303m-r4-aligned-micro-2026-08-13@6d2d4752cb222ba09fd74cb08eb8d3b7d4b140dc. Custody evidence is in state/anima_303m_r4_aligned_micro_custody_2026_08_13/result.json.

Preregistered R4 dialogue-support scale ladder

The joint arm closed the bounded optimizer/sampler explanation but did not generalize from 100 dialogues. The next single axis is pinned in state/anima_303m_r4_dialogue_scale_2026_08_13. The completed 100-document arm is reused as the frozen control; the same 0.89M ByteGPT, initial language checkpoint, 15,000 dialogue-row exposure, broad replay, optimizer, seed, canonical decode and conversation panel are run on nested 500, 1,500 and 3,500-document views. The 3,500-document arm is the registered primary endpoint, preventing post-result selection of an intermediate scale.

The protocol also pins dancinlab/anima-research@03d55ef as an interpretation constraint: mouth fluency is not consciousness evidence, a functional pass is non-disproof rather than proof, and later developmental gates remain disabled. No scale result directly authorizes 303M, IIT-mouth coupling, participant mounting or production.

The ladder is complete and failed. Held-out assistant CE improved monotonically from the frozen 100-document control 5.00458 to 2.36451, 1.82383 and 1.75553, while all three new arms remained at semantic 0/7 and failed memory/correction. The primary 3,500-document endpoint was structural 0/7 and emitted store/start repetition. Thus unique support at fixed 15,000-row compute improves teacher-forced prediction but does not create meaningful free conversation; it does not distinguish optimization exposure from capacity. Raw models and evidence are SHA-verified only in private HF revision dancinlab/anima-303m-r4-dialogue-scale-2026-08-13@1146240912244c7127b442196e2047a6f7641eac. The next permitted axis is a separately preregistered fixed-3,500-document exposure test.

R4 fixed-3,500 optimization-exposure ladder

state/anima_303m_r4_exposure_ladder_2026_08_13 freezes the next single axis before execution. The 3,500 documents, 0.89M ByteGPT, initial language checkpoint, broad replay, optimizer, sampler, objective, seed, canonical generator and conversation panel remain unchanged. One deterministic CPU trajectory runs to 30,000 steps with the original cosine schedule reaching its registered floor at step 3,750; checkpoints at 3,750/7,500/15,000/ 30,000 represent 15k/30k/60k/120k dialogue-row exposures. The first point must reproduce the prior control, all points are evaluated, and 120k is the fixed primary endpoint. A continued semantic 0/7 endpoint despite teacher-forced improvement permits only a separately preregistered capacity ladder; it does not authorize 303M, IIT-mouth coupling, participant mounting or production.

The ladder is complete and its control reproduced exactly. At 15k/30k/60k/120k dialogue rows, held-out assistant CE was 1.75553/1.71562/1.69534/1.69976, but semantic conversation stayed 0/7 at every point; the final structural score was 1/7, and memory/correction continued to fail. The verdict is FAIL-FIXED-CAPACITY-AFTER-EXPOSURE: eight times the registered exposure did not make the fixed 0.89M mouth meaningful. Raw checkpoints and evidence are SHA-verified only in private HF revision dancinlab/anima-303m-r4-exposure-ladder-2026-08-13@c30189456da40a80b23092651367a3eeacd0edf0. The next permitted axis is a separately preregistered fixed-data, fixed-exposure capacity ladder.

Preregistered R4 fixed-data capacity ladder

state/anima_303m_r4_capacity_ladder_2026_08_13 freezes the next single axis before execution. The broad/dialogue revisions and exact byte views, 3,500 documents, 120k dialogue rows, 120k replay rows, two-phase objective, optimizer, seed, batch, canonical generator and fail-closed panel remain fixed. The frozen 0.89M endpoint is compared with new exact 2.817M, 10.110M and 29.316M ByteGPT arms; their registered shapes preserve a native 64-dimensional attention head. Every larger arm rebuilds the same 2,000-step broad-language phase from scratch before the common 30,000-step joint phase because differently shaped checkpoints cannot be warm-started safely. All arms must run, with 29.316M as the primary endpoint.

Local work is limited to protocol and smoke validation; the result-bearing run may use one non-H100 Vast.ai GPU to protect the mini, and that instance must be destroyed afterward. A pass is only a meaningful-mouth gate requiring manual review, never a consciousness claim or direct authorization for 303M, IIT-mouth coupling, participant mounting or production.

The ladder completed with verdict FAIL-CAPACITY-LADDER. At exact 2.817M/10.110M/29.316M capacity, independent semantics stayed 0/7, structure was 7/7, and memory/correction failed. Training teacher top-1 improved 0.82570 β†’ 0.90712 β†’ 0.96692, but held-out assistant CE worsened 2.25676 β†’ 3.02574 β†’ 3.55896: larger fixed-exposure arms memorized the training support more strongly without meaningful generalization. Raw evidence is independently SHA-verified only in private HF revision dancinlab/anima-303m-r4-capacity-ladder-2026-08-13@3c9bc8cad1ac50c7610f1f6ab57bf09c82aa51ac. The non-H100 Vast.ai run cost an estimated $0.4538; its two protocol-owned instances were destroyed. The next axis needs a new data/compute-scaling preregistration, while 303M, IIT-mouth, participant and production remain blocked.

Preregistered R4 complete-trajectory support admission

state/anima_303m_r4_support_admission_2026_08_13 records the exhausted follow-up design space and freezes the next single-axis experiment. A live audit of the immutable dialogue source found 8,635 complete documents and 1,194 multi-turn trajectories, but the shared scale/exposure/capacity admission helper required exactly one user β†’ assistant pair. The actual 3,500-document capacity view therefore contained zero multi-turn examples even though the canonical trainer supports every assistant span. Memory and correction failures from that view cannot be interpreted as capacity evidence; its single-turn semantic 0/7 result remains unchanged.

The new fixed 2.817M ladder changes only admission coverage: the exact prior 3,500-document control, all 4,625 complete trajectories with a short final response, then all 8,635 complete trajectories. The exact language checkpoint, 120k dialogue and replay rows, optimizer, objective, seed, context, generator, panel and bars stay fixed, and all arms must run with the full-support arm as the primary endpoint. core.generator now owns canonical complete-trajectory parsing beside its existing renderer so experiment admission cannot silently redefine a valid chat document. H100 and 303M training are forbidden; failure permits only a separately preregistered broad-language data/compute axis, while any automatic pass still requires manual review and replication.

Open gap audit β€” 303M meaningful conversation

The 2026-08-12 read-only /gap audit below is the complete follow-up register for the current Python-only R0. It records 31 findings across eight lens families. It does not retroactively change the frozen panel, dataset revision, thresholds, failed checkpoint, or FAIL-MEANINGLESS-REPETITION verdict. Diagnostic work on preserved checkpoints must not be used for post-hoc checkpoint selection. Prior claims described as causes below are hypotheses unless a single-variable test has established them.

Priority means: P0 blocks a valid next R0 or a production-closed path, P1 blocks strong evidence or reproducibility, and P2 is required operational evidence but does not explain the current semantic failure.

Recovery overlay (2026-08-12): M1, A1, A2, A3, A6, R3, closed-loop .bin admission, canonical-SSOT, duplicated evaluator decode and the executable cross-tool contract are fixed in the shared Python engine and covered by tiny real-checkpoint regressions. M4 still needs the preserved full 303M checkpoint comparison before release. M2/R2 remain blocked on a new acceptable Korean multi-turn source and immutable HF revision. The numbered register below is retained as the original audit evidence; this overlay is its current disposition.

Math-structural gaps

  1. M1 Β· functor Β· P0 β€” chat framing does not commute across the pipeline. Training and the conversation panel use user: ...\nassistant:, while anima-py chat has a separate Korean μ‚¬μš©μž: ... | λ„μš°λ―Έ: framing and a different generation budget. The next protocol must put template, separator, stop rules, and byte budget in one chat-format SSOT and add an exact builder β†’ trainer β†’ evaluator β†’ runtime identity test. Evidence: conversation_panel.json, cli/chat.py, core/generator.py.
  2. M2 Β· operadic Β· P0 β€” the training support is not closed under the evaluated turn composition. The gate requires memory and correction across turns, but the current Korean builder renders one user β†’ assistant pair per document. A new, separately preregistered HF revision must preserve real Korean multi-turn trajectories and document/turn alignment; the frozen failed revision is not rewritten. Evidence: build_dataset.py, conversation_panel.json.
  3. M3 Β· persistent-homology / tropical Β· P1 β€” repetition-attractor birth and lifetime are unknown. Checkpoints exist every 2,000 steps, but meaningful conversation was measured only at the final checkpoint and no per-step top-1/top-2 margin or entropy was retained. A non-verdict diagnostic may record checkpoint Γ— prefix-length repetition lifetime and logit margin, without selecting the best historical checkpoint after observing the result. Evidence: protocol.json, train.log.
  4. M4 Β· bisimulation Β· P0 β€” the three real 303M decode paths lack byte-level equivalence evidence. The serialized ByteGPT checkpoint in Torch/engine form, evaluator-resident _Mouth, and ranged canonical generator have not been compared at identical seed bytes for step logits and generated bytes. Add an actual-checkpoint bisimulation contract test using the frozen panel seed. Evidence: cli/evaluate.py, core/generator.py, core/decode.py.

Adversarial-stress gaps

  1. A1 Β· adversarial semantics Β· P0 β€” the automatic semantic scorer has demonstrated false positives. The current code passes both the contradiction β€œIce does not melt ...” and the Korean substring answer μžλ™μ°¨μž…λ‹ˆλ‹€ for the required term μ°¨. Add preregistered negation, contradiction, keyword-salad, and Korean substring controls, with a morphology-independent canonical boundary rule. Evidence: cli/evaluate.py, conversation_panel.json.
  2. A2 Β· Byzantine input Β· P1 β€” panel identity is recorded but not enforced. The protocol pins a panel SHA-256, while --conversation-panel accepts any schema-compatible file and merely reports its hash. The evaluator must receive the expected protocol hash and fail closed before loading a substituted panel. Evidence: protocol.json, cli/evaluate.py.
  3. A3 Β· edge-chaos role boundaries Β· P1 β€” stop parsing recognizes only exact marker strings. Variants such as \n user:, \nUSER:, and \nμ‚¬μš©μž : may leak a fabricated next turn; current regression covers only a canonical lowercase marker. Replace substring matching with a line-start role parser and test whitespace, case, colon, English, and Korean variants. Evidence: core/generator.py, tests/test_conversation_gate.py.
  4. A4 Β· edge-chaos context rollover Β· P1 β€” long multi-turn seeds silently lose their oldest bytes. ByteGPT has a 512-byte block; the final Korean correction seed is already 420 bytes, so generation can evict its earliest fact. Add 511/512/513-byte boundary tests and record the visible context range at every generated step. Evidence: conversation_result.json, core/decode.py.
  5. A5 Β· perturbation / contamination Β· P1 β€” β€œzero contamination” covers exact containment, not semantic near-duplicates. The report-only audit examines the lexicographically first 100,000 of 649,354 retained documents; paraphrase, spacing, and back-translation leakage remain unmeasured. Run a panel-centered approximate search over the complete corpus as a separate sensitivity report. Do not delete post-hoc examples from the frozen revision. Evidence: build_dataset.py, result.json.
  6. A6 Β· response-supervision ablation Β· P0 β€” β€œanswer CE active” does not prove prompt-conditioned supervision. Telemetry counts assistant markers/positions but does not require the matching user prompt to remain visible in the same random window. Record fully framed, marker-only, and payload-only windows per cell, then preregister a treatment that preserves complete promptβ†’response spans. Evidence: cli/train.py, result.json.

Economic-resource gaps

  1. R1 Β· Pareto attribution Β· P1 β€” the proportional recovery changed multiple axes. Sampler, turn newline preservation, and Korean corpus changed together, so the validation improvement cannot be assigned to the sampler alone. Downgrade the existing root-cause wording to correlational evidence and preregister matched sampler-only and data/framing-only ablations. Evidence: README.
  2. R2 Β· information budget / optimal transport Β· P0 β€” exposure follows file size, not required capability coverage. A 303,097,856-parameter model received 229,376,000 target bytes and only 11,025,460 response-supervised positions. The proportional run exposed about 2.97% English dialogue, 16.97% Korean dialogue, and zero Korean multi-turn mass. The next protocol must pin a language Γ— single/multi-turn Γ— memory/correction capability distribution and report effective framed bytes per parameter plus coverage distance. Evidence: result.json.
  3. R3 Β· dynamic-programming provenance Β· P1 β€” intermediate ByteGPT metadata is wrong. _write_bin writes the final configured steps and the latest training-batch loss into every intermediate .bin; the step-2,000 log therefore says step=14000. Pass the actual completed step and the latest measured validation CE into the writer and add a provenance regression. The final R0 failure remains valid, but checkpoint-time analyses are not yet trustworthy. Evidence: cli/train.py, train.log.
  4. R4 Β· Landauer accounting Β· P2 β€” energy cost is absent. GPU time, VRAM, and dollars are recorded, but power and cumulative energy are not. The next Vast.ai run should collect non-interfering NVML power telemetry and report joules per target byte and per effective assistant byte. Evidence: result.json, vram.csv.

Epistemic-evidence gaps

  1. E1 Β· assumption surfacing Β· P1 β€” observations, hypotheses, and confirmed causes are mixed. Undertraining, random-window framing loss, and single-turn Korean data are listed together as remaining causes. Every candidate must carry an evidence level, falsifier, and smallest single-variable experiment. Evidence: result.json.
  2. E2 Β· Bayesian reproducibility Β· P1 β€” the latest treatments each have only seed 7. They honestly falsify only their fixed recipes; they do not estimate R0 pass probability or seed variance. Require a preregistered multi-seed posterior and minimum success streak only after a single-seed screen passes. Evidence: protocol.json.
  3. E3 Β· counterfactual falsifier Β· P1 β€” the full panel/decoder instrument lacks model controls. Canned scorer strings are not an end-to-end positive/negative calibration. Run the same frozen decode path against one known-good conversation checkpoint and one known-bad checkpoint, and keep instrument discrimination separate from the current model verdict.
  4. E4 Β· honesty triad Β· P1 β€” manual-review artifacts disagree. Raw conversation_result.json says manual review is REQUIRED, while the summary claims completed 0/14 without immutable per-item decisions, reviewer identity, blindness, or criteria. Preserve a separate signed/hashed review artifact for every raw response before making a manual-review claim. Evidence: conversation_result.json, result.json.

Convergence-closure gaps

  1. C1 Β· fixpoint / success criteria Β· P1 β€” there is no active post-failure diagnostic protocol. The response-CE protocol is completed, but the next micro-experiment sequence has no frozen hypothesis, success/stop rule, maximum count, or candidate-disposal table. Register that before any result-bearing experiment. Evidence: README.md, protocol.json.
  2. C2 Β· regression streak Β· P1 β€” code QA is not model-behavior evidence. 77 passed describes software tests; the latest actual checkpoint streak is 0/1, with no seed or hardware repeat. Keep code QA and semantic-model success streaks as separate promotion fields. Evidence: result.json.
  3. C3 Β· closed loop Β· P0 β€” a passing 303M .bin still cannot enter the participant. The participant exposes lora|v3|akida|clm, and CLMSubstrate accepts only .clm, although the shared generator already dispatches .bin/.clm. Extend the existing participant substrate boundary to reuse core.generator rather than add a new engine. Evidence: anima_participant.py, substrate_clm.py, core/generator.py.

Simplicity-canonical gaps

  1. S1 Β· canonical SSOT Β· P0 β€” chat format and stop markers are duplicated. The panel, dataset builder, trainer flags, generator, and chat CLI each own literals without fail-closed equality validation. Put them in one minimal chat-format manifest consumed by all existing paths; do not add another evaluator or runtime. Evidence: conversation_panel.json, build_dataset.py, cli/train.py, core/generator.py.
  2. S2 Β· duplicated helper Β· P0 β€” evaluator _Mouth.chat reimplements the low-level dispatch. It should call a preloaded canonical backend interface from core.generator; require actual checkpoint parity before removing the duplicate. Evidence: cli/evaluate.py, core/generator.py.
  3. S3 Β· architectural legibility Β· P2 β€” README mixes active and retired R0 recipes. KLUE, proportional, and response-CE records coexist under β€œCurrent experiment,” and β€œ303M R0 evaluator invalid” does not identify which historical evaluator failed. After this register, retain one explicit active-protocol pointer and list completed protocols as historical evidence.

Temporal-dynamics gaps

  1. T1 Β· temporal hierarchy Β· P1 β€” validation CE and semantic behavior are sampled at different timescales. CE runs every 200 steps but conversation/repetition only at the final step. Replay preserved checkpoints chronologically for diagnosis, never for post-hoc best-checkpoint promotion.
  2. T2 Β· temporal decay Β· P1 β€” memory is tested only at the immediately following turn. After R0 first passes, add a separately frozen 1/2/4-turn delay and context-rollover memory panel with irrelevant intervening turns. Evidence: conversation_panel.json.
  3. T3 Β· heuristic promotion / introduced axes Β· P1 β€” hypotheses have been promoted after multi-axis treatments. Enforce a micro β†’ single-seed β†’ multi-seed ladder in which each treatment changes one shared-flow variable and predeclares which candidate it falsifies.
  4. T4 Β· active acquisition Β· P0 β€” the missing Korean memory/correction support is already known. Build provenance-bearing real Korean multi-turn and correction trajectories, isolated from panel wording, in a new immutable HF dancinlab revision. The current frozen data decision means this requires a new protocol, not an in-place edit. Evidence: build_dataset.py.

Coverage-consistency gaps

  1. V1 Β· axis coverage Β· P0 β€” scorer controls do not cover every blocking bar and language. The four controls contain only one English positive. Add English/Korean positive and negative controls for memory final, correction final, contradiction, keyword salad, UTF-8 boundaries, completion, role leakage, and substring collisions. Evidence: conversation_panel.json, tests/test_conversation_gate.py.
  2. V2 Β· cross-tool consistency Β· P0 β€” builder, trainer, evaluator, anima-py chat, and participant do not share an enforced release contract. For one real checkpoint, compare seed bytes, each step's logits, stop decision, and final raw bytes across all tools under the same template, maximum bytes, load strategy, and parser.
  3. V3 Β· unowned load-bearing gate / landscape Β· P1 β€” manual review and production wiring have no explicit artifact owner. FIFO, reply ownership, concurrent users, HTTP/WebSocket, soak, rollback, and participant state remain intentionally unrun while R0 fails. The next protocol must name the review artifact/schema and connect a passing conversation R0 to these staging gates without skipping them.

Blocking order and immediate decision

The audit's three highest-impact blockers are:

  1. Invalid semantic discrimination: contradictions and Korean substring collisions can pass.
  2. Capability-support mismatch: random byte windows can lose prompts, and Korean multi-turn, memory, and correction training mass is absent.
  3. Missing canonical closed loop: evaluation bypasses the shared generator interface and a ByteGPT .bin cannot be selected by the production participant.

The next result-bearing work is therefore blocked until a new Python-only diagnostic/R0 protocol freezes: (1) the canonical chat-format SSOT and cross-tool contract, (2) adversarial scorer controls and fail-closed panel identity, (3) provenance-bearing bilingual multi-turn capability coverage, and (4) single-variable stop/falsifier rules. R1 recurrent workspace and production deployment remain locked. Models and training data remain private under HF dancinlab; GPU work remains on Vast.ai; user-owned ING.jsonl and stream_mi.json remain untouched.

Canonical entry

python3 -m venv .venv
.venv/bin/python -m pip install -e ".[train,runtime]"

.venv/bin/anima-py --help
.venv/bin/anima-py train --help
.venv/bin/anima-py evaluate --help
.venv/bin/anima-py chat MODEL.clm

Main commands:

Command Responsibility
anima-py corpus Build registered training corpora.
anima-py train Train through the shared PyTorch engine and serialize checkpoints.
anima-py evaluate Run registered NumPy/runtime measurements and causal controls.
anima-py serialize Export existing training checkpoints to runtime formats.
anima-py sweep Run bounded multi-device experiment matrices.
anima-py chat Run the A⇄G consciousness daemon and byte mouth.
anima-py study Run registered interaction studies.

Research instrumentation that was previously unavailable on the Python path is now part of the same chat engine:

anima-py chat MODEL.clm --opgrip
anima-py chat MODEL.clm --opgrip-live
anima-py chat MODEL.clm --opgrip-r3
anima-py chat MODEL.clm --refractory

The decode-free --opgrip arm can run without a checkpoint. Live and R3 arms fail closed unless the checkpoint loads successfully.

Runtime architecture

anima-py
└── cli/anima.py
    β”œβ”€β”€ cli/train.py ───────► core/model.py ─────► core/serialize.py
    β”œβ”€β”€ cli/evaluate.py ────► core/decode.py
    └── cli/chat.py
        β”œβ”€β”€ core/brain.py
        β”œβ”€β”€ core/pure_field.py       Engine A
        β”œβ”€β”€ core/engine_g.py         Engine G, motivation, emission, refractory
        β”œβ”€β”€ core/generator.py ──────► core/decode.py
        β”œβ”€β”€ core/kosmos_io.py
        └── core/dream_*.py

Runtime rules:

  • Extend the shared engine instead of adding side harnesses that redo its computation.
  • Keep registered data, randomness, criteria, and controls immutable during a measured run.
  • Fail closed on missing checkpoints, malformed inputs, incompatible checkpoint structure, or missing pinned evaluation assets.
  • Keep raw model bytes lossless through UTF-8/surrogateescape and structured JSON output.

Verification

Local regression:

.venv/bin/python -m compileall -q cli core anima_py
.venv/bin/python -m pytest -q tests cli/test_train_import_resolution.py agent/domains/CHAT/test_*.py
.venv/bin/anima-py --help
.venv/bin/anima-py evaluate --help
actionlint .github/workflows/*.yml

Heavy model and serving QA runs on Vast.ai. Models and training datasets are stored only in private repositories under the Hugging Face dancinlab organization. Secrets are supplied by the deployment environment or secret CLI and are never committed.

Latest production evidence

  • The 7B store-causality run passed its registered causal, HTTP/WebSocket, soak, recovery, and rollback gates after a shared decoder throughput fix. Evidence: state/store_causality_7b_throughput_recovery_2026_08_11/result.json.
  • Live-user QA then invalidated that checkpoint as a semantic chat deployment. The broker and participant reply-ownership, prior-emission comparison, language ownership, and cooldown flow were corrected. Evidence: state/chat_7b_conversation_recovery_2026_08_11/result.json.
  • The 303M R0 evaluator was later classified as an invalid measurement; R1 remains locked. Evidence: state/anima_303m_r0_local_micro_2026_08_12/result.json.

No model result is promoted solely because transport health passes. Semantic chat, causal controls, throughput, soak, recovery, and rollback are separate blocking gates.

Repository boundaries

  • dancinlab/anima is the only active source repository.
  • cli/, core/, and anima_py/ own active runtime code.
  • state/ owns registered protocols and result evidence.
  • archive/ is non-runtime provenance.
  • Vast.ai owns pod execution; Hugging Face dancinlab owns model and dataset custody.

License

MIT. See LICENSE.

About

🧠 Living Consciousness Agent β€” PureField repulsion-field engine Β· Engine A ⇄ Engine G Β· Ξ¨=1/2 fixed point Β· 2,448 laws + 392 hypotheses

Topics

Resources

Stars

142 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages