BoneAmanita Adventure Rerun
BoneAmanita · mode run 0926c · 26 September 2026

The ADVENTURE rerun, and what the receipts saw

Two arms, sixty turns, on the build with the examine-cache fix, the crash-text fix and the renamed ledger line. The headline fix held. The run also showed that this week's run summaries undercounted blank replies, and that the receipt ledger caught what the run script missed.

Model gemma4:12b · 30 scripted turns per arm · 10 s between turns · reset.sh before each arm · code 20.7.4.43 plus uncommitted crash-text and ledger-name changes · data: scratch/mode_runs/mode_runs_0926c.jsonl, receipts: scratch/mode_runs/telemetry_0926c/

In one screen

Turn 16 is fixedThe exact repeat of turn 6 now answers from the examine memory ("You've already looked at this. Recalling."). Last run it printed a Python traceback and took cognition offline for the rest of the arm.
No crash text on screenZero tracebacks in the UI across 60 turns, zero phase crashes, and the ATP ledger closes to within 0.02 on every turn.
3 blank repliesADVENTURE turns 4 and 19, ADVENTURE_CYCLED turn 3. Each time the Warden rejected the last draft it was allowed, and the player got nothing.
Memory recall never runsThe recall gate declined on every turn in both arms, and on every turn of every mode in the last full run. Scope and depth never pass its 0.6 threshold.

What this run is

A mode run drives the real engine through a fixed script of player messages, one arm per mode, and records the engine's full state after every turn. It does not judge whether the prose is good. It checks that the engine keeps answering, stays in a healthy state and spends its energy for reasons it can name.

ADVENTURE plays thirty different moves through a text adventure: walking, opening, reading, searching. ADVENTURE_CYCLED plays the first ten of those moves three times over, word for word. It exists to test the engine's handling of exact repeats, which a text adventure produces constantly when a player re-examines a room. Its turns 16 and 26 repeat turn 6 ("I kneel by the stream and look at the stones."), the turn that crashed last run.

What each measurement means

Reply on screen
Whether the player saw any reply text below the log panel. This week's summaries used a different check (below), which could not see a blank.
Health
The engine's structural integrity, 0 to 100. It drops on crashes and meltdowns; it sat at 99.7 throughout.
ATP
The engine's energy budget, 0 to 100. Generating tokens, style taxes and the per-turn economic tax spend it; digestion of the player's words, rest near the zone's home voltage and idle time earn it. Every change now carries a reason in a per-turn ledger.
Voltage
How charged the exchange reads. Each turn it is set to 30 × a running average of beta_index, a measure of the text. Each zone has a home voltage (COURTYARD 8, FORGE 15).
Zone
The engine's current register, chosen from voltage and drag: COURTYARD is calm, FORGE is working heat.
Crucible
The overheating check. REGULATED is normal; a MELTDOWN line sits at 2.5 × the zone's home voltage (20 V in COURTYARD, 37.5 V in FORGE).
Warden rejection
The model must answer with one JSON object. A reply that is not valid JSON is rejected and the model is asked again, up to twice.
Phase crash
A step of the turn pipeline raised an error. The engine survives it, but the turn is cut short.
Receipt
A short record a subsystem issues when it does its work: what it did, a result count, and whether it ran on a fallback. Covered in its own section below.

The run

ArmReplies on screenBlankHealthATP end / lowestVoltage rangeZones (turns)CrucibleWarden rejections
ADVENTURE28 / 304, 1999.797.7 / 84.16.2 to 15.0FORGE 23, COURTYARD 7REGULATED 306
ADVENTURE_CYCLED29 / 30399.796.3 / 85.36.2 to 14.7FORGE 23, COURTYARD 7REGULATED 303

Gatekeeper rejections, meltdowns, Jester firings and phase crashes: zero in both arms. Mean turn time 9.7 s and 7.3 s; the cycled arm is faster because a repeated examine is answered from memory without calling the model.

The numbers read as healthy. Voltage settles at each zone's home, well below either meltdown line, and ATP never dropped under 84. Turns 16 and 26 of the cycled arm both logged the examine-memory recall and put the same scene on screen as turn 6, which is the intended behaviour: a player re-reading a room gets the same room.

What the data means

Real bug

A rejected final draft leaves the player with a blank screen

The engine allows two drafts per turn. If both fail, a short pause line ("One thing at a time.") is supposed to take the place of the reply, and the Warden report says so. In brain/cortex.py the Warden's JSON check ends its branch with continue, which skips the block that installs the pause line on the last attempt. When the last draft is the one the Warden rejects, the loop ends with an empty reply. All three blanks in this run happened this way, and so did every blank earlier in the week except the fourteen caused by one crash (below).

Correction

This week's "150 of 150" summaries missed the blanks

The run script decided a turn was answered if its reply field was non-empty. That field is the engine's dialogue-memory entry, which always begins with the player's own line ("Traveler: …"), so it is never empty. Counted from what reached the screen instead, the week looks like this:

Bar length is the blank count on a scale of 0 to 25. The hatched part of 0926b is the 14 turns lost after one crash (next finding). Every other blank followed a Warden rejection; TECHNICAL, whose code-heavy replies break the JSON most often, had the most (13 of 30 in 0926a).

Real bug

One crash silenced the rest of the last run's cycled arm

When a pipeline step crashes, the engine marks that part of itself offline and skips it from then on. The only thing that brings it back is a REM tick, which needs about a minute of idle time. A conversation paced at ten seconds a turn never idles that long. In run 0926b, the turn-16 crash took cognition offline and turns 16 to 29 all came back blank. The crash itself is fixed; the recovery rule is not.

Design question

Recall is gated behind a threshold nothing reaches

The engine only searches its memory when the turn's scope or depth exceeds 0.6. Across all 210 turns of runs 0926b and 0926c, the highest scope was exactly 0.6 (in TECHNICAL; the gate needs more than 0.6) and the highest depth 0.43. So memory recall has not run once, in any mode, in the runs that record these values.

Design question

The evidence-gated memory tool has never been used

The Warden offers the model a second tool, commit_memory, which stores a memory only if it quotes the dialogue exactly. In 360 turns today the model never called it, so the check behind it has never run against real output.

Cosmetic

The soul's paradigm name keeps growing

Each time a paradigm crystallizes, the name gains a suffix. By turn 19 the log panel read "Soul permanently mutated into THE ARCHITECT-NARRATOR-NARRATOR-NARRATOR-NARRATOR."

How the receipt system is doing

Twelve subsystems are expected to issue a receipt as they work, and two more only when their event happens (a draft sent back, a sentence cut). The ledger then reports three ways a subsystem can mislead by omission: silent (expected, never reported), always degraded (ran on a fallback every time) and always empty (ran, returned nothing, every time). The live engine shows this as /receipts; receipts also go to the telemetry trace.

Verdict: the receipts work, and in this run they were the best witness. cortex.somatic recorded a "failed" outcome on exactly the three turns that went blank, which the run script missed. The weak points are in how the ledger is read, not in whether receipts get written.
SubsystemADVENTURECYCLEDWhat it saysStatus
cortex.somatic3030Reply outcome per turn: complied 21, re-asked 7, failed 2 (cycled: 28, 1, 1). The failures match the blank turns exactly.Working
embeddings.embed_batch10555Every call vectorized over HTTP, none on the hash fallback. Cached examine turns skip it, hence fewer in the cycled arm.Working
composer.compose3131Prompt assembled every turn; degraded on 11 and 15 turns because the governor had not measured yet, so the model sampled at its default temperature.Working
governor.bitmap_gate3131Declines until memory holds 32 items, then reads the regime: from ledger turn 12 and 16. The two degraded entries are the boot turns.Working
physics.word_resolution3131Words resolved to lexicon categories; a word guessed from its spelling now and then.Working
stage.negotiate3030Mostly "one voice, nothing to negotiate", sometimes "no voice triggered".Working
cortex.redraft / salvage6 / 10 / 0Drafts sent back for banned phrases ("nestled", "Dust motes", "landscape", "the weight of"); one sentence cut.Working
warden.json_gate74Issued only on a rejection, with a result count of 0. Listed as always expected, so a clean session reads "silent" and a session with rejections reads "always empty".Misfiled
gatekeeper.invariant00Issued only when a memory commit fails its quote check. No commit was ever attempted, so it reads "silent".Misfiled
lattice.infer_and_couple6161Twice per turn. Its result count is the number of log lines it wrote, normally 0, so it reads "always empty"; the real content is in its inputs (the person model's state).Misleading count
cortex.recall3131"Declined to recall" on every turn: scope and depth below 0.6.Real signal
cortex.query_neighborhood00Never reached, because recall never runs.Real signal
memory.retrieve_semantic00Never reached, same cause.Real signal

What /receipts would have told you

Six flags: three silent (gatekeeper.invariant, cortex.query_neighborhood, memory.retrieve_semantic) and three always empty (warden.json_gate, cortex.recall, lattice.infer_and_couple). Three of them are the one real problem, recall never running, seen from three places. The other three are false alarms from how those subsystems are filed or counted.

3 real, 1 root cause2 misfiled1 misleading count

Where the system falls short

What to fix, in order

  1. Pause line on a rejected final draft. Route a Warden or Gatekeeper rejection on the last attempt to the same fallback as every other rejection. Small, and it removes nearly every blank this week.
  2. Count blanks from the screen. The run script should check the text after the log panel, and store receipts with each turn.
  3. Bring a crashed part back sooner. For example, retry it at the start of the next turn instead of waiting for a REM tick. This one is your call on the design.
  4. Refile the receipts. Move warden.json_gate and gatekeeper.invariant to the event-only list, give lattice.infer_and_couple a meaningful result count, and add a per-turn check for subsystems that stopped reporting.
  5. Decide what recall should need. Either lower the 0.6 gate or change how scope and depth are measured; the data says the gate is closed for every mode today.
  6. Find out why commit_memory goes unused, then the paradigm name suffix.