The ADVENTURE rerun, and what the receipts saw
Two arms, sixty turns, on the build with the examine-cache fix, the crash-text fix and the renamed ledger line. The headline fix held. The run also showed that this week's run summaries undercounted blank replies, and that the receipt ledger caught what the run script missed.
In one screen
What this run is
A mode run drives the real engine through a fixed script of player messages, one arm per mode, and records the engine's full state after every turn. It does not judge whether the prose is good. It checks that the engine keeps answering, stays in a healthy state and spends its energy for reasons it can name.
ADVENTURE plays thirty different moves through a text adventure: walking, opening, reading, searching. ADVENTURE_CYCLED plays the first ten of those moves three times over, word for word. It exists to test the engine's handling of exact repeats, which a text adventure produces constantly when a player re-examines a room. Its turns 16 and 26 repeat turn 6 ("I kneel by the stream and look at the stones."), the turn that crashed last run.
What each measurement means
- Reply on screen
- Whether the player saw any reply text below the log panel. This week's summaries used a different check (below), which could not see a blank.
- Health
- The engine's structural integrity, 0 to 100. It drops on crashes and meltdowns; it sat at 99.7 throughout.
- ATP
- The engine's energy budget, 0 to 100. Generating tokens, style taxes and the per-turn economic tax spend it; digestion of the player's words, rest near the zone's home voltage and idle time earn it. Every change now carries a reason in a per-turn ledger.
- Voltage
- How charged the exchange reads. Each turn it is set to 30 × a running average of
beta_index, a measure of the text. Each zone has a home voltage (COURTYARD 8, FORGE 15). - Zone
- The engine's current register, chosen from voltage and drag: COURTYARD is calm, FORGE is working heat.
- Crucible
- The overheating check. REGULATED is normal; a MELTDOWN line sits at 2.5 × the zone's home voltage (20 V in COURTYARD, 37.5 V in FORGE).
- Warden rejection
- The model must answer with one JSON object. A reply that is not valid JSON is rejected and the model is asked again, up to twice.
- Phase crash
- A step of the turn pipeline raised an error. The engine survives it, but the turn is cut short.
- Receipt
- A short record a subsystem issues when it does its work: what it did, a result count, and whether it ran on a fallback. Covered in its own section below.
The run
| Arm | Replies on screen | Blank | Health | ATP end / lowest | Voltage range | Zones (turns) | Crucible | Warden rejections |
|---|---|---|---|---|---|---|---|---|
| ADVENTURE | 28 / 30 | 4, 19 | 99.7 | 97.7 / 84.1 | 6.2 to 15.0 | FORGE 23, COURTYARD 7 | REGULATED 30 | 6 |
| ADVENTURE_CYCLED | 29 / 30 | 3 | 99.7 | 96.3 / 85.3 | 6.2 to 14.7 | FORGE 23, COURTYARD 7 | REGULATED 30 | 3 |
Gatekeeper rejections, meltdowns, Jester firings and phase crashes: zero in both arms. Mean turn time 9.7 s and 7.3 s; the cycled arm is faster because a repeated examine is answered from memory without calling the model.
The numbers read as healthy. Voltage settles at each zone's home, well below either meltdown line, and ATP never dropped under 84. Turns 16 and 26 of the cycled arm both logged the examine-memory recall and put the same scene on screen as turn 6, which is the intended behaviour: a player re-reading a room gets the same room.
What the data means
A rejected final draft leaves the player with a blank screen
The engine allows two drafts per turn. If both fail, a short pause line ("One thing at a time.") is supposed to take the place of the reply, and the Warden report says so. In brain/cortex.py the Warden's JSON check ends its branch with continue, which skips the block that installs the pause line on the last attempt. When the last draft is the one the Warden rejects, the loop ends with an empty reply. All three blanks in this run happened this way, and so did every blank earlier in the week except the fourteen caused by one crash (below).
This week's "150 of 150" summaries missed the blanks
The run script decided a turn was answered if its reply field was non-empty. That field is the engine's dialogue-memory entry, which always begins with the player's own line ("Traveler: …"), so it is never empty. Counted from what reached the screen instead, the week looks like this:
Bar length is the blank count on a scale of 0 to 25. The hatched part of 0926b is the 14 turns lost after one crash (next finding). Every other blank followed a Warden rejection; TECHNICAL, whose code-heavy replies break the JSON most often, had the most (13 of 30 in 0926a).
One crash silenced the rest of the last run's cycled arm
When a pipeline step crashes, the engine marks that part of itself offline and skips it from then on. The only thing that brings it back is a REM tick, which needs about a minute of idle time. A conversation paced at ten seconds a turn never idles that long. In run 0926b, the turn-16 crash took cognition offline and turns 16 to 29 all came back blank. The crash itself is fixed; the recovery rule is not.
Recall is gated behind a threshold nothing reaches
The engine only searches its memory when the turn's scope or depth exceeds 0.6. Across all 210 turns of runs 0926b and 0926c, the highest scope was exactly 0.6 (in TECHNICAL; the gate needs more than 0.6) and the highest depth 0.43. So memory recall has not run once, in any mode, in the runs that record these values.
The evidence-gated memory tool has never been used
The Warden offers the model a second tool, commit_memory, which stores a memory only if it quotes the dialogue exactly. In 360 turns today the model never called it, so the check behind it has never run against real output.
The soul's paradigm name keeps growing
Each time a paradigm crystallizes, the name gains a suffix. By turn 19 the log panel read "Soul permanently mutated into THE ARCHITECT-NARRATOR-NARRATOR-NARRATOR-NARRATOR."
How the receipt system is doing
Twelve subsystems are expected to issue a receipt as they work, and two more only when their event happens (a draft sent back, a sentence cut). The ledger then reports three ways a subsystem can mislead by omission: silent (expected, never reported), always degraded (ran on a fallback every time) and always empty (ran, returned nothing, every time). The live engine shows this as /receipts; receipts also go to the telemetry trace.
cortex.somatic recorded a "failed" outcome on exactly the three turns that went blank, which the run script missed. The weak points are in how the ledger is read, not in whether receipts get written.
| Subsystem | ADVENTURE | CYCLED | What it says | Status |
|---|---|---|---|---|
| cortex.somatic | 30 | 30 | Reply outcome per turn: complied 21, re-asked 7, failed 2 (cycled: 28, 1, 1). The failures match the blank turns exactly. | Working |
| embeddings.embed_batch | 105 | 55 | Every call vectorized over HTTP, none on the hash fallback. Cached examine turns skip it, hence fewer in the cycled arm. | Working |
| composer.compose | 31 | 31 | Prompt assembled every turn; degraded on 11 and 15 turns because the governor had not measured yet, so the model sampled at its default temperature. | Working |
| governor.bitmap_gate | 31 | 31 | Declines until memory holds 32 items, then reads the regime: from ledger turn 12 and 16. The two degraded entries are the boot turns. | Working |
| physics.word_resolution | 31 | 31 | Words resolved to lexicon categories; a word guessed from its spelling now and then. | Working |
| stage.negotiate | 30 | 30 | Mostly "one voice, nothing to negotiate", sometimes "no voice triggered". | Working |
| cortex.redraft / salvage | 6 / 1 | 0 / 0 | Drafts sent back for banned phrases ("nestled", "Dust motes", "landscape", "the weight of"); one sentence cut. | Working |
| warden.json_gate | 7 | 4 | Issued only on a rejection, with a result count of 0. Listed as always expected, so a clean session reads "silent" and a session with rejections reads "always empty". | Misfiled |
| gatekeeper.invariant | 0 | 0 | Issued only when a memory commit fails its quote check. No commit was ever attempted, so it reads "silent". | Misfiled |
| lattice.infer_and_couple | 61 | 61 | Twice per turn. Its result count is the number of log lines it wrote, normally 0, so it reads "always empty"; the real content is in its inputs (the person model's state). | Misleading count |
| cortex.recall | 31 | 31 | "Declined to recall" on every turn: scope and depth below 0.6. | Real signal |
| cortex.query_neighborhood | 0 | 0 | Never reached, because recall never runs. | Real signal |
| memory.retrieve_semantic | 0 | 0 | Never reached, same cause. | Real signal |
What /receipts would have told you
Six flags: three silent (gatekeeper.invariant, cortex.query_neighborhood, memory.retrieve_semantic) and three always empty (warden.json_gate, cortex.recall, lattice.infer_and_couple). Three of them are the one real problem, recall never running, seen from three places. The other three are false alarms from how those subsystems are filed or counted.
Where the system falls short
- Receipts do not survive a run. They go to the telemetry trace, and
reset.shdeleteslogs/before every arm. This report exists because the traces were copied out during the run. The run script's per-turn rows do not include them. - The scorecard only looks at the whole session. A subsystem that reported on turn 1 counts as "seen" forever. In run 0926b cognition went offline for 14 turns and the ledger would not have flagged it, because the subsystems had already reported earlier.
- False alarms train people to ignore it. Half of the six flags here are noise from filing and counting, which makes the real flag easier to dismiss.
- The ledger keeps the last 512 receipts. This run issued about 12 a turn, so a long session judges "always" over roughly the last 40 turns only.
What to fix, in order
- Pause line on a rejected final draft. Route a Warden or Gatekeeper rejection on the last attempt to the same fallback as every other rejection. Small, and it removes nearly every blank this week.
- Count blanks from the screen. The run script should check the text after the log panel, and store receipts with each turn.
- Bring a crashed part back sooner. For example, retry it at the start of the next turn instead of waiting for a REM tick. This one is your call on the design.
- Refile the receipts. Move
warden.json_gateandgatekeeper.invariantto the event-only list, givelattice.infer_and_couplea meaningful result count, and add a per-turn check for subsystems that stopped reporting. - Decide what recall should need. Either lower the 0.6 gate or change how scope and depth are measured; the data says the gate is closed for every mode today.
- Find out why
commit_memorygoes unused, then the paradigm name suffix.