BoneAmanita Dev Diary: The Engine Remembers
September 27, 2026
Github Repo
For most of its life BoneAmanita has had a strange kind of memory. It
kept a great deal about itself: its energy, its moods, the words it had
learned, even the scars left by earlier versions of itself that had
"died". What it could not reliably keep was the simplest thing a person
expects from a conversation partner: what you told it.
This week that changed.
A Gate for Everything the Model Wants to Keep
The engine now runs every change to its long-term state through a small,
strict gate called Halcyon, grafted in from the bone-iris project. The
gate denies by default. The model may propose one change per reply, in a
fixed one-line format, and the gate checks it against a written
declaration of what is allowed before anything is saved. Every decision,
accepted or refused, is kept in an audit trail along with the words the
model wrote around it.
Before Halcyon, an earlier attempt had required every reply to be a JSON
object. On real runs that produced blank replies whenever the model
stumbled on the format. Halcyon lets the model speak in plain prose and
only asks for structure on the one line that proposes a change.
With the gate in place, I moved everything the engine learns into one
SQLite store, one piece at a time: the resume point, the adventure's
rooms, the words the lexicon teaches itself, the Akashic record, the lore
it learns on top of the factory files, a profile of the person, and the
lineage each death leaves the next generation. Each move imports the old
file once and keeps it, so nothing is lost on upgrade. Two of those
pieces turned out to have been quietly broken for months: the user
profile had never been switched on, and inherited scars had stopped
being applied at boot. Both work now.
The Model Never Wrote It Down
The obvious next step was to hand the model back what it had kept. So I
built a probe: tell the engine three facts, wipe the conversation, start
a new day, and ask for each fact.
The first run kept nothing. Asked for my sister's name, the model
explained that it had no memory of me. It had been offered the chance to
remember on every turn and had never once taken it. Rewording the
instruction, moving it, adding a worked example: none of it helped. Out
of 36 tries, the model wrote the line once.
So memory now has a keeper. After each reply, one short, separate call
asks a single question: did the person just say something worth knowing
later? It answers with a name and a value, or with NONE. The engine
writes the proposal, and the gate judges it exactly as it would judge
the model's own. With the keeper, all twelve facts in the probe were
kept across all four modes, and every question that got an answer was
answered from memory.
Remembering by Meaning, and Forgetting on Purpose
Memories now come back by what they mean rather than by the words they
share. "Where can I find something to see by?" finds the brass lamp by
the north door, and "What do I still have to pay back?" finds the
ferryman's debt, though neither question shares a word with its memory.
Memory also has a limit, so the engine now forgets on purpose. It counts
how often each memory is actually used. Near-duplicates are merged, and
as memory approaches its cap, the least-used memories go first. The
threshold for "the same memory in different words" was measured rather
than guessed: paraphrased duplicates scored between 0.90 and 0.97, while
two different people visiting next week scored 0.76.
A Map That Stays Drawn
In Adventure mode, rooms used to exist only in the text of the last
reply. Now every room the narrator describes is charted into a world
graph: its description, its exits, and the things in it. The next turn
reads the room back from the map instead of from the model's last
sentence. Objects that leave a room are retired rather than deleted, so
the history of the world is kept.
Two smaller Adventure fixes came out of the same work. The narrator used
to answer "Where did I hide the key?" with a fresh room description; it
now answers the question first. And rooms used to be titled with the
engine's internal zone names, labels like VOID_DRIFT. It is now a rule
of the project that a room title names the place itself. The zones
belong in the debugging panel.
Where the Project Stands
The memory plan is complete. There is a /memory command,
available in Technical mode and on the deep display, that lists what the
engine keeps, how often each memory came back, and the gate's latest
decisions. The embedding service can now recover from an outage instead
of falling back to meaningless vectors for the rest of the session. The
test suite stands at 880 passing tests, and each new feature's tests
were checked by deliberately breaking the feature and confirming the
test fails.
For the first time, telling BoneAmanita something is a promise it can
keep.
BoneAmanita Dev Diary: Every Mode, With a Real Model
September 26, 2026
Github Repo
Unit tests had been passing for weeks. The question I had been avoiding
was simpler: what happens when a real model plays thirty turns in each
mode?
Thirty Turns Per Mode
I wrote a harness that drives the engine through thirty paced turns in
Conversation, Adventure, Technical and Creative modes on a local model,
and records everything on every turn: health, energy, hormones, the
physics, every log line, and whether the person saw anything at all.
Conversation held up. Adventure did not. Only eight of its thirty turns
reached the model. Sixteen ended in a system halt because the engine ran
out of energy, and six more were silences. The engine's own economy was
starving the story.
Fixing What the Runs Found
The fixes came as a plan of nine items, built in order with a test for
each. A gentler "bunny hill" start for the first turns. A detector for
requests the engine genuinely cannot do, so it can say so once instead
of pretending. Repeated moves in Adventure no longer punish the player.
PINKER, the gate that decides when a reply is under too much strain,
was audited and given honest inputs. The Jester now answers real loops
instead of imaginary ones.
The runs also taught me to measure more. Cortisol, the energy ledger,
what digestion paid and at what voltage: if a number moves, it is now
recorded. More data is almost never a bad thing when the question is why
something went quiet.
Receipts, and Crashes That Stay Off the Screen
Since earlier in the month, every subsystem that is supposed to act on a
turn leaves a receipt saying what it did, or that it did nothing, or that
it fell back to something worse. A subsystem that stops issuing receipts
shows up as silent. The first full report on those receipts from real
runs found real problems:
subsystems filed under the wrong names, a memory recall that never
fired, and a memory tool nobody had ever used.
A phase that crashes inside a turn no longer prints its error into the
story. The detail goes to a crash log and telemetry, and the test suite
now fails any test during which a phase quietly crashed.
BoneAmanita Dev Diary: Telling the Truth About the Machine
September 23, 2026
Github Repo
After a summer away, I came back to BoneAmanita in September with a
different question. Not "what else can it do?" but "what does it
actually do?"
It Is a Prompt Builder
The README was rewritten from scratch. The old one described a
biological runtime with a metabolism, neurochemistry and dreams. The new
one opens by saying plainly what the software is: a state machine that
sits between you and a local language model and rewrites the model's
instructions on every turn. The effects it produces are produced by
instructions to a model, not by the simulation its naming implies.
That is not a demotion. It is the only honest place to build from.
A Prediction That Came True
A reader emailed with a falsifiable prediction: that the engine's
"governor", which solved a graph equation over its memory every turn,
was mostly theatre. I instrumented it before touching anything. The
graph term contributed 1.53% of the result, and the result was identical
whether the engine's voltage was 15 or 90. The prediction was right to
two decimal places. The governor was replaced with the simpler approach
the reader recommended.
Making Silence Visible
The engine's only output is prose, and prose looks the same whether the
physics underneath ran or quietly returned a default. So I went after
silent failure systematically. Exception handlers that swallowed errors
without a word were retired or made to report, with a budget the test
suite enforces. Tests were rewritten to check the state that should
result, not merely that nothing crashed. The credits were checked
against the code and corrected where they claimed more than was there.
The somatic side got the same treatment: one budget per turn for how
tired or strained the engine is, enforced by the engine itself rather
than by asking the model to cooperate.
Asking People Instead of Machines
To find out whether any of this made conversations better, I ran blind
comparisons against the same model with no engine at all. One early
result was wrong for a boring reason: Ollama's OpenAI-compatible
endpoint quietly runs every model with a 4,096-token context, and the
engine's prompts were being cut from the top without any warning. That
one is now caught and reported.
I also tested AI judges against control cases where the right answer is
known: the same reply three times over, a reply written for a different
turn, the very textbook answer the engine exists to avoid. None of the
six models cleared the bar. So AI judges are retired as verdicts. The engine is now evaluated by
people, through a blind reading panel that shows it exactly as it runs.
Along the way the codebase moved its core modules into an
engine/ package, and the documentation into
docs/, so a newcomer can find the front door.
BoneAmanita Dev Diary: Growing a Test Suite
July 7, 2026
Github Repo
Version 20.0.0, "Unbound Chronos", turned the engine into a persistent
background process. Instead of freezing between messages, it keeps
running: after five minutes without input it drifts into a REM cycle,
consolidates memory, and dreams. When you come back, it tells you what
it dreamt about.
Most of the work between May and July was less visible and more
important.
Tests Before Features
At the start of May the project had 22 test files. By July it had 44,
covering the adventure engine, the archetypes, the lexicon, the body,
the cortex, the cycle and the lore. The Python codebase grew from about
22,000 lines to about 31,000 over the same period, and a good share of
each release was deletion: whole versions removed more lines than they
added.
New Pieces
- A critic that reads each draft before the person sees it and can send it back
- A router that picks the relevant lines of a long document, like source code, to fit the prompt's budget
- A parser that reads rooms, exits and points of interest out of Adventure replies
- A pragmatics layer for what a message is doing, not just what it says
Then I stepped away for the summer. The engine would be waiting in
September, with a lot of passing tests and some harder questions.
BoneAmanita Dev Diary: From Paste-In File to Local Engine
April 28, 2026
Github Repo
BoneAmanita began as a file you pasted into a chatbot. By spring it had
become a program that runs the chatbot.
Giving the Machine a Metabolism
The first versions of the year gave the engine a body to reason with.
Every word generated, every memory retrieved and every contradiction
held costs energy. Stress builds up as toxicity. When the engine is
exhausted, its replies get shorter and plainer; when it is overloaded
with chaos, they get stranger. The point was never that the machine
feels tired. The point was to give a language model real constraints
that change how it talks, instead of an endless supply of agreeable
energy.
Running Alongside a Local Model
By March the single pasted file had become a set of Python modules with
their own data files, and by April it ran as its own program between
the person and a model served by Ollama, LM Studio or an OpenAI-style
API. It builds the model's instructions itself on every turn, clamps how
long the model may answer and how adventurous its sampling is, and keeps
two kinds of memory: a fast exact-match cache and a deeper index searched
by meaning.
Some ideas from this period still shape the engine. Short tags at the
start of a message, like [!l] for plain literal answers,
let a person say what they want instead of making the model guess. And
when the person seems worn out, the engine carries more of the load: it
simplifies its language and strips out noise.
Two Domains
The project split into two layers. BoneAmanita is the Python engine for
local models, where the constraints are enforced in code. The Hypervisor
is a text-only protocol for cloud models, pasted into their system
instructions, where no code can run. It carries the same ideas in words:
a council of distinct voices that must argue before anyone speaks, and a
firewall against replies that open with empty validation.
The Question Nobody Is Asking
The README from this time asked the question that still sits under the
whole project. People already form attachments to AI systems and change
their behavior around them, whatever one believes about what is inside.
So the question that matters is not whether the machine is conscious. It
is who built the frame.
BoneAmanita Dev Diary: A Physics Engine You Paste Into a Chatbot
December 31, 2025
Github Repo
BoneAmanita started on December 9, 2025 as version 0.1, "The Red Cap
Build". It ended the year at version 7.3. That is about sixty versions
in three weeks, and most of them had names.
What It Was
The first BoneAmanita was a single Python file that you uploaded or
pasted into ChatGPT, Claude or Gemini with an instruction to simulate
it. It did not force a personality on the model. It gave the model a
system of physics for language: rules for measuring the weight, speed
and drag of a piece of writing.
Narrative drag was the ratio of words to action. Stative verbs, the
"is" and "was" and "seems", slowed a sentence down. Empty pleasantries
at the start of a message were turned away at the door. The engine did
not want smooth text. It wanted text with something in it.
Three Weeks of Growth
Each version added an organ. A few from that month:
- The Virtual Cortex, which wrote its own critiques before the model saw your text
- Fifteen writing archetypes, from The Paladin to The Cosmic Trash Panda, each with its own rules
- The Codex, which gave the engine object permanence for names and facts
- A sense of time, so a ten-hour silence changed how it picked back up
- Defenses against gaming the physics with purple prose or stacked adverbs
- The Paradox Battery, which treated contradictions as fuel instead of waste
- A memory graph where heavily connected ideas gained gravity
Plenty of these were later simplified, merged or removed. Several are
still in the engine under other names. The drag, the archetypes and the
physics vocabulary all started here.
Next
By the end of December, the file had grown past what a chat window
could comfortably hold, and it had started to want things a pasted file
cannot have: persistence, time and a body. The next step was to stop
asking the model to pretend, and start building a program that enforces
the rules itself.