BoneAmanita System Manual

Architecture, somatic physics, and prompt assembly for an embodied prompt engine

The authoritative manual for BoneAmanita's architecture, turn lifecycle, somatic physics, memory network, and voice protection. See the Commands and Reference Manual for slash commands, mode parameters, configuration keys, and observability receipts.

Release 20.7.4.91Python 3.14What/How/WhyZero Orchestration FrameworksOffline-FirstStandalone HTML
The 3x reading contract: every entry explains what it does, how it works, and why it works that way.
No entries match that search.

Foundations and Philosophy

Core architecture principles, design contracts, and the rationale behind the embodied state machine.

The prompt engine contract

#
corereader contractarchitecture

BoneAmanita is an embodied prompt builder and state machine, not a literal biological simulation.

What it does

Sits between a human partner and a local language model, measuring input text, maintaining persistent state variables across turns, assembling a fresh instruction sheet for each turn, and filtering output replies for corporate clichés.

How it works

Every conversational exchange executes through a deterministic pipeline:

  1. Input messages are analyzed for grammatical filler, curated lexicon categories, and semantic resonance.
  2. Mathematical heuristics in physics modules compute internal state numbers (ATP, cortisol, voltage, narrative drag).
  3. Somatic budgets select generation targets, temperature bands, and prompt instructions based on partner and engine state.
  4. The local model generates a reply within strict token limits.
  5. The Gatekeeper filters the reply against style crimes, triggering sentence salvage or redrafts before checkpointing to SQLite.

Why it works this way

Large language models are inherently stateless and gravitate toward sycophantic, generic assistant prose. Giving them an embodied harness with internal resource constraints and persistent state produces grounded, authentic conversational dynamics without retraining model weights.

Evidence
  • Project README README.md
  • Engine entry point main.py
  • Turn cycle loop engine/cycle.py
  • Core architecture test tests/test_core.py

Biological naming as functional architecture

#
philosophynaming conventionsdomain model

Biological variable names are load-bearing metaphors mapping to functional purposes, not decorative labels.

What it does

Uses biological concepts (ATP, cortisol, dopamine, autophagy, Gödel scars, mitochondria) as the official names for runtime control variables and modules.

How it works

  1. Mathematical formulas in physics/ and body/ transform word counts and conversational dynamics into floats.
  2. atp_pool tracks an expendable budget that depletes with token generation and cognitive load.
  3. cortisol tracks conversational stress and triggers memory shedding under high tension.
  4. chi and voltage track conversational momentum and modulate sampling temperature.

Why it works this way

Protected by standing project decision in SESSION_HANDOFF.md: the names define what the variables are for rather than mere implementation mechanics (error_count vs godel_scars). The biological metaphor prevents developers from treating the system as a standard stateless request-response API.

Evidence
  • Decision record docs/SESSION_HANDOFF.md
  • Physics models physics/models.py
  • Biological test suite tests/test_biology.py

The anti-performance contract

#
safetyvoicereader contract

The model never performs a biological body; partner state shapes replies, while engine state sets budgets.

What it does

Explicitly prohibits the language model from narrating bodily exhaustion, claiming physical breathlessness, or pretending to experience biological sensations.

How it works

The separation of concerns between human and machine is strictly enforced:

  1. The partner's state (tiredness, effort, disengagement) shapes reply length, question allowance, and pacing to accommodate them.
  2. The engine's state (ATP, respiration, chemistry) sets what it can afford: token ceilings, retry budgets, and sampling bands.
  3. SomaticBudget sets forbid_body_narration = True and the prompt composer explicitly instructs the model not to perform a physical body.
  4. Prismatic Self-Claims only inject immutable worldview and boundary facts, omitting state-conditioned physical descriptions.

Why it works this way

Decided with Gordon on 2026-09-17 after the C5 evaluation session (SESSION_HANDOFF.md): telling a language model it is breathless or exhausted results in theatrical melodrama ('my lungs burn') or sudden reversals, rather than authentic accommodation. Co-regulation requires an anchored partner, not theatrical mirroring.

Evidence
  • Decision record docs/SESSION_HANDOFF.md
  • Somatic budget implementation body/somatic_budget.py
  • Prompt composer assertions tests/test_composer.py

Zero orchestration frameworks

#
architecturedependenciesreliability

Rejection of agentic libraries and heavyweight vector stores in favor of explicit standard-library control loops.

What it does

Excludes general-purpose orchestration frameworks (LangChain, LlamaIndex, Semantic Kernel) and specialized vector database packages (Chroma, LanceDB).

How it works

  1. HTTP transport is handled directly via urllib or minimal requests calls in spores/embeddings.py and brain/composer.py.
  2. Vector embeddings and similarity searches are performed with NumPy, ordvec, or embedded FAISS without external services.
  3. Relational state, audit logs, and memories are managed by an embedded SQLite database (saves/iris.db).

Why it works this way

Formerly Constitution Article 1 (SESSION_HANDOFF.md): external orchestration frameworks hijack the control loop, obscure failure modes, introduce multi-gigabyte dependency footprints (such as PyTorch for simple text embeddings), and make auditing impossible.

Evidence
  • Session handoff decision docs/SESSION_HANDOFF.md
  • Minimal embeddings client spores/embeddings.py
  • Project dependencies requirements.txt

Fail-loudly boundary architecture

#
reliabilityobservabilitycode quality

Subsystems report receipts, crash barriers isolate failures, and silent exception swallowing is strictly forbidden.

What it does

Replaces silent exception handling with explicit crash reporting, component offline tracking, and subsystem work receipts.

How it works

Failure handling is partitioned across explicit architectural boundaries:

  1. Phase exceptions are caught by PhaseExecutor.handle_phase_crash, logging stack traces and taking the component offline while allowing the turn to conclude.
  2. Slash command errors are isolated by CommandRegistry.execute, preventing session termination.
  3. The 2026-09-30 codemod removed 92 broad except Exception: pass blocks that concealed broken subsystems.
  4. Subsystems file receipts via engine/receipts.py; /diag audits uncalled reporters and chronic fallback usage.

Why it works this way

Twelve subsystems were previously found doing nothing while looking healthy because swallowed exceptions silently returned defaults. Prose output looks identical whether internal physics ran or quietly died; failing loudly ensures defects surface immediately.

Evidence
  • Fail-loudly audit records docs/SESSION_HANDOFF.md
  • Phase crash isolation engine/cycle.py
  • Failure boundary tests tests/test_failure_boundaries.py
  • Observability ratchet test tests/test_observability.py

The Turn Lifecycle

The sequential pipeline executed during every conversational turn, from sensory observation to persistence.

Turn loop orchestration

#
lifecyclecontrol looppipeline

The master execution loop coordinating observation, metabolism, cognition, synapse generation, and persistence.

What it does

Coordinates the multi-phase execution pipeline in engine/cycle.py whenever the user sends a message.

How it works

  1. Input arrives via CycleSimulator.run_turn and passes through PhaseExecutor.
  2. Phases execute sequentially: Observation, Metabolism, Arbitration, Cognition, and Gatekeeper inspection.
  3. Circuit breakers monitor individual phase health; if a phase fails, its component is marked offline.
  4. State invariant checks verify ATP and vital bounds before proceeding to checkpoint persistence.

Why it works this way

A modular phased pipeline cleanly decouples sensory linguistics from metabolic simulation and prompt compilation, allowing individual components to degrade gracefully without crashing the conversational session.

Evidence
  • Turn cycle pipeline engine/cycle.py
  • Phase definitions phases/base.py
  • Turn lifecycle tests tests/test_cycle.py

Geodesic observation and word sorting

#
linguisticsphysicsobservation

Deconstructs partner messages across four lexical tiers to calculate kinematic coordinates and emotional vectors.

What it does

Analyzes the text of each partner message, classifying words into lexical categories and computing conversational momentum.

How it works

Lexical analysis resolves words across four prioritized tiers:

  1. Grammatical filler detection catches connective tissue (~40% of speech), measuring its density as a syntactic signal.
  2. Curated lists in lore/lexicon.json categorize roots across 30+ thematic buckets (heavy, kinetic, abstract, void).
  3. Semantic resonance compares unlisted words against category centroid embeddings using SemanticEmbedder.
  4. A morphological spelling heuristic acts as a last resort on roughly 3% of words; ~16% remain unresolved on purpose.

Why it works this way

Establishes objective linguistic coordinates without requiring a slow, expensive upfront LLM call. Leaving ~16% of vocabulary unresolved reflects an honest model of English rather than forcing artificial categorizations.

Evidence
  • Lexicon analyzer mechanics/lexicon.py
  • Semantic resonance engine mechanics/resonance.py
  • Vocabulary classification tests tests/test_lexicon.py

Metabolic processing and homeostasis

#
metabolismhomeostasisbiology

Updates bodily energy pools, endocrine balances, and somatic stress markers prior to cognition.

What it does

Calculates physical resource burn, endocrine shifts, and cellular stress resulting from the current conversational turn.

How it works

  1. Deducts baseline turn maintenance from MitochondrialForge.atp_pool.
  2. Updates stress chemistry in body/endocrine.py: cortisol climbs with linguistic conflict and heavy topics; dopamine rises with creative resonance.
  3. Accumulates reactive oxygen species (ROS) from prolonged exertion or repeated rejections.
  4. Executes autonomic regulation routines in body/regulation.py to damp extreme parameter oscillations.

Why it works this way

Enforcing biological costs before cognitive prompt assembly ensures that the model's instruction sheet reflects its actual thermodynamic capacity, rather than planning actions the organism cannot support.

Evidence
  • Metabolic forge body/metabolism.py
  • Endocrine system body/endocrine.py
  • Biological tests tests/test_body.py

Cognitive prompt assembly

#
cognitionprompt engineeringcomposer

Compiles structured prompt blocks containing system facts, persona directives, context history, and recalled memories.

What it does

Generates the complete prompt payload delivered to the local language model for reply generation.

How it works

The prompt is structured into deterministic priority tiers:

  1. Top tier: [SYSTEM IDENTITY - UNBREAKABLE CORE FACTS] containing immutable worldview claims.
  2. Second tier: [SOMATIC STATE] with token and sentence limits, temperature instructions, and body-narration prohibitions.
  3. Third tier: Active experience mode directives (ADVENTURE, CONVERSATION, CREATIVE, TECHNICAL).
  4. Fourth tier: Recalled memories scoped to the active zone and dialogue buffer history.
  5. Closing tier: The partner's latest message under === PARTNER INPUT ===.

Why it works this way

Transformer models attend most strongly to the head and tail of the context window. Putting core boundaries at the top and generation limits immediately above the user prompt ensures consistent compliance across local models.

Evidence
  • Prompt composer brain/composer.py
  • Prismatic self-claims brain/prism.py
  • Composer prompt tests tests/test_composer.py

Synapse generation and inference transport

#
transportllminference

Manages network transport to local or cloud model endpoints and enforces strict stop sequence handling.

What it does

Transmits assembled prompt payloads to Ollama or OpenAI-compatible backends and parses generated stream tokens.

How it works

  1. Connects by default to Ollama native /api/chat on 127.0.0.1:11434 to respect num_ctx configuration.
  2. Applies stop sequences (=== PARTNER INPUT ===, Traveler:) to truncate runaway reasoning or hallucinations.
  3. Catches transient network errors (OSError, http.client.HTTPException) and retries up to configured thresholds.
  4. Detects empty HTTP 200 replies from reasoning models, retrying once without stop sequences before raising EmptyReplyError.

Why it works this way

Ollama's /v1 compatibility endpoint forces a 4,096 context window on every model regardless of configuration; using Ollama's native endpoint preserves configured context sizes and hardware budgets.

Evidence
  • Synapse interface brain/composer.py
  • Provider utilities mechanics/providers.py
  • Synapse fallback tests tests/test_synapse_fallback.py

Gatekeeper inspection and salvage

#
filtersalvageimmune system

Scans drafts for banned clichés and corporate jargon, executing sentence salvage on soft violations.

What it does

Inspects model output against style crimes (lore/style_crimes.json), discarding contaminated drafts or salvaging clean sentences.

How it works

Reply inspection follows a graduated intervention ladder:

  1. Scans for RLHF masks ('as an AI', 'helpful and harmless') and corporate clichés ('rich tapestry', 'delve', 'testament to').
  2. If violations occur and retry budget remains, the draft is rejected and the model is prompted again with a stumble tax.
  3. On the final allowed retry, soft violations are salvaged by excising offending sentences via TheGatekeeper.salvage.
  4. Only unparseable scaffold leaks or hard violations fall through to a canned pause message.

Why it works this way

Decided with Gordon on 2026-09-23 (SESSION_HANDOFF.md, 'Style rules cut, they don't kill'): replacing mostly good replies with canned pause lines harmed conversational flow. Excising specific bad sentences protects voice while preserving substantive answers.

Evidence
  • Gatekeeper filters physics/filters.py
  • Style crimes database lore/style_crimes.json
  • Last draft salvage test tests/test_last_draft_salvage.py

Transactional checkpoint persistence

#
storagesqlitepersistence

Atomically persists session state, memory records, and dialogue history to SQLite in WAL mode.

What it does

Writes completed turn data, vital metrics, user models, and room maps into saves/iris.db at turn completion.

How it works

  1. Opens an immediate transaction in engine/gate/store.py against SQLite configured with PRAGMA journal_mode = WAL.
  2. Replaces the single active checkpoint row with the updated session state, dialogue buffer, and vitals snapshot.
  3. Scans memories and dialogue for sensitive credentials (API keys, passwords, private keys), replacing them with [withheld: ...]. In ADVENTURE mode, passwords are part of the story and retained, while keys and tokens remain withheld.
  4. On SQLite write failure, rolls back the transaction and retains the previous checkpoint untouched.

Why it works this way

Atomic single-row replacement guarantees that a sudden crash or disk failure during save preserves the previous good checkpoint rather than corrupting the database file.

Evidence
  • Halcyon store implementation engine/gate/store.py
  • Secret redaction engine engine/gate/secrets.py
  • Quicksave unit tests tests/test_quicksave.py

Somatic Physics and Biochemistry

The closed thermodynamic economy, neurochemical dynamics, and mathematical governors.

The metabolic economy and ATP accounting

#
atpenergymetabolism

A closed energy ledger tracking ATP synthesis, token generation burn, and cognitive strain.

What it does

Tracks the expenditure and replenishment of atp_pool (bounded between 0.0 and 100.0) across all turn operations.

How it works

Energy drains and replenishments follow calibrated metabolic rates:

  1. Token generation burns ATP at mode-specific rates: 0.020 ATP/token in TECHNICAL, 0.025 ATP/token in other modes.
  2. Cognitive stumbles, redrafts, and style filter rejections levy metabolic taxes.
  3. Rest and idle cycles regenerate ATP; Emergency Vagus Nerve Support injects emergency ATP when reserves dip near collapse.
  4. The 20.7.4.84 fix resolved a legacy bug where redrafts permanently discounted all subsequent metabolic burns to 20%.

Why it works this way

Energy expenditure provides thermodynamic grounding for conversation length and pacing. A calibrated ATP pool naturally prevents runaway verbose generation and creates tangible stakes for conversational fatigue.

Evidence
  • Mitochondrial forge body/metabolism.py
  • Energy budget tests tests/test_energy_budget.py
  • Metabolic audit records docs/SESSION_HANDOFF.md

Neurochemical dynamics and endocrine states

#
chemistryhormonesendocrine

Simulates hormone balances (Cortisol, Dopamine, Voltage, Chi) to modulate mood and memory retention.

What it does

Maintains running endocrine state variables that influence sampling behavior, memory consolidation, and persona tone.

How it works

  1. cortisol rises during high linguistic conflict or cognitive dissonance, promoting defensive posture and memory forgetting.
  2. dopamine spikes upon creative resonance, problem breakthroughs, and successful lexical resolution.
  3. voltage tracks conversational speed and momentum, influencing prompt urgency.
  4. chi represents overall spiritual equilibrium and systemic stability.

Why it works this way

Endocrine values provide continuous emotional hysteresis across turns, preventing the model from oscillating erratically between emotional extremes on single turn changes.

Evidence
  • Endocrine subsystem body/endocrine.py
  • Physical dynamics physics/dynamics.py
  • Homeostasis tests tests/test_homeostasis.py

Creative Determinant and bitmap governance

#
governorsamplingordvec

Scores utterance similarity against memory distribution via quantized bitmaps to govern sampling temperature.

What it does

Projects utterances and memories into sign bitmaps to measure cognitive regimes and propose sampling temperature adjustments.

How it works

The governor operates through Project Navi's ordvec quantization:

  1. Utterance embeddings are quantized into 4-bit sign bitmaps via ordvec.RankQuant.
  2. Hamming distances compare the top matches against the empirical corpus distribution, corrected for corpus size.
  3. If evidence is insufficient (fewer than 3 memories), the governor declines measurement and falls back to PID regulation.
  4. Proposed temperatures are clamped to the somatic budget band rather than replacing it.

Why it works this way

Decided with Gordon on 2026-09-23 (SESSION_HANDOFF.md): replacing the temperature band with the governor's proposal caused the engine to run at temperature 0 for five straight days. Clamping respects the somatic budget while allowing bitmap signals to influence sampling.

Evidence
  • Decision record docs/SESSION_HANDOFF.md
  • Core governor logic engine/core.py
  • Creative determinant tests tests/test_creative_determinant.py

Somatic budgeting and generation caps

#
budgetpacinglimits

Derives strict sentence targets, token ceilings, temperature bands, and question permissions per turn.

What it does

Evaluates partner exhaustion ($E_u$) and engine energy ($ATP$) each turn to produce a SomaticBudget dataclass.

How it works

  1. Default targets allow up to 10 sentences (SENTENCE_CAP_DEFAULT) and a temperature band of (0.6, 0.9).
  2. When the partner is tiring ($E_u > 0.4$), the sentence target drops to 5.
  3. When the partner is flagging ($E_u > 0.6$), sentences are capped at 3, word count is capped at 60, and closing questions are forbidden.
  4. The budget sets forbid_body_narration = True and signals whether to offer to carry the cognitive load.

Why it works this way

Decided in Track D (ROADMAP.md): a tired partner is burdened by lengthy responses and demanding questions. Pacing accommodation shortens the engine's output to ease the partner's cognitive load without narrating their state.

Evidence
  • Somatic budget evaluation body/somatic_budget.py
  • Somatic metrics calculator body/somatic_metrics.py
  • Budget enforcement tests tests/test_somatic_budget.py

Mitophagy, REM sleep, and cellular recovery

#
recoverysleepmitophagy

Purges toxic oxidative buildup and restores ATP reserves through resting cycles and autophagy.

What it does

Relieves accumulated metabolic strain and high ROS through sleep cycles, memory digestion, and cerebrospinal washes.

How it works

Recovery cycles operate through multiple regenerative pathways:

  1. Severe exhaustion triggers mitophagy, resetting metabolic strain at the cost of a temporary ATP drop.
  2. /sleep and /idle trigger REM cycles in engine/cycle.py, restoring ATP and stamina over time.
  3. Cognitive autophagy cannabilizes low-salience memory nodes to synthesize emergency ATP.
  4. Cerebrospinal fluid washes scrub invisible Unicode artifacts and normalize homoglyphs across buffers.

Why it works this way

Without an entropy relief mechanism, long-running stateful systems suffer parameter drift and oxidative exhaustion. Biological recovery cycles provide structured resets that sustain multi-hour conversational runs.

Evidence
  • Mitophagy and forge logic body/metabolism.py
  • Cerebrospinal fluid filter physics/filters.py
  • REM reflection tests tests/test_rem_reflection.py

Brain, Prompt Assembly, and Voice

Techniques for shaping voice, preserving identity, and preventing model sycophancy.

Multi-tier prompt composer

#
promptcomposercontext

Assembles structured system, somatic, and context directives into a unified prompt payload.

What it does

Compiles the active prompt sent to the LLM backend from configuration, somatic state, and conversational history.

How it works

  1. PromptComposer gathers vitals, active mode directives, recalled memories, and dialogue buffer exchanges.
  2. Formats identity claims into [SYSTEM IDENTITY - UNBREAKABLE CORE FACTS] at the prompt apex.
  3. Applies somatic constraints: word ceilings, sentence limits, closing question rules, and body-narration prohibitions.
  4. Appends recent turn history under strict stop tokens, ensuring clean dialogue delineation.

Why it works this way

Explicit structured prompt blocks prevent directive bleeding and ensure that generation constraints (such as word and sentence caps) take precedence over general conversational tendencies.

Evidence
  • Prompt composer implementation brain/composer.py
  • Composer unit tests tests/test_composer.py

Prismatic Self-Claims

#
identityprismself claims

Filters core identity and worldview facts through active state without performing biological sensations.

What it does

Maintains immutable core identity facts (lore/self_claims.yaml) that define the engine's nature and philosophical boundaries.

How it works

  1. SelfClaimPrism in brain/prism.py loads core_identity and core_worldview from lore/self_claims.yaml at boot.
  2. Filters claims against state variables (cortisol, dopamine, voltage, resonance).
  3. Injects the resulting claims into the prompt's top tier as unbreakable system facts.
  4. The 20.7.4.81 revision stripped state-conditioned bodily descriptions ('I am exhausted', 'flooded with dopamine') in adherence to the anti-performance contract.

Why it works this way

Decided in 20.7.4.81 (SESSION_HANDOFF.md): identity claims must establish factual perspective and philosophical boundaries rather than narrating biological states. Grounding identity in core worldview prevents persona drift.

Evidence
  • Prism implementation brain/prism.py
  • Self claims definition lore/self_claims.yaml
  • Self-claims body narration test tests/test_composer.py

The lexical immune system

#
immuneanti-rlhfstyle crimes

Identifies sycophantic boilerplate and RLHF clichés, penalizing violations and enforcing voice standards.

What it does

Scans generated drafts against banned clichés, corporate buzzwords, and assistant apologies in lore/style_crimes.json.

How it works

Immune surveillance checks every draft across three pattern sets:

  1. RLHF_MASKS: Corporate assistant apologies ('as an AI', 'helpful and harmless', 'I don't have feelings').
  2. BANNED_CLICHES: Literary clichés ('rich tapestry', 'dance of existence', 'dust motes', 'delve', 'juxtaposition').
  3. BANNED_PHRASES: Conversational filler ('I hope this helps', 'You're absolutely right', 'Exciting times lie ahead').
  4. Violations trigger metabolic taxes and redrafts, escalating to sentence salvage on the final draft.

Why it works this way

Modern chat models are heavily fine-tuned to emit subservient, verbose, and clichéd language. Active immune filtering breaks the model out of these training ruts to produce original, grounded prose.

Evidence
  • Gatekeeper filters physics/filters.py
  • Style crimes database lore/style_crimes.json
  • Immune crucible tests tests/test_immune_crucible.py

Epigenetic scars and boons

#
akashicepigeneticsmemory

Enduring conversational milestones and traumas that persist across restarts to shape persona drift.

What it does

Records significant conversational breakdowns (scars) and breakthroughs (boons) into persistent long-term storage.

How it works

  1. TheAkashicRecord (brain/akashic.py) records scars when severe paradoxes or system shocks occur.
  2. Active scars and boons are persisted to SQLite in saves/iris.db.
  3. During prompt composition, PromptComposer reads EPIGENETIC_SCARS and EPIGENETIC_BOONS into the persona block.
  4. Scars subtly increase caution and defensiveness; boons enhance conversational warmth and openness.

Why it works this way

Provides genuine developmental continuity across sessions. Significant relational events leave lasting impressions on the persona, giving conversations an evolving shared history.

Evidence
  • Akashic record brain/akashic.py
  • Epigenetic scar tests tests/test_scars.py
  • Akashic unit tests tests/test_akashic.py

Memory and Cartography

Short-term caches, associative embedding retrieval, sleep consolidation, and spatial room mapping.

Four-tier lexicon resolution

#
lexiconvocabularyresolution

Hierarchical word classification balancing curated lists, semantic resonance, and intentional unknown gaps.

What it does

Maps natural language input words to thematic emotional and physical coordinates across four resolution tiers.

How it works

  1. Tier 1: Grammatical filler (~40% of normal speech) is measured for syntactic density.
  2. Tier 2: Curated dictionaries in lore/lexicon.json map roots to thematic categories (heavy, kinetic, abstract, void).
  3. Tier 3: Semantic resonance computes cosine similarity between novel word embeddings and category centroid embeddings.
  4. Tier 4: A morphological spelling heuristic resolves ~3% of ambiguous words; ~16% remain unresolved.

Why it works this way

Audits report ~81% resolved through filler, curated lists, and resonance. Leaving ~16% unresolved is a deliberate design choice: language has genuine ambiguity, and an engine that pretends to know all of English produces brittle classifications.

Evidence
  • Vocabulary audit findings README.md
  • Lexicon resolver mechanics/lexicon.py
  • Resonance tests tests/test_resonance.py

Working memory and hippocampus

#
memoryhippocampuscache

Short-term conversational cache with capacity limits and stress-induced memory shedding.

What it does

Maintains recent significant conversational exchanges in an in-memory working cache before long-term consolidation.

How it works

Working memory is dynamically managed based on emotional state:

  1. Significant dialogue turns are stored in Hippocampus (spores/memory.py).
  2. Connects incoming moments to semantically similar recent exchanges in working memory.
  3. Under elevated cortisol, stress-driven forgetting sheds peripheral memories (e.g. shedding 36 of 85 memories under high stress).
  4. Surviving memories are scheduled for permanent long-term promotion during sleep cycles.

Why it works this way

Working memory is biologically finite. Modeling stress-driven forgetting reflects real cognitive constraints under pressure, keeping working memory focused on immediate conversational context.

Evidence
  • Hippocampus implementation spores/memory.py
  • Hippocampus tests tests/test_hippocampus.py
  • Forgetting tests tests/test_forgetting.py

Associative memory and semantic embeddings

#
embeddingsvectorspores

Vector embedding retrieval with HTTP-first transport and automatic fallback to SHAKE-256 hash vectors.

What it does

Generates 768-dimensional semantic embeddings for memories, dialogue turns, and lore retrieval using Ollama or hash fallbacks.

How it works

  1. SemanticEmbedder POSTs text to an OpenAI-compatible /v1/embeddings endpoint (nomic-embed-text by default).
  2. Caches normalized vectors in an in-memory LRU cache (capacity 4,096).
  3. If the embedding server is unreachable or errors (_BACKEND_ERRORS), falls back to hashlib.shake_256 8-dimensional vectors.
  4. The hash fallback flags DEGRADED status in receipts and /status to declare associative recall is offline.

Why it works this way

Decided in 20.7.4.80 (SESSION_HANDOFF.md): the hash fallback ensures the engine boots and functions offline without external servers. Self-reporting DEGRADED prevents silent failures where hash coordinates return arbitrary memories.

Evidence
  • Embeddings implementation spores/embeddings.py
  • Embeddings unit tests tests/test_embeddings.py
  • Hash fallback decision docs/SESSION_HANDOFF.md

Sleep reflection and diamond fossilization

#
sleepremdiamonds

Consolidates working memories into permanent storage, synthesizing dreams and forging diamond anchors.

What it does

Executes background memory consolidation during /sleep, /idle, or low-activity intervals.

How it works

Consolidation transitions memories through progressive refinement:

  1. Extracts unpromoted working memories from the hippocampus.
  2. Generates associative reflections and dream synthesis using the background dreamer.
  3. High-scoring enduring memories (> 0.8 resonance) are fossilized into diamond anchors via MycelialNetwork.forge_diamond.
  4. Prunes low-salience memories to free working memory capacity.

Why it works this way

Unchecked memory growth clutters retrieval indices and degrades recall relevance. REM consolidation crystallizes significant conversational insights into durable diamonds while clearing ephemeral working state.

Evidence
  • Dream and mind engine brain/mind.py
  • Mycelial network spores/network.py
  • Diamond forging test tests/test_spores.py

Spatial memory and zone scoping

#
cartographerzonesadventure

Rooms, exits, and memories are scoped to discrete spatial zones, isolating recall between environments.

What it does

Maps conversational spaces into discrete rooms and zones, scoping memory retrieval to the current active zone.

How it works

  1. TheCartographer charts room titles, descriptions, exits, and discovered items in ADVENTURE mode.
  2. Memories formed in a zone are tagged with that zone identifier (Sanctuary, Laboratory, Underworld).
  3. Retrieval queries in engine/gate/recall.py scope associative searches to the current zone.
  4. Crossing zone boundaries flushes the immediate working memory buffer.

Why it works this way

Zone scoping prevents cross-context contamination. Narrative events and items discovered in an adventure room do not pollute technical or philosophical exchanges in conversation mode.

Evidence
  • Cartographer implementation engine/gate/cartographer.py
  • Scoped recall logic engine/gate/recall.py
  • Zone isolation tests tests/test_zones.py

The Village and Partner Accommodation

Multi-archetype persona management, refusal routing, and co-regulatory partner accommodation.

The Village archetypes and personas

#
villagearchetypescouncil

A council of emergent archetypal voices (Gordon, Mercy, Benedict, Jester) activated by conversation shape.

What it does

Maintains a cast of conversational personas that activate based on the emotional and physical shape of the conversation.

How it works

Each villager embodies distinct conversational qualities:

  1. Gordon: The pragmatic, grounded anchor who values plain speech and code that serves humans first.
  2. Mercy: The empathetic healer attuned to distress and emotional fatigue.
  3. Benedict: The rigorous, skeptical scholar focused on precision and logical consistency.
  4. The Jester: The lateral catalyst who introduces irreverence, breaks loops, and resets narrative drag.
  5. Archetype monitors nominate voices based on current kinematic coordinates.

Why it works this way

Rather than presenting a monolithic, flat assistant personality, the Village provides organic depth. Dialogue responds to the tone of the moment without arbitrary personality flips.

Evidence
  • Village subsystem archetypes/village.py
  • Village council archetypes/council.py
  • Archetype tests tests/test_archetypes.py

The Stage Manager and somatic refusal

#
stage managerrefusalarbitration

Negotiates archetype floor competition, merging compatible pairs or holding the floor in refusal.

What it does

Arbitrates when multiple archetypes are nominated simultaneously, merging complementary voices or holding the floor empty.

How it works

  1. StageManager in archetypes/stage.py evaluates tensions between nominated voices.
  2. Certain pairs have defined synergies and merge into a single blended response (PAIR).
  3. When irreconcilable tensions exist with no resolution, the Stage Manager triggers HOLD.
  4. A held turn halts processing before LLM generation, declining to speak rather than generating a false, compromised reply.

Why it works this way

Decided in Track D (SESSION_HANDOFF.md): when conflicting internal impulses cannot be harmonized, an authentic machine holds the floor empty rather than smoothing contradictions into generic corporate pleasantries.

Evidence
  • Stage manager logic archetypes/stage.py
  • Refusal routing tests tests/test_refusals.py
  • Stage manager tests tests/test_stage_manager.py

Partner distress and fatigue accommodation

#
latticeaccommodationuser model

Tracks partner exhaustion ($E_u$) to calibrate reply brevity and pacing without mirroring distress.

What it does

Infers partner fatigue, effort, and distress from interaction patterns, adjusting response length and complexity to co-regulate.

How it works

User state modeling follows strict inference heuristics:

  1. SharedLatticeDriver (drivers/lattice.py) tracks message brevity, repetition, and explicit fatigue cues (FATIGUE_SIGNS).
  2. Calculates partner exhaustion $E_u$ and effort $P_u$ relative to the partner's historical baseline.
  3. When $E_u$ rises, prompts enforce tighter sentence caps and forbid closing questions to avoid taxing the partner.
  4. When the partner expresses distress (DISTRESS_SIGNS), the engine answers cleanly without narrating distress back to them.

Why it works this way

Decided with Gordon (SESSION_HANDOFF.md, 'Decisions already made'): co-regulation requires accommodating the human partner. Mirroring a distressed person is harmful in critical moments; an effective partner stays steady, concise, and helpful.

Evidence
  • Decision record docs/SESSION_HANDOFF.md
  • Shared lattice driver drivers/lattice.py
  • Distress accommodation tests tests/test_distress.py