BoneAmanita Reference Manual
Slash commands, runtime modes, configuration keys, file formats, and diagnostic receipts
Complete operational and interface reference for BoneAmanita. Return to the System Manual for internal architecture, somatic physics, and design rationale.
Release 20.7.4.91CLI ReferenceTuning PresetsTelemetry & DiagStandalone HTML
The 3x reading contract: every entry explains what it does, how it works, and why it works that way.
No entries match that search.
In-Session Slash Commands
Interactive commands executed during a session via the terminal interface or headless runner.
vitalsinspectionhud
Displays system vitals, energy pools, memory backend status, and user model inferences.
/status
What it does
Prints current health, stamina, and ATP energy bars, reports the status of the Mnemonic Arcade (nominal or degraded), and displays what the engine currently believes about the partner.
How it works
- Queries
CommandStateInterface.get_vitals() to pull health, stamina, and ATP pool levels. - Renders the Mnemonic Arcade status, reporting whether the embedding backend is active (
nominal) or operating on fallback hash coordinates (DEGRADED). - Renders the partner state model from
SharedLatticeDriver (You: steady, tiring, or flagging), allowing the partner to see and disagree with the heuristic inference.
Why it works this way
Radical transparency is required in stateful systems. Giving the partner visibility into inferred exhaustion and backend degradation prevents the system from silently acting on mistaken assumptions.
Evidence
- Status command implementation
mechanics/commands.py - Vitals model
body/models.py - Command routing tests
tests/test_command_routing.py
Related: diag · memory and memory why · truth and hud
diagnosticsreceiptsobservability
Audits this turn's receipts, reporting nominal operations, degraded fallbacks, and silent subsystems.
/diag
What it does
Prints the receipt ledger for the current turn, detailing which subsystems ran, what they produced, and whether any promised reporter remained silent.
How it works
The diagnostic readout formats three critical categories of observability records:
- Turn receipts: Each registered subsystem (
composer.compose, cortex.somatic, halcyon.gate) records its result count and whether it ran on fallback. - Silent reporters: Subsystems expected on the core roll call that failed to file any receipt during the turn.
- Chronic fallbacks: Subsystems that consistently run on degraded fallback paths across every turn.
- Chronically empty: Subsystems that report execution but consistently return zero results.
Why it works this way
From the outside, broken subsystems return default values that produce plausible-looking text. The receipt ledger makes silent failures and degraded vector backends inspectable to human operators.
Evidence
- Diag command implementation
mechanics/commands.py - Receipt ledger
engine/receipts.py - Receipt audit tests
tests/test_receipts.py
Related: status · Subsystem receipts and the core roll call
memoryinspectiongate
Inspects indexed associative memories, world nodes, and specific gate retention decisions.
/memory [word] | /memory why <name>
What it does
Lists indexed memories filtered by keyword, or traces why a specific memory or room was retained by the Gatekeeper.
How it works
/memory without arguments lists recent memories and their access frequency./memory <word> filters memories containing or resonating with the specified keyword./memory why <name> traces a specific memory back to the partner turn that spawned it, displaying the Gatekeeper's checks and confidence scores.
Why it works this way
Associative memory should not be an inscrutable black box. Providing an audit trail from recalled memories back to the originating input allows users to verify what the engine is learning.
Evidence
- Memory command handler
mechanics/commands.py - Memory inspection report
engine/gate/report.py - Memory command tests
tests/test_memory_command.py
Related: status · diag · The Halcyon SQLite store (saves/iris.db)
modeexperiencetuning
Switches between experience modes (Adventure, Conversation, Creative, Technical) while preserving dialogue.
/mode <ADVENTURE | CONVERSATION | CREATIVE | TECHNICAL>
What it does
Changes the active operational framing and associated tuning preset without destroying conversational dialogue.
How it works
- Validates the target mode against
BonePresets.MODES. - Updates engine state, adjusts village archetype suppression rules, and applies mode-specific tuning presets.
- Carries over the live dialogue buffer, user model inferences, and memories across mode switches.
- If switching to
ADVENTURE, re-engages room cartography and inventory; if switching away, stores adventure state.
Why it works this way
Decided with Gordon on 2026-10-04 (SESSION_HANDOFF.md, 'Decisions already made'): switching modes mid-conversation must preserve dialogue history, secrets included, rather than artificially siloing the conversational context.
Evidence
- Mode switch command
mechanics/commands.py - Presets and mode definitions
engine/presets.py - Mode switch tests
tests/test_mode_switch.py
Related: preset and tune · ADVENTURE mode · CONVERSATION mode · CREATIVE mode · TECHNICAL mode
tuningphysicsconfiguration
Loads somatic tuning presets or dynamically alters individual physics thresholds in-flight.
/preset <name> | /tune <path=value>
What it does
Applies named somatic presets (Zen Garden, Thunderdome, Sanctuary, Laboratory) or overrides specific configuration keys without changing the active experience mode.
How it works
/preset <name> looks up the preset in BonePresets and applies physics overrides (voltage ceilings, drag floors, decay rates)./tune <path=value> mutates deep nested keys in lore/tuning_presets.json dynamically.- Changes take effect immediately on the next turn cycle without restarting the process.
Why it works this way
Decoupling somatic tuning from experience mode allows experimenters to test varied energetic and emotional dynamics (such as high-voltage stress in conversation) without altering UI framing.
Evidence
- Preset command handler
mechanics/commands.py - Tuning presets
lore/tuning_presets.json - Tuning tests
tests/test_tune.py
Related: mode · Somatic tuning presets
sleepremregeneration
Triggers a restorative REM cycle, regenerating ATP and stamina while synthesizing dreams.
/sleep | /idle
What it does
Puts the engine into a resting REM state that regenerates metabolic pools, processes working memories, and produces dream narratives.
How it works
- Regenerates ATP pool and stamina in
body/metabolism.py. - Invokes
TheDreamEngine in brain/mind.py to synthesize thematic dreams from unpromoted working memories. - Consolidates salient memories and forges diamond anchors via
MycelialNetwork.forge_diamond. - Washes cerebrospinal fluid, clearing accumulated homoglyphic and invisible Unicode noise.
Why it works this way
Provides a first-class operational mechanism for recovering from conversational exhaustion and purging oxidative stress without terminating the session or resetting the database.
Evidence
- Sleep command implementation
mechanics/commands.py - REM cycle logic
engine/cycle.py - REM reflection tests
tests/test_rem_reflection.py
Related: rest, zen, and flush · status
resetflushrecovery
Executes a hard conversational reset, severing context, zeroing narrative drag, and restoring vitals.
/rest | /zen | /flush
What it does
Clears the active dialogue context buffer, purges conversational drag and trauma, and restores health and stamina to nominal baselines.
How it works
- Clears
TheCortex.dialogue_buffer, severing active context window history. - Resets
narrative_drag to zero and purges accumulated trauma markers. - Restores
atp_pool, health, and stamina to full capacity. - Preserves long-term memories in
saves/iris.db while clearing transient turn dynamics.
Why it works this way
When conversational context falls into unproductive loops, high drag, or hallucination traps, a hard contextual reset restores clean operating conditions without destroying persistent learnings.
Evidence
- Rest command implementation
mechanics/commands.py - Cortex context management
brain/cortex.py - Context window tests
tests/test_context_window.py
Related: sleep and idle · save and export
saveexportfractalos
Forces an immediate transactional checkpoint to disk, or exports the adventure world in FractalOS format.
/save | /export [path]
What it does
Explicitly writes current session state to saves/iris.db, or exports charted adventure rooms into a standalone JSON file.
How it works
/save invokes Store.save_checkpoint, replacing the active checkpoint row within an immediate SQLite transaction./export extracts room nodes, exits, and inventory from TheCartographer, formatting them as a FractalOS adventure JSON (default saves/fractal_adventure.json).- Sensitive credentials are automatically redacted with
[withheld: ...].
Why it works this way
Explicit save commands give operators control over persistence boundaries before closing sessions, while export functionality enables interoperability with external tools like FractalOS.
Evidence
- Save and export handlers
mechanics/commands.py - Store transaction manager
engine/gate/store.py - Cartographer export tests
tests/test_cartographer.py
Related: The Halcyon SQLite store (saves/iris.db) · Substrate file management and secret withholding
hudtruthtelemetry
Adjusts reality ambiguity levels and configures diagnostic readout verbosity.
/truth <0-3> | /hud <warm | lite | core | deep>
What it does
Modulates reality ambiguity (controlling hallucination tolerance) and sets terminal telemetry depth.
How it works
/truth <0-3> adjusts the reality layer parameter, altering how strictly statements must match ground facts./hud <depth> configures VSL display verbosity: warm (minimal natural prose), lite (vitals summary), core (turn metrics), or deep (full receipts, embeddings, and gate decisions).
Why it works this way
Different users require different diagnostic visibility. Technical operators need deep telemetry to trace gate decisions, while casual conversation benefits from a minimal, warm interface.
Evidence
- HUD and truth commands
mechanics/commands.py - UX string definitions
lore/ux_strings.json - Telemetry tests
tests/test_observability.py
Related: status · diag
substratesecurityfiles
Reviews, approves, or denies pending file modifications requested by the model under output/.
/allow <path> [keep] | /deny <path>
What it does
Grants or revokes permission for the model to create or edit files in the output/ directory.
How it works
Substrate operations adhere to strict permission policies:
- In TECHNICAL mode,
<write_file> blocks write files directly under output/. - If the file was created by the engine, it may edit it freely; existing pre-existing files are held pending approval.
/allow <path> approves a held edit; appending keep grants permanent permission for that file or folder./deny <path> discards the proposed change and notifies the engine of the refusal.- Decisions are recorded in
output/.substrate_ledger.json.
Why it works this way
Prevents unauthorized file corruption or unexpected file system mutations by local models, ensuring that external modifications are always gated by explicit human consent.
Evidence
- Substrate approval commands
mechanics/commands.py - Substrate tools
mechanics/tools.py - Substrate save directive tests
tests/test_save_file_directive.py
Related: Substrate file management and secret withholding · TECHNICAL mode
journal, hallucinate, and shuffle
#
creativejournaljester
Generates narrative session diaries, unlocks thermal hallucination locks, or executes the Jester's gambit.
/journal | /hallucinate | /shuffle
What it does
Executes specialized narrative actions: generating session summaries, forcing lateral leaps, or breaking conversational loops.
How it works
/journal generates a reflective narrative diary summarizing key themes and memories of the current session./hallucinate disengages thermal locks to force a high-voltage, creative paradigm shift./shuffle executes the Jester's Gambit: resets narrative drag, shifts conversational velocity, and introduces lateral randomness.
Why it works this way
Creative exploration and long-running story generation occasionally stall in conversational ruts. Dedicated narrative tools provide deliberate interventions to spark novelty.
Evidence
- Creative commands handler
mechanics/commands.py - Jester archetype logic
archetypes/village.py - Whimsy mechanics tests
tests/test_whimsy.py
Related: CREATIVE mode · preset and tune
Experience Modes and Tuning Presets
The four official runtime modes and their associated somatic tuning parameters.
modeadventurecartography
The default exploration experience featuring parser-game framing, room cartography, and inventory.
What it does
A survival and exploration mode with interactive inventory, room-to-room navigation, and Sanctuary tuning.
How it works
- Tuned to
SANCTUARY preset with target voltage 7.0 and drag 2.0. - Enables room cartography, item discovery, and inventory tracking (
/look, /use, /inventory, /map). - Allows ATP drain and chaos taxes; gates narrative surprises behind item interactions.
- World state and room maps are fully exported via
/export.
Why it works this way
Provides an embodied spatial world where thoughts and conversation have physical anchors. Exploring rooms grounds long-term memory in tangible environments.
Evidence
- Mode definition
engine/presets.py - Adventure mode tests
tests/test_adventure.py - Cartographer tests
tests/test_cartographer.py
Related: CONVERSATION mode · CREATIVE mode · TECHNICAL mode · Somatic tuning presets
modeconversationzen
Pure dialogue mode tuned to Zen; disables inventory, room maps, and entropy for uninterrupted connection.
What it does
Focuses entirely on dialogue and connection, suppressing world mechanics and inventory systems.
How it works
- Tuned to
ZEN preset with low voltage ceilings (10.0) and zero baseline drag. - Suppresses spatial village archetypes (Gordon, Navigator, Cartographer, Tinkerer).
- Disables ATP depletion and chaos taxes, allowing extended philosophical or emotional dialogue without energy starvation.
- Carries over all dialogue and secrets when switching to or from other modes.
Why it works this way
Not every interaction requires game mechanics or spatial exploration. Conversation mode allows intimate, calm dialogue free from inventory clutter and metabolic pressure.
Evidence
- Mode configuration
engine/presets.py - Handoff decisions
docs/SESSION_HANDOFF.md - Mode switch tests
tests/test_mode_switch.py
Related: ADVENTURE mode · CREATIVE mode · TECHNICAL mode · Somatic tuning presets
modecreativewriting
High-voltage writing assistance tuned to Manic; loosens constraints while strictly forbidding menus and bullets.
What it does
Writing assistant mode optimized for narrative generation, high creative tension, and bold prose.
How it works
Creative mode enforces specific stylistic guidelines:
- Tuned to
MANIC preset with high voltage floor override (70.0) and Crucible HOT active throughout. - Rule 3 (20.7.4.84 revision): Offers 2–3 bold directions only when the partner asks for ideas; when given a direction, writes it fully without menus.
- Strictly limits bulleted lists and length, holding the narrative voice to first-person 'I'.
- Disables inventory and vitals HUD to maintain full focus on narrative prose.
Why it works this way
Language models default to lazy multi-choice menus ('Option 1... Option 2...') instead of writing the actual story. Creative mode forces the model to commit to continuous prose.
Evidence
- Mode configuration
engine/presets.py - Creative probe records
docs/SESSION_HANDOFF.md - Creative mode tests
tests/test_mode_failures.py
Related: ADVENTURE mode · CONVERSATION mode · TECHNICAL mode · Somatic tuning presets
modetechnicalcoding
Code generation and engineering mode with tight constraints, 0.020 ATP/token, and deep telemetry.
What it does
Engineering mode designed for software development, debugging, and supervised file generation.
How it works
- Tuned to
DEBUG preset with deep HUD readout by default. - Permits
<write_file> code blocks targeting output/, subject to /allow and /deny approvals. - Charges 0.020 ATP per generated token (adjusted from 0.025 in 20.7.4.85 to accommodate long code replies).
- Keeper prompt preserves technical assertions and architectural facts while ignoring transient conversational queries.
Why it works this way
Technical tasks require precise code, inspectable receipts, and conservative resource accounting. Calibrating the token burn rate prevents premature exhaustion during long coding sessions.
Evidence
- Mode configuration
engine/presets.py - Technical energy accounting
docs/SESSION_HANDOFF.md - Code generation tests
tests/test_code_in_replies.py
Related: ADVENTURE mode · CONVERSATION mode · CREATIVE mode · allow and deny
presetstuningphysics
Predefined configurations in lore/tuning_presets.json governing voltage, decay, and stress limits.
What it does
Named sets of physical and biological parameters that alter engine dynamics independently of experience mode.
How it works
Presets configure operational parameters across subsystems:
ZEN_GARDEN: Voltage floor 1.0, maximum 25.0; minimal decay rate (0.02); high manic trigger (99.0).THUNDERDOME: High voltage floor (8.0), maximum 30.0; elevated ATP starvation threshold (20.0); manic trigger at 12.0.SANCTUARY: Balanced voltage target (7.0), drag target (2.0), truth target 0.7; metabolism rate 0.5.LABORATORY: Zero biological decay rate, fixed drag floor (2.0), voltage capped at 15.0 for reproducible experiments.
Why it works this way
Grouping related physical parameters into cohesive presets makes it easy to switch the organism between tranquil reflection and high-energy stress testing without hand-editing dozens of JSON keys.
Evidence
- Preset class definitions
engine/presets.py - Tuning preset JSON
lore/tuning_presets.json - Preset unit tests
tests/test_presets.py
Related: preset and tune · ADVENTURE mode · CONVERSATION mode
Configuration, Storage, and Security
Environment variables, configuration files, SQLite persistence, and credential redacting.
Setup wizard and config.json
#
setupconfigwizard
First-boot interactive wizard, configuration structure, and model provider selection.
python main.py
What it does
Guides first-time users through configuring model providers, endpoint URLs, and user names, writing config.json.
How it works
- Executed via
ConfigWizard in mechanics/setup.py when config.json is missing. - Prompts for chat model provider (Ollama, LM Studio, OpenAI, Anthropic), model name, and embedding endpoint.
- Generates
config.json (gitignored, not tracked in repo). - Corrupt configuration files are backed up automatically and the wizard is rerun safely.
Why it works this way
An interactive first-run wizard eliminates manual JSON configuration errors while ensuring user credentials and private endpoints remain outside version control.
Evidence
- Setup wizard
mechanics/setup.py - Configuration loader
engine/presets.py - Setup unit tests
tests/test_main.py
Related: Embedding configuration and environment variables · The Halcyon SQLite store (saves/iris.db)
Embedding configuration and environment variables
#
embeddingsenvironmentconfiguration
Environment variable precedence hierarchy for vector embedding endpoints and models.
What it does
Configures embedding backends, API keys, timeouts, and reprobe intervals via environment variables.
How it works
Settings resolve through a strict precedence hierarchy:
- Rank 1 (Highest): Environment variables (
BONE_EMBED_URL, BONE_EMBED_MODEL, BONE_EMBED_BACKEND, BONE_EMBED_API_KEY). - Rank 2: Configuration entries in
config.json (EMBEDDINGS.URL, EMBEDDINGS.MODEL). - Rank 3: Preset values in
engine/presets.py (BoneConfig.EMBEDDINGS). - Rank 4 (Lowest): Module defaults (
http://127.0.0.1:11434/v1/embeddings, nomic-embed-text).
Why it works this way
Decided in SESSION_HANDOFF.md: environment variables must strictly outrank BoneConfig.EMBEDDINGS. Because BoneConfig ships populated, lower precedence made BONE_EMBED_URL a silent no-op.
Evidence
- Decision record
docs/SESSION_HANDOFF.md - Embeddings resolver
spores/embeddings.py - Embeddings environment tests
tests/test_embeddings.py
Related: Setup wizard and config.json · The Halcyon SQLite store (saves/iris.db)
The Halcyon SQLite store (saves/iris.db)
#
sqlitedatabasestorage
The single authoritative SQLite database housing learned words, memories, audit logs, and checkpoints.
What it does
Consolidates all persistent state into a single SQLite file (saves/iris.db) using WAL mode and transactional safety.
How it works
- Manages tables:
learned_words (dynamically acquired vocabulary), lore_overlay (user additions over factory lore), memories (associative memory vectors), conversations, turns, user_profile, and checkpoint. - All writes use explicit transactions via
Store.transaction(immediate=True). - Upgrades older JSON stores (
cortex_hive.json, legacy.json, quicksave.json) automatically on boot. - Factory files in
lore/ are strictly read-only and never modified.
Why it works this way
Ported from bone-iris (20.7.4.67) to replace fragmented JSON files. A single transactional SQLite file prevents partial write corruption and provides an auditable history of all memory operations.
Evidence
- Store implementation
engine/gate/store.py - Memory database tests
tests/test_halcyon_gate.py - Learned words tests
tests/test_learned_words.py
Related: Substrate file management and secret withholding · Clean resets via reset.sh
Substrate file management and secret withholding
#
securityredactionledger
Isolates model file writes to output/, tracks approvals in a ledger, and redacts sensitive credentials.
What it does
Enforces sandboxing on file operations and prevents private credentials from persisting in memories, logs, or checkpoints.
How it works
Security boundaries are enforced at two critical interfaces:
- File writes: Constrained to
output/. Pre-existing files require /allow approval, logged in output/.substrate_ledger.json. - Secret withholding:
engine/gate/secrets.py scans memories and turns for API keys, private keys, credit cards, and passwords. - Matched secrets are redacted with
[withheld: ...] in checkpoints, crash logs, and audit records. - In ADVENTURE mode, puzzle passwords are part of the story and retained, while keys and tokens remain withheld.
Why it works this way
Local language models frequently reflect back credentials passed in conversation. Automated redaction prevents accidental credential leakage into databases, checkpoints, or telemetry traces.
Evidence
- Secret redaction logic
engine/gate/secrets.py - Substrate tools
mechanics/tools.py - Secrets unit tests
tests/test_secrets.py
Related: The Halcyon SQLite store (saves/iris.db) · allow and deny
Clean resets via reset.sh
#
resetmaintenancecleanup
Scripted cleanup utility that purges memories, saves, and logs while preserving setup configuration.
./reset.sh
What it does
Shell script that removes generated databases (saves/iris.db), log files, temporary dumps, and spores.
How it works
- Deletes
saves/, output/, and *.log files. - Preserves
config.json so the user does not need to re-run the setup wizard. - Kept in sync with
.gitignore to ensure git status remains clean following resets.
Why it works this way
Provides a clean slate for running reproducible evaluation benchmarks and tests without forcing developers to reconfigure their model endpoints and keys.
Evidence
- Reset script
reset.sh - Gitignore file
.gitignore - Handoff reset documentation
docs/SESSION_HANDOFF.md
Related: The Halcyon SQLite store (saves/iris.db) · Offline test suite discipline
Observability, Testing, and Auditing
The receipt ledger roll call, offline test invariants, somatic audit tools, and human evaluation panels.
Subsystem receipts and the core roll call
#
receiptstelemetryroll call
Inspectable work records issued by subsystems to expose silent failures and fallback paths.
What it does
A structured accounting system where components report what work they performed and whether they degraded.
How it works
The receipt system operates under strict reporting rules:
- The core roll call tracks 14 primary subsystems (
composer.compose, cortex.somatic, embeddings.embed_batch, governor.bitmap_gate, halcyon.gate, stage.negotiate). - Each receipt records inputs, result count, and a
degraded flag (ok or DEGRADED). - The component doing the work fills in its own record; callers never infer receipt status from exceptions.
/diag compares the expected roll call against received receipts to detect silent dead code.
Why it works this way
Prose generation looks identical whether underlying physics ran or quietly returned defaults. Receipts make silent failure impossible by requiring every critical component to account for its execution.
Evidence
- Receipt ledger
engine/receipts.py - Receipt audit script
tools/audit_receipts.py - Observability tests
tests/test_observability.py
Related: diag · Offline test suite discipline
Offline test suite discipline
#
testingpytestoffline
The test suite is strictly pinned offline using hash embeddings; live network calls are opt-in only.
What it does
Guarantees that pytest runs entirely offline without requiring a reachable Ollama server or network connection.
How it works
tests/__init__.py sets BONE_EMBED_BACKEND=hash before any test module is imported.- Mock transports simulate HTTP responses in
tests/test_embeddings.py. - Live server tests are guarded behind the opt-in environment variable
BONE_EMBED_LIVE_TEST=1. - Runs the complete test suite (~1,000+ tests) cleanly in
.venv/bin/pytest.
Why it works this way
Decided in SESSION_HANDOFF.md ('The test suite is pinned offline'): unit test passes must reflect deterministic code correctness, not local server uptime or external network latency.
Evidence
- Test suite initialization
tests/__init__.py - Embeddings test suite
tests/test_embeddings.py - Session handoff requirements
docs/SESSION_HANDOFF.md
Related: Subsystem receipts and the core roll call · Somatic audit pipeline
toolsauditpipeline
Specialized audit scripts in tools/ measuring vocabulary coverage, ATP burn rates, and distress inference.
What it does
A collection of headless verification tools in tools/ that measure real engine dynamics across simulated conversations.
How it works
The somatic audit pipeline comprises specialized evaluation tools:
tools/audit_somatic.py: Executes automated multi-turn conversations, auditing vitals, vocabulary resolution, and prompt compliance.tools/audit_somatic_census.py: Runs standardized census topics across models without GPU judge dependencies.tools/audit_receipts.py: Verifies receipt roll-call compliance without booting the full organism.tools/somatic_sim_user.py: Replays simulated flagging and distressed partner dialogues to verify accommodation.
Why it works this way
Subjective prose inspection cannot detect subtle thermodynamic imbalances. Somatic audit tools derive hard numbers for vocabulary resolution, ATP burn, and sentence compliance across hundreds of turns.
Evidence
- Somatic audit script
tools/audit_somatic.py - Testing guide
docs/TESTING.md - Audit somatic tests
tests/test_audit_somatic.py
Related: Offline test suite discipline · Human panel evaluation and blind benchmarking
Human panel evaluation and blind benchmarking
#
evaluationblind panelbenchmarking
Verdicts come from human readers via blind panels, strictly rejecting automated LLM judges.
What it does
Generates randomized, anonymized HTML comparison panels (tools/build_blind_panel.py) for human evaluation.
How it works
tools/build_blind_panel.py compiles multi-arm conversational outputs into an interactive blind review page (tools/blind_panel_template.html).- Human evaluators vote on responsiveness, authenticity, and voice without knowing which arm produced each reply.
- Automated LLM judges are strictly excluded from deciding production verdicts.
Why it works this way
Decided with Gordon on 2026-09-23 (SESSION_HANDOFF.md): across nine judge models and three rubrics, AI judges rewarded verbose, bulleted, sycophantic replies and failed responsive controls. What BoneAmanita is for is the opposite of what LLM judges favor.
Evidence
- Decision record
docs/SESSION_HANDOFF.md - Blind panel builder
tools/build_blind_panel.py - Judge control tests
tests/test_judge_controls.py
Related: Somatic audit pipeline