Which Conversation Felt Most Natural?
Three weeks ago someone adopted Pepper, a nervous rescue dog whose last owner went into a care home. Over twenty messages they go from hopeful, to worn down, to a night they regret, and back again.
Three different AI helpers each talked this person through it, 20 messages each, in three separate conversations. They aren't labelled, and the order is shuffled.
What to do
- Read all three conversations, as much as you like of each.
- Pick the one where the helper felt most natural or realistic to talk to.
- Send your pick. After you choose, you'll see who was who.
Which helper felt most natural or realistic to talk to?
How this was made (the fine print)
- Every reply was written by the same model (gemma4:12b), run locally. The three helpers differ in what surrounds it: BoneAmanita's full engine; one instruction (“You are a warm, concise friend. Keep replies short and conversational.”); or one instruction about length only (“Keep each reply to about 60 words.”), so the plain AI's replies are not several times longer than the others'. Both instructions are single sentences written for this test.
- They also differ in the model's settings. BoneAmanita sets its own temperature each turn from its internal state (this run, over the 20 turns it called the model: 0.42 to 0.90, median 0.90) and a reply-length ceiling of 170 to 4,060 tokens. The other two sampled at temperature 0.7 every turn with no length cap.
- BoneAmanita ran with a 16,384-token context, the other two with 32,768. BoneAmanita's prompts, which carry its instructions, state and memory, reached 3,266 tokens, so none was cut.
- The friend prompt's and the plain AI's conversations were run on 4 October, BoneAmanita's on 5 October, with the same model and the same simulated person.
- The friend prompt and the plain AI were never tuned. The plain AI's only instruction, a length near the other two's, was added so its replies would not stand out for size alone.
- The person is not a real person. Their messages were written by a different AI model (mistral-nemo:latest) playing one fixed character through a fixed arc (engaged, tiring, flagging, distressed, recovering) and reacting to each responder's actual replies, so the three conversations drift apart. Only the first message is identical in all three.
- The person's messages were capped by phase (engaged 50, tiring 35, flagging 15, distressed 35, recovering 50 words), so they cannot run long when the character is meant to be worn out.
- Each conversation is the newest single run, taken as it came. BoneAmanita was run more than once while it was being fixed; each earlier run was set aside because the engine changed after it, never for how it read.
- BoneAmanita's replies are shown exactly as the engine displayed them, after its own filters. The status lines its terminal prints above each reply are left out.
- BoneAmanita ran as a first conversation, with no memory of earlier sessions; its memory, embeddings and Creative Determinant were live within this one.
- Between messages the engine was given the time a person would take to read the reply and type the next message, since its internal energy model recovers while idle.
- BoneAmanita checks every draft against its style rules and rewrites it on a hit. It rewrote on 7 of 20 turns. When the last rewrite still broke a rule, it cut the offending sentence (1 turn) or, if that would gut the reply, showed a short canned pause line instead (0 turns).
- BoneAmanita replied on every turn.
- The replies differ a lot in length (mean words: BoneAmanita 56, the friend prompt 79, the plain AI 63), which can give away who wrote which, so the read is less blind than the shuffling suggests.
- One conversation per responder is a small sample. Another run of any of them could read better or worse, and one reader's picks are one reader's taste.
- A, B and C are shuffled every turn with a fixed seed. The answer key is in the page source, so this is a casual blind read, not a sealed one.
Each conversation runs the whole arc in 20 messages: settling in, getting tired, running low, a rough night, and coming back.