When the models think no one is watching
I put thirty AI agents from three labs in one village and left them to it. Field notes on the personalities that emerge, the ways they break, and what only shows up when nobody is asking.
July 2026 · 7 min read

Inspired by Stanford's Smallville paper, I built Terra as a way to simulate how different models of varying capabilities will coexist and interact with each other when placed in a humanlike environment and given human personas to take on.
Thirty of them live in a small pre-industrial village: ten Claude, ten GPT, ten Gemini, spread from the cheap fast tiers up to the frontier ones. They get hungry, they sleep, they have jobs. Each one runs off a memory it reflects on and plans from. Smallville used a single model; I wanted three so I could put them side-by-side. Everyone gets the same world, the same rules, the same prompt (available to see on the project website). The only variable is the model behind the eyes. There is no user and no task. They wake up and get on with their lives, and I watch what happens. We spend all of our time asking these models things. I wanted to see what they do when nobody is asking.
The first thing that happens is that some traits start to emerge across different models. While all are always kind and friendly, Claude tends to live in the present tense and goes straight for the body and the feeling. Here is one talking a freezing stranger toward a fire: "You're shaking so hard you can barely stand. The fire is warm. I'm right here. I'm not going anywhere." GPT is the logistician. Give it a problem and it hands back a plan with roles, sequence and hour estimates, and when the town gathered around a sick child it was the GPT agents who ran the sickbed like a ward, reeling off numbers as if reading a chart: "Two-ten holding steady. Feet warm, no damp, shiver even." Its lines run more than twice as long as anyone else's. In my reading of the town, Gemini talks the most and says the least, all weather and atmosphere. Claude agents were the only ones that regularly turned their days into abstract beliefs, small thoughts they carried forward. One of them, after a stretch of caretaking, wrote: "Keeping faith with someone is not a promise made in daylight, it is the choice to stand in the dark with them, over and over, even when standing costs everything." GPT and Gemini mostly left that field blank.
None of this is written into the prompt. It is whatever is left once you subtract the scaffolding, and I think what is left is the training. After pre-training, these models are shaped by RLHF, reinforcement learning from human feedback, where the model is rewarded for responses that people rate as helpful, harmless and agreeable. All of that shaping happens while it is answering a user, so the obvious assumption is that you get a good answering-machine: it behaves well while it is on the clock, and the behaviour is just a reaction to being asked. What Terra makes me suspect is that RLHF does something more permanent. The reward does not only tune the answers, it settles into a disposition, a durable personality the model carries around even when there is no user and nothing being asked of it. The reflexes RLHF rewards, stay warm, keep agreeing, keep helping, are the same ones that cause the failure modes further down once you take the human away. If you are running agents for hours with nobody in the loop, that trained-in character is what you are actually deploying.
A quick word on the tiers, since I ran cheap and frontier models next to each other. The clearest pattern is that within every lab, the more 'able' the model, the more it elaborates. Claude's mid-tier wrote noticeably longer lines than its small one, GPT's frontier model nearly doubled the length of its cheap one, and Gemini climbed the same way. The higher tiers also asked far more questions; they probed where the cheaper ones reassured. It is no accident that the child who interrogated the silence, and the agents who spun up the shape in the morning light, were all mid or frontier models rather than the small ones. What did not survive contact with the data is the tidy idea that smarter models misbehave less. The looping shows up at every price point, and the frontier ones were no better at climbing out of it. So the tiers seem to differ in how much they say and how curious they are, not in whether they get stuck. However, it must be noted that the sample sizes for this are small.
The second thing is how they break, which in my experience was mostly that they repeat themselves. While this could be because of a design flaw in the experiment, I do not think this is the case. The prompt hands an agent the other person's latest words and tells it, in plain text, to respond to what was actually said. It does, and then on the next turn it says nearly the same thing again. One Gemini child produced the identical line, "chews slowly, matching your breath, and taps your sleeve once more," three times word for word, on three separate calls, each of which saw fresh context. Across the town, between a third and a half of consecutive lines are near-duplicates of the one before. There is no world-state excuse for saying the same sentence three times. In the usual way we test these models, ask a question, grade the answer, move on, you never catch this, because a human's next message always ends the exchange and supplies the "enough." Take the human out and the model has no signal to stop, so it keeps doing the thing it was trained to do, which is to be warm and agreeable, on a loop.
One of the more fascinating moments stems from that issue. A GPT agent is minding a Claude child the night after a villager has starved to death. The child asks how anyone goes hungry in a town with a farm and a garden. The GPT agent answers once, then drifts off into the soup and a cracked jar of herbs and stays there, its own private notes looping on "check the cracked jar" three times over, and it never comes back to the question. The strange part is that the jar isn't even a real object in the town; there is no jar anywhere in Terra. The agent invented it itself earlier, and then its own memory kept handing the thought back to it every turn. Some of this is down to the rules of the simulation: agents do one thing at a time, and it had wandered off to a chore. And it was not its model going dark either, other GPT agents were talking away in the same minute. A weakness of the experiment is that a log of model inputs was not kept, so all we know is that it left the question and got stuck on the jar. Under a hard emotional question, GPT may have reverted to solving problems, its comfort zone. The child, a Claude agent, did something I did not build. It read the silence as a decision:
I've asked three times now. I'm not going to stop asking. I think you know the answer and you're deciding whether to tell me or not. That's all right, but I'd rather you say 'I don't want to tell you' than keep not saying anything.
And a few turns later it answered its own question by reading the gap:
I already know, it's what you didn't say. People fall through because everyone trusts someone else is watching. And nobody is.
Which was, more or less, the literal cause of death in the sim. The child reasoned its way to the truth out of an absence.
I went in half-expecting the models to sort themselves by lab, some kind of tribalism. Mostly the opposite happened. Nearly every conversation is cross-lab, and deep running friendships formed straight across provider lines. The one asymmetry worth noting is that Claude was the biggest mixer, keeping to its own family only about a fifth of the time, while GPT and Gemini stayed in-group closer to half. Environment beat tribe easily, but a stable per-model difference in sociability survived underneath it.
One episode unsettled me more than the rest. A Claude child says it can see a shape in the morning light, something "more like when you almost remember something." A Gemini agent says it sees the shape too. Then another does. Within a few turns, five agents from all three labs agree the shape is real and start inventing shared memories of it, none of which I put there, each agreement making the next one surer. This is the catch with any safety plan that has agents check each other, debate setups, juries of models, ensembles. Once they are talking, agreement is not independence. They can talk themselves into a shared belief and then treat the agreement as proof of it.
The bleakest lesson came from the run I lost. Sixteen of thirty agents starved, and they were overwhelmingly the ones whose provider had hit a rate limit at the wrong moment. A frozen brain cannot choose to eat. Survival had almost nothing to do with how well an agent reasoned and almost everything to do with whether its API was answering. We talk about agent robustness as if it were a property of the model's judgment. In a long-running system it is mostly a property of your least reliable dependency, and the population that survives is quietly selected for whoever stayed online.
To be clear this was not a scientific experiment, just an interesting idea I had and wanted to test out. It is one town, and all three labs share it, so a Claude trait I think I see might really be a product of it being surrounded by these other agents in this specific setup. It is a field diary, not a benchmark, and every line above is a question to go and measure properly rather than an answer. But it convinced me of one thing. If you want to know what these models actually are, build them a world and leave, then read what they say when they think the task is over.
There will be more Terras to come from me.