John's World

John's World: The Long Walk

Every image in this article is generated or article-supporting media from the original Substack post. Synthetic people and worlds here are research artifacts, not documentary claims.

This page mirrors my original Substack article inside EricRhea.com. The original remains available at advisoryhour.substack.com.

John begins to suspect something is off.

I’ve run multiple experiments since the first harness that I described John’s virtual world. If you happened to miss that article, please start there.

image

Eric Rhea[OpenAI's gpt-image-2 has a secret world[I messed around and discovered an entire HOA inspired world inside gpt-image-2…[Read more[24 days ago · Eric Rhea

John’s Late Night Walk

John’s late night walk began like any other-with what I noticed was his cellphone.

John gets a message
John gets a message
John gets a message

John gets a message

John Walks Into the Procedural Night

One of the stranger failure modes starts in the daytime, which feels important. The system does not begin in some obviously broken haunted state. It begins in a garage, or garage-adjacent domestic space, with John doing something perfectly normal: receiving a reminder.

The reminder says to call Marc about the mower.

That is a beautifully mundane sentence. Nobody in science fiction is ever called to destiny by a calendar appointment about lawn equipment, but maybe that is only because science fiction has lacked courage. The real agentic future may not arrive through a glowing portal. It may arrive through a little circle on a reminder app, asking a man named John to call Marc because the mower has become, in some unknowable suburban way, administratively important.

At first, the scene is actually impressive. John taps the reminder. The little completion circle clears. He keeps the phone in his hand and starts walking through the apartment. Then the phone locks itself.

This is the sort of detail that makes these systems feel uncanny in the useful sense. The phone is not merely a rectangle being held by an avatar. It appears to understand that phones go idle. It returns to the home screen. It shows the time as 10:42, which is also funny because I have no idea where 10:42 came from. Somewhere in the stack, a tiny clock goblin woke up, stamped the scene with bureaucratic confidence, and went back to sleep.

This is the kind of thing that keeps pulling me back into these experiments. Not because they are correct, exactly. They are not. Correct is too strong a word. They are doing something more interesting than correctness. They are assembling scraps of worldly behavior into a scene that is almost coherent enough to fool you, and then failing in a way that shows you the seams.

The first … seam? appears when John walks outside.

image
image
image
image
image
image

He is holding the phone, but we can also observe John’s full character body in the stage at the same time. That is new in this particular run. The camera has produced a kind of suburban astral projection: John as first-person phone-holder and John as visible third-person pedestrian. He is both subject and object. He is the actor and the little man in the dollhouse. It is less a continuity error than an ontology error, which is much more expensive and therefore sounds like innovation.

Then the phone disappears from the frame.

John walks down to the street. He turns right. It is still daylight. Nothing about the moment announces catastrophe. There is no monster, no explosion, no red error overlay. Just a man leaving the house after completing a reminder about a mower.

And then he keeps walking.

This is where the whole thing becomes funny in the way only broken simulations can be funny. John walks down the sidewalk and the camera follows. Then he continues walking. Then he continues walking some more. The generation keeps advancing. The day begins to drain out of the sky. Cars turn their headlights on. The pavement grows dark and reflective. The neighborhood stretches forward with the quiet menace of a Windows progress bar.

There is no destination. The endless walk.

This is the part that reminds me of what I wrote about in the Solivane piece, where the real bottleneck in agentic engineering was not output, but review. The system can produce more motion than a human can comfortably interpret. Here, the model is not failing by stopping. It is failing by continuing. It has found a plausible next action and latched onto it like a very diligent intern who has mistaken endurance for purpose. There are tricks to automate review, but the world isn’t ready for them-let’s instead talk about John.

Do you know why John was walking? I do. It’s in the logs.

Walk away from the observer along a wet reflective suburban sidewalk.

That is, apparently, what the logs say John’s goal is.

And to the system’s credit, that is exactly what John does. This is the terrible little joke. The model has not disobeyed. It has obeyed too literally. It has taken the instruction as a fate. John is not going to the store. He is not going to Marc’s house. He is not calling anyone about the mower. He is simply walking away from the observer along a wet reflective suburban sidewalk, because the wet reflective suburban sidewalk has become the entire universe.

This is also why these failures are more instructive than normal bugs. A normal bug says, “Something crashed.” This one says, “Something understood the wrong layer of the world.” In the robotic shell article, I was dealing with the physical version of that problem. The robot could be impressive and still be constrained by batteries, servos, Raspberry Pi behavior, and the brute fact that reality charges by the amp-hour. In the virtual version, the constraints are stranger. The battery does not die. The sidewalk does not run out. The night can keep rendering. John can keep walking forever, because nobody remembered to make the world ask why.

That is the difference between a state machine and a world engine, which I wrote about in the disembodied Claw piece. A state machine knows what happens next. A world engine knows what exists, what changed, what matters, and what should eventually become absurd if repeated too long. A state machine can say: John is walking, therefore next frame John is still walking. A world engine should eventually say: John has walked for three hours because of a mower reminder and this is now either a medical event, a spiritual crisis, or a very local form of exile.

The system needs that second layer.

It needs a sense of proportion. It needs the ability to notice that an action which was reasonable for twelve seconds becomes surreal at twelve minutes and possibly mythological after two hours. It needs to understand that suburban sidewalks are not infinite just because the camera has not found the end of one yet. Otherwise every scene becomes a tiny procedural Discworld turtle, carrying another sidewalk on its back, and John becomes its most faithful pilgrim.

This is also the lesson from the Mazefall experiments, though in a different costume. In Mazefall, when the art changed, the game did not merely become prettier. The readability of the world changed. The player could suddenly understand intention, danger, space, and mood differently. A world is not just assets and motion. It is a contract with the observer about what matters. When that contract breaks, the system can still generate beautiful frames while the meaning quietly falls down an elevator shaft.

Is John, art?

The John walk is beautiful in that sense. It is a clean little parable about agentic systems. The model gets the phone behavior right. It gets the reminder interaction right. It gets the transition from interior to exterior mostly right. It even gets the daylight-to-night progression and car headlights right, which is frankly rude because those are the details that make you want to forgive it.

But it loses the plot.

Not dramatically. Not all at once. It loses the plot by walking past it.

That is the failure mode worth paying attention to. We keep talking about AI systems as if the scary thing is hallucination in the obvious sense: wrong facts, fake citations, invented commands, a confident answer wearing a rented suit. Those matter. But there is another class of failure that feels more important for agents and simulations. It is the failure of continuing a locally plausible behavior after the global meaning has expired.

John walking down the sidewalk is plausible.

John walking down the sidewalk until afternoon becomes evening and evening becomes night, still carrying the ghost of a phone interaction and still obeying a prompt fragment about walking away from the observer, is something else.

It is not nonsense. It is worse than nonsense. It is coherent nonsense. However, I’ve also done the same thing-there can be simple delight in it.

And coherent nonsense is where agentic systems get interesting, because it forces the uncomfortable question: what part of the system is supposed to know when the story has become stupid?

Right now, too often, the answer is still: me.

Which brings us back to review. I can watch John walk. I can inspect the logs. I can notice the duplicate embodiment, the vanishing phone, the orphaned mower reminder, the endless sidewalk, the procedural nightfall. I can laugh at it, because honestly, a man being summoned into eternal suburban exile by a calendar appointment about lawn care is objectively funny.

But I also have to review it.

That means the bottleneck has moved again. The model is no longer merely producing text or images. It is producing little worlds with their own broken causality. The reviewer is not just checking whether a sentence is true. The reviewer has to notice whether the world still makes sense after the system has been allowed to keep going.

That is a much harder job.

And somewhere out there, John is still walking.

image
image
image
image

I can’t wait to share with you about John driving his vehicle around town, which is a real treat.

Thanks for reading! Subscribe for free to receive new posts and support my work.