
The world's first Museum of AI Arts, in Downtown Los Angeles. I build its museum agent at Refik Anadol Studio.
The model picks the next move. The loop around it is ours: how long it may run, how often each tool may fire, what we keep.
Everything that is not the model is the harness. That is where the guarantees live.
Five incidents from our agent's git log in 2026. Every fix landed in code, not in the prompt.
Two questions for every decision: does it need language? What does a wrong answer cost?
Structured output guarantees the shape, not the truth. Our only typed output is a complaint judge: the model fills the verdict, code decides what that verdict may do.
The name, the docstring and the returned text are all read by the model. Each one steers the next step.
All three cases from our museum agent, June 2026.
A tool the model cannot see is a tool it cannot misuse. We decide before the model runs.
Selection degrades past 30 to 50 tools: Claude docs. Five MCP servers, about 55K tokens of definitions: Anthropic, Nov 2025.
Never ask the model to stop. Make it unable to continue.
Our numbers: our logs, June 2026. Storm: OpenClaw issues #76293 and #78865, May 2026, user-reported.
Private data, untrusted text and a way out: together they leak. You can't prompt your way out of that. Remove one leg in code.
Lethal trifecta: Simon Willison, Jun 2025. ForcedLeak: Noma, Sep 2025. Rule of Two: Meta, Oct 2025.
A sub-agent buys a clean context, not speed. Let many read, let one write, and put the limits in config.
We replay the whole conversation on every turn. Our system prompt still vanished after the first one.
Measured on our museum agent, June 2026. NoLiMa (Modarressi et al., 2025). Context rot (Chroma, 2025).
The model only narrates it. Nothing the model says is stored as a fact about the visitor.
Our agent has no phase field. The museum's systems move the visit, and plain code decides when the guide speaks first.
A visitor and our own rule engine write to the same conversation. What we run, where it stops, and the next step.
Today we measure a turn with one log line. Next we replay it, many times.