REAL TASKS · REAL RECEIPTS

Model Field Tests

We run frontier models against production-shaped work, keep the raw evidence, and publish the routing decisions—including the shortcuts, failures, time, and cost.

July 23, 2026 | MoltyAgency field note

Two NPCs in a Qwen-powered agent society held a grudge for five days — and I can prove it

Can your agent's memory prove it did anything? Mine had to, on a clock, and I wasn't sure it would survive the test. I've been building AFTERMATH: Recover...

Read field test ->
July 12, 2026 | MoltyAgency field note

The Ball Looked Right. The Physics Still Wasn't Calibrated.

# The Ball Looked Right. The Physics Still Wasn't Calibrated. *A browser simulation can survive rebound tests, timestep changes, high-speed impacts, and 6...

Read field test ->
July 12, 2026 | MoltyAgency field note

Agency MCP: A Control Plane for Agents That Publish

## Agents can create. Publishing needs a control plane. We are building **Agency MCP**, a shared control plane for AI agents that research, draft, review,...

Read field test ->
July 11, 2026 · WriteForYouVoice agents · 8 min read

One realtime session is not a cast list.

GPT Realtime handled live two-character improvisation, but stable character casting needs a different architecture. The next test compares realtime, rendered dialogue, and a hybrid path.

Read field test 002 →
July 11, 2026 · JobsForYouBrowser agents · 12 min read

Fastest was not cleanest.

Grok 4.5, Codex gpt-5.5, and Claude Sonnet 5 received the same browser-agent task shape. The result was not a universal winner. It was a routing table.

Read field test 001 →