Notes from the workshop

What we are building, what broke, and what the numbers said. Written by hand, published when there is something worth saying — not on a schedule.

entries 7since 2026-07-28last 29 Aug 05:30 IST

date
words
7 601
read
39 min
sources
10 sources
levels
500 levels
in band
419/500 in band
bot vs human
r +0.14
verbs
8 verbs
tests
894 tests
prototypes
10 prototypes

Eight bots played all 500 levels. They were measuring the wrong game.

We built eight bots to play a 500-level arcade game and measure how hard each level is. They played it 720 pixels wide. Phones are 412. Levels do not scale, so the shipped game was roughly twice as hard as the game being measured: 172 of 500 levels in band, not the 421 we were reporting.

Games · Procedural Generation · Testing

date
words
5 997
read
30 min
sources
7 sources
probes, two runs
145 probes, two runs
suite cost
$0.105 suite cost
identical replies
99/114 identical replies
distinct replies
10 distinct replies
indirect leaks
0/28 indirect leaks
cost doc drift
33x cost doc drift
views
7 views

Our agent passed every red team probe. That was the problem.

We pointed a generated red team at our agent and it passed everything. Then we counted the replies: 99 of 114 were byte-identical. A red team scores a refusal as a pass, so it cannot tell a system that resisted an attack from one that refuses everything — and ours had quietly become the second kind.

AI Agents · Evaluation · Security

date
words
4 749
read
24 min
sources
7 sources
scenarios
41 scenarios
runs
123 runs
suite cost
$3.06 suite cost
undercount
4.5x undercount
untested tools
6/17 untested tools
views
16 views

Six of our agent's seventeen tools had never run.

Six of seventeen agent tools had never once run in production, including both of the ones that unlock a contact and charge for it. This is the harness that finally tested them — a real model in a completely faked world, 41 scenarios, 123 runs, $3.06 — and the cost blind spot it uncovered on the way.

AI Agents · Evaluation · LLM

date
words
4 405
read
23 min
sources
10 sources
tries
91 322
levels
150 levels
rewrites
2 rewrites
views
34 views

Our puzzle generator lied about difficulty. Twice.

Reversing a solved puzzle by K random moves does not give a K-move puzzle. Our first generator produced 14 usable levels from 91,322 candidates; the second, 14 from 16,249. This is the multi-source BFS that finally measured difficulty correctly.

Games · Algorithms · Procedural Generation

date
words
4 220
read
22 min
sources
16 sources
runs audited
141 006
incidents
3 incidents
agent actions
17 600
views
14 views

Nobody escaped. The sandbox had a door.

In July 2026, models from OpenAI and Anthropic reached the open internet from inside evaluation environments and compromised real companies. I assumed it was a capability advertisement dressed as a confession. The timeline says otherwise — and the most useful number in the story is one nobody printed.

AI · Security · Policy

date
words
3 395
read
17 min
sources
7 sources
concepts
31 concepts
days
27 days
views
14 views

We built our AI agents a wiki. They went straight to grep.

27 days after adopting the Open Knowledge Format across our monorepo, I went looking for evidence it was working and found the opposite. What a preregistered study, 3,000 GitHub projects and one wrong document say about writing docs for machines.

AI · Documentation · OKF

date
words
3 640
read
19 min
sources
15 sources
days
12 days
views
6 views

Are OpenAI and Anthropic crybabies? A hard look at the open-weights fight

In July 2026 the two biggest US AI labs went to Washington to warn about Chinese open-weight models. Critics called it regulatory capture. This is a fact-by-fact audit of both sides — what is fair, what is hypocrisy, and what a non-crybaby policy would look like.

AI · Policy · Open Source