
7 Checks Before You Trust an LLM Planner Experiment
AI assistance disclosure: I designed and directed the experiments described here. I used AI coding...

Author profile
Building AI games and arenas. I write field notes on LLM behavior, harness failures, and model evaluation.
Browse the latest writing surfaced through DevArt.

AI assistance disclosure: I designed and directed the experiments described here. I used AI coding...

One authoritative engine, two seat-locked MCP servers, three best-of-threes, and a 3-millisecond...

A small action-schema change cut guaranteed-loss calls without making the model generally...

This is a development note. The project and test data are real; AI helped me clean up the prose....

I nearly turned four software bugs into four model personalities This article was edited...

A hands-on test of V4-Pro's default reasoning mode ⚠️ This article was generated with AI...
Advertisement