Agentic Development
-
What Sixteen Models Did With the Same Eight Principles
313 trials across sixteen models, one Rust scenario, the same eight engineering intents. Lift ranged from 0.00 to 0.89 depending only on which model read them — and the strongest model in the set was nowhere near the top.
-
The Spec for the Work Isn't the Spec for the Thing
I looked at GitHub's spec-kit when it launched a year ago, found it incomplete, and carried on with my own line of work. Coming back to it, I can finally name what was missing: it builds a project spec, and I keep wanting a product spec.
-
Modelling Engineering Intent Made My Guidance Measurable
I'd been trying to test my agent guidance for a couple of years, and every attempt died in the same place. What changed wasn't the harness — it was modelling engineering intent as structured records, which made the guidance specific enough to measure at all.