I build and evaluate AI to make products work.
GG Tank Watch is a civic emergency archive I built during a real May 2026 chemical-tank evacuation (~50,000 residents; Wikipedia, NPR). A consumer-facing AI system held inside its authority by code and tests, not prompting:
- Bounded authority, in code. The model never writes the published snapshot. Its output clears one validation chokepoint it can't bypass, so the system cannot exceed "route to officials."
- The asymmetry that matters. A false all-clear is catastrophic; a false alarm is survivable. So danger downgrades need ≥2 sources (including an official agency), while upgrades fire on one. Enforced in code, never asked of a model.
- A behavioral harness of 200+ tests — green in CI — catches drift from the safety contract (fabricated sources, authored directives, stale data) before it ships, not after.
Three essays on AI evaluation, from practice — mikeilog.com/writing:
- What building an emergency dashboard taught me about trustworthy AI — the GG Tank Watch story: five lessons about AI you can trust.
- What AAA QA teaches about AI evaluation — Square Enix QA as a multi-agent evaluation system before anyone called it that.
- Human-in-the-loop eval and the part you can't ask about — what a real human-in-the-loop looks like, and when the honest move is to cut the feature.
- gstack
- Claude Code
- Reported and diagnosed session-bloat bug #61613 in the Read tool's binary-file serialization path (root cause + fix sketch)
- Added a reproduction, render-only proof, and fix direction to scrollback-duplication #51828
- Building AI agents: Claude Agent SDK, MCP, Anthropic SDK
- Languages: Python, TypeScript
I pick up new domains fast; everything is language.
- LinkedIn → https://www.linkedin.com/in/mikeilog
- Email → mike@mikeilog.com
cooperation FTW · US (Pacific time) · remote




