
Long-form takes on applied AI, eval-driven development, multi-LLM strategy, and what production AI actually looks like once the demo is over.
The skeptic-but-pragmatic stance, in long form. How I run the discovery audit, what questions actually matter, and the heuristic for "is this an AI problem or a process problem?"
A concrete walkthrough of building eval harnesses for LLM apps — synthetic test sets, golden answers, regression checks, and the boring discipline that separates "it worked once on stage" from "it works on Tuesday at 3 AM."
Why single-provider AI is a strategic risk, what production multi-LLM routing actually looks like, and the architecture patterns that let you swap providers without rewriting your application.
FIRST POSTS DROPPING SOON
Want them in your inbox? Reach out →