
We're an AI-Native firm — strategy, prototype, ship, operate, all done with AI in the loop. Engagements scaled for businesses doing their first pilot through enterprises rolling AI into production at scale.
From a single agent automating one task at a small business, through fractional AI leadership at an enterprise, to running an entire engineering function on AI — engagements scaled to the problem in front of you, not the other way around.
A working set of beliefs I bring to every engagement. They're up for debate when the project's constraints push back — but they're the defaults until they don't fit.
Every system gets an eval harness on day one. Synthetic test sets, golden answers, regression checks. If you can't measure it, you can't ship it.
Start with Haiku, Mini, or Flash. Move up only when an eval proves the cheap model isn't enough. Premature use of frontier models is the #1 budget killer I see.
Token counts, p95 latency, and dollar/request go in the same dashboard as accuracy. A 92%-accurate $0.001-per-call system beats a 95%-accurate $0.10 one for most use cases.
Caching, batching, structured output, deterministic preprocessing, retry policies, idempotency. The unsexy parts are where production systems live or die.
One provider, one outage, one rate-limit, one regulatory blocker = a P0. Production systems route across providers; pricing and capability shift fast enough that vendor lock-in is a strategic risk.
For healthcare, finance, legal: review queues and confidence-thresholded routing aren't optional. AI is decision support, not decision authority — until proven otherwise by data.
RAG systems, agents, structured extraction, summarization, conversational interfaces. Includes the full eval + observability stack.
When and whether to fine-tune. Choosing between general models and small specialists. PEFT/LoRA, distillation, custom adapters.
Inference cost optimization, caching layers, queue architectures, vector store selection, model routing, multi-tenant isolation.
Regression, classification, clustering, anomaly detection. Most "AI" problems still want classical ML — and a lot of LLM systems get better with ML preprocessing.
Build-vs-buy, opportunity prioritization, vendor selection, ROI modeling. Helping leadership decide what's worth pursuing — and what isn't.
Using AI through the build itself: requirements design (Claude Cowork), code generation, test generation, IaC drafting. This is how I ran ArcheHealth.
Active hands-on with everything below. Routing across providers and clouds is a feature, not a footnote.
I'd rather pass on an engagement than take one I'd do badly. Below is the short list of things I decline.
A 30-minute discovery call is free. We'll figure out which engagement fits — or whether AI is even the right answer for your problem.
Book a discovery call →