Notes on production systems
Writing
Teardowns of systems I built, engineer to engineer. Each one leads with the problem, gives one concrete number, and names the cost as well as the win. Dated by when the work happened, newest first.
- Across systemsAugust 2026Three agent platforms: what I'd do differentlyThree things I got right for reasons I couldn't measure, and one I'd rebuild.
- Across systemsJuly 2026What "production-ready" means for an agent systemFour questions. None of them are about the model.
- Builders StudioJune 2026Multi-tenant isolation with Postgres RLS — why single-deployment wonPer-tenant deployments are easier to reason about and multiply every operation you'll ever perform.
- Across systemsJune 20263,072 dimensions vs 384: choosing embedding size by use caseI used both, eight months apart, and the difference wasn't budget.
- Builders StudioMay 202624 tools in one function-calling registry: organizing a namespacePast about a dozen tools, the model's problem stops being capability and starts being selection.
- Builders StudioApril 2026Autoscaling on backlog-per-task instead of CPUQueue depth alone is the wrong signal and CPU is a lagging one. The honest unit is messages per worker.
- Builders StudioApril 2026A 7-stage speaker-identification cascadeOrder the checks by confidence, not by cleverness, and the expensive one runs almost never.
- Builders StudioMarch 2026Cutting LLM pipeline compute ~94% — ECS on EC2 vs FargateIt wasn't a clever model choice. It was noticing that one capacity strategy across three workload shapes means overpaying on at least two.
- Builders StudioFebruary 2026A promotion ladder that never demotesMonotonicity turned a graph traversal into a threshold comparison, and made every downstream cache safe.
- Builders StudioJanuary 20267 queues, 7 DLQs: what failure handling actually looks likeOne queue with a type field is simpler right up until one slow stage starves every other one.
- Sunny LabsDecember 2025Agent memory across sessions: reprocessing vs retention windowsA window forgets by recency. Reprocessing refuses to forget at all. Both are sorting on the wrong axis.
- Sunny LabsNovember 2025Where the latency actually goes in a sub-500ms voice pipelineThree different numbers get rounded to "half a second" — a hard budget, a tuned target, and a median. Only one of them is a promise.
- Cere NetworkOctober 2025Pipelines as infrastructure-as-codeA new pipeline shipped as a stack file and deployed without recompiling any service.
- Cere NetworkOctober 2025Content-hash idempotency — making an agent pipeline replayableMost LLM pipelines can't be re-run. One decision at the storage layer made ours fully replayable and turned Postgres into a disposable cache.
- Cere NetworkSeptember 2025Exposing pipeline queries as MCP tools automaticallyThe integration you don't write is the one that can't drift.
- Cere NetworkAugust 2025Fan-out reads, single-writer stateTen agents processing the same batch in parallel. The pattern that keeps them from corrupting shared state is older than any of this.