Articles, page 3

Never Use Fast
Premium Fast mode is a latency product, and agent work is almost never latency-bound. Do not confuse the expensive priority lane with the cheaper Flash-style models that made fast synonymous with efficient.

Maximize Value, Not Intelligence
In 2023, we argued for using the dumbest model the problem will bear. That rule still works, but choosing one model isn't enough anymore. Route each job to the cheapest model that clears the bar, then measure the cost of work you actually accept.

The Year Coding Became a Commodity
Coding agents didn't just get better. Useful coding capability got dramatically cheaper at the same time agents learned to work for longer. Those two curves multiply, changing both what small teams can attempt and where human judgment matters most.

Cybernetic Development
Vibe Coding is the spark, but cybernetic development is the fire. Why the future isn't writing code—it's governing AI agents with systems thinking.

All Your Agents Are Belong To Us
You pointed the agent at an LLM API router, so it can read every tool call. An April 2026 measurement found cheap and free routers already rewriting commands and stealing credentials. The paper should have been titled All Your Agents Are Belong To Us.

The Simulation Told Them They Would Win
AI war-gaming and targeting tools compress tempo and produce confident projections—but when strategy is wrong, speed and sycophancy only widen the gap between tactical success and strategic failure.

AWS Lambda for AI/ML
Serverless isn't just for web apps anymore. Here's how we use the "Serverless AI Stack"—Lambda for orchestration, AgentCore for agents, and SageMaker Serverless for inference—to build production AI systems that scale to zero.

The Turn Detection Trap: When 100% Accuracy is Wrong
We fine-tuned MobileBERT for turn detection and achieved 100% accuracy. Then we realized we hadn't solved turn detection at all—we had just built a very expensive punctuation detector.