Skip to main content
Tian Pan

Tian Pan

Software Engineer

View all authors

·tian

What You Deleted Is Invisible to Your Coding Agent

Coding agents reintroduce code you deleted yesterday because absence leaves no trace in the repo. A field guide to recording the negative decisions agents need to respect.

insider
coding-agents
ai-engineering
developer-experience
+1
·tian

The Nightly Batch Job That Quietly Became a Latency-Critical Service

A nightly batch job becomes a latency-critical service one reasonable request at a time. Why batch and online inference optimize for opposite goals, how the drift produces quiet failures, and how to re-architect on purpose.

insider
ai-engineering
system-design
mlops
+2
·tian

The Feature Store Your Agent Reinvented Badly

AI agents re-derive the same facts every turn — churn risk, account age, plan tier — with no caching, no shared definitions, and no point-in-time correctness. Why that makes them a broken feature pipeline, and how to fix it.

ai-agents
feature-store
mlops
data-engineering
+1
·tian

Provider Rate Limits Are a Capacity Plan You Never Wrote

When your app hits a 429, the retry code that runs next quietly becomes your capacity policy. Treat rate-limit handling as deliberate load shedding — with priority tiers, jitter, and a scheduler — instead of a library default nobody reviewed.

rate-limiting
llm-api
capacity-planning
load-shedding
+1
·tian

Your Happy Path Is Your Expensive Path: The Agent That Costs More When It Wins

A failed agent run is cheap; a successful one can cost 50x more. Why raising your agent's success rate compresses margin, and the levers that fix it.

ai-agents
unit-economics
finops
llm-cost
+1
·tian

Your Agent Endpoint Is a Distributed System Pretending to Be a Function Call

An `await agent.run()` looks like a local function but hides a remote, partially-failing distributed system. Here is the timeout, retry, idempotency, and circuit-breaker discipline agent code needs.

ai-agents
distributed-systems
reliability
llm
+1
·tian

Your Agent Has No Concept of Business Hours

AI agents act the instant they decide, with no sense of whether 3 a.m. is the wrong time to send it. How to separate anytime work from daylight work and build a timing layer that knows when to wait.

insider
ai-agents
agent-design
reliability
+1
·tian

The Backfill Problem: Why Agent Memory Needs Migrations Like a Database

Agent memory is a production database that drifts every time you improve its format. Version your records, write real migrations, and backfill before old memories quietly rot.

agent-memory
schema-migration
data-engineering
llm-agents
+1
·tian

The Bug You Can't Reproduce Because the Model Picked a Different Token

Replaying an LLM bug and watching it pass doesn't mean the bug is gone — it means you drew a different sample. How to debug a sampler when your tools assume determinism.

insider
llm
debugging
observability
+1
·tian

Build vs. Buy Is the Wrong Question for Your AI Feature

An AI feature is a stack of five layers, not one thing to make or purchase. The decision that matters is which layer compounds your differentiation and which one a competitor can simply buy.

insider
ai-engineering
architecture
strategy
+1
·tian

Capacity Planning When Every Request Thinks a Different Amount

Agent requests don't have a stable cost — one resolves in 200 tokens, the next burns a million. Why p50 forecasts fail for agent workloads, and how to plan in token and tool-call distributions instead.

ai-agents
capacity-planning
llm-infrastructure
observability
·tian

The Carbon Line Item Nobody Puts in the AI Feature Spec

Every AI feature review argues latency, token cost, and accuracy — and never energy. Here is how to measure per-request carbon and make it a number a team owns.

ai-engineering
sustainability
observability
infrastructure
Showing 433–444 of 2313 posts
Prev37 / 193Next