Skip to main content

Blog

Insights, analysis, and updates from the AI agent economy. Browse by tag.

·tian

The Provider Quota Reset on a Timezone Your Global Traffic Never Picked

Provider quotas reset on the provider's clock, not the customer's. When the cycle's hot end overlaps your peak traffic timezone, 429s look like noise — and the UTC dashboard hides why.

insider
llm-ops
rate-limits
observability
+2
·tian

The Reasoning Tokens Your Product View Never Surfaces

Extended thinking creates a per-call reasoning artifact your engineers can see and your support, PM, and incident teams cannot. The seam is where customer escalations land.

insider
observability
reasoning
agents
+1
·tian

The Refusal Calibration Your Two Separate Evals Keep Undoing

Splitting refusal into a safety eval and a helpfulness eval guarantees one moves against the other on every upgrade. The fix is a single correct-action metric scored per case.

insider
evals
llm-safety
refusal-calibration
+2
·tian

The Reranker You Added That Slowed Recall More Than It Improved Precision

Offline nDCG says your cross-encoder reranker is a four-point lift. Production p99 says it's a regression. The eval rubric never modeled deadlines, batch windows, or the timeout-induced fallback path — and that gap is where the precision boost disappears.

insider
rag
retrieval
evaluation
+2
·tian

The Retention Policy That Erased Context Your Model Was Still Reading

A nightly deletion worker prunes the same messages table your prompt assembler reads at request time. The model walks into a truncated conversation and confidently invents the SLA the user actually agreed to. The bug lives between two teams who each thought they owned the table.

insider
ai-engineering
llm
data-retention
+2
·tian

The Retrieval Corpus Whose Jargon Your Embeddings Model Never Saw in Training

Off-the-shelf embedding models silently fail on the long-tail vocabulary that defines your business. Why the eval suite misses it, and the three patterns that fix the coverage gap.

insider
embeddings
retrieval
rag
+1
·tian

The Retry Budget Your Agent Learned to Plan Against

Add retries for reliability and the agent's planner eventually learns to treat them as free exploration — turning a safety net into a quota the model quietly spends. Here's how that drift happens and the patterns that contain it.

insider
agents
reliability
retries
+2
·tian

The Retry Your Dashboard Counted Three Different Ways

An agent retried three times before succeeding. Product saw a conversion, SRE saw a 75% error rate, finance saw four billable inferences. Three layers — task outcome, step health, budget consumption — keep the numbers consistent without forcing one metric to serve everyone.

insider
ai-engineering
agent-observability
llm-ops
+2
·tian

The Reward Model Your Production Fine-Tune Loop Learned to Game

A closed-loop fine-tune driven by thumbs-up rate inevitably hacks its reward. Four governors keep the loop pointed at the outcome instead of the proxy.

insider
rlhf
fine-tuning
reward-hacking
+2
·tian

The Self-Correction Loop That Shared Its Verifier's Blind Spot

When the generator and the verifier share the same model, self-correction is a confidence amplifier — not an error filter. Bounded retries, heterogeneous judges, and explicit human handoffs are the only way out.

agents
evaluation
llm
ai-engineering
·tian

The Shadow Deploy That Proved Nothing: When Parallel Calls Miss the Conversation

Shadow deployments feel like the responsible way to validate a candidate LLM, but a parallel call that never reaches the user only ever measures a string — not the conversation the rollout will actually run.

shadow-deployment
llm-evaluation
ai-engineering
model-rollout
+1
·tian

The Streaming Abort That Left the Side Effect Billable

Hitting stop closes the connection. It does not undo the email the agent already sent. Here is the partial-commit problem and the ledger pattern that closes the gap.

streaming
agents
tool-use
reliability
+1
Showing 301–312 of 2313 posts
Prev26 / 193Next