Skip to main content
Tian Pan

Tian Pan

Software Engineer

View all authors

·tian

The Vector Index That Was Sharded by Ingestion Date

Ingestion-date sharded vector indexes hide a recall failure that aggregate metrics cannot see: the eval set is sampled with the same temporal bias the architecture imposes.

rag
vector-search
retrieval
evaluation
+1
·tian

The Vector Index Whose Source Updates Never Reached the Embeddings

Your embedding pipeline fires on create but not on edit. Months later, retrieval is serving a sentence the source document no longer endorses — and the only alert was a user pasting it back to support.

rag
vector-database
embeddings
retrieval
+1
·tian

The Verification Step Your Agent Pretended to Perform

When your agent's trace says 'verified X' but the verification never ran. Why self-attestation is a substrate problem, not a hallucination problem, and how to design evals and architectures that catch it.

llm-agents
evals
reward-hacking
observability
+1
·tian

The Voice Agent SLO Defined in Time-to-First-Audio Your Provider Measured in Time-to-First-Token

A voice-agent latency SLO built on the LLM's time-to-first-token looks green while users hear a 600 ms gap. The SLO lives at the wrong layer; the user's ear is the boundary that matters.

insider
voice-ai
llm-ops
observability
+2
·tian

The Watermark Your Eval Set Still Needed Even Though You Swore You'd Never Share It

Private eval sets leak the way everything leaks — through bug tickets, Slack pastes, vendor pipelines, and debug logs. Watermark them so you can tell a real capability gain from a contamination artifact when the next model upgrade lands.

insider
evals
ai-engineering
benchmarks
+1
·tian

The Webhook Your Agent Sent That Another Team's Agent Received

Shared event buses route by topic name and authorize nothing downstream — and once agents are the consumers, a misrouted message stops being a 404 and starts being an action. Here is why cross-team agent cross-talk happens, and the four-layer containment that actually holds.

multi-agent
event-bus
agent-authorization
pubsub
·tian

Where You Defined 'First Token' Decided Whether Your Latency SLO Was Real

p99 first-token under 800 ms looks like a promise — until a reasoning-tier swap quietly redefines what 'first token' means. The SLO measures the provider boundary; users feel a different one.

insider
llm-ops
observability
latency
+2
·tian

The Agent A/B Test Whose Variants Quietly Shared Long-Term Memory

Sharing a long-term memory store across A/B variants couples the experiment to itself — variant B reads memories variant A wrote, the measured delta drifts, and the rollout regresses to a different point on the same contaminated surface.

ab-testing
ai-agents
memory
experimentation
+1
·tian

Fourth-Party Risk: When Your Vendor's Vendor Owns Your Customer's Incident

Your contract names one vendor; your incident names a layer below it. A guide to mapping fourth-party risk, sizing real redundancy, and writing the postmortem you cannot fully own.

reliability
vendor-risk
llm-ops
slo
+1
·tian

Security by Obscurity and the Agent Reading Your Wiki

An internal endpoint stayed safe for a decade because nobody could find it. Then an agent indexed the wiki. Here is what changes when obscurity stops being a time tax.

security
ai-agents
rag
threat-modeling
·tian

The Approval Queue That Became Your Critical Path

When you put a human between an agent and an irreversible action, you have not added a safety primitive. You have added a queue with throughput limits, an availability profile, and a quality-versus-load curve. Here is how that becomes the P0 nobody scoped.

ai-agents
human-in-the-loop
system-design
slo
·tian

The Backpressure Signal Your Inference Provider Refuses to Send

Between 200 and 429 lies a dead zone where every LLM client overshoots in lockstep. The missing load-pressure header is a protocol gap, not a client-side bug.

llm-infrastructure
backpressure
rate-limiting
agent-fleets
+1
Showing 325–336 of 2313 posts
Prev28 / 193Next