Skip to main content
Tian Pan

Tian Pan

Software Engineer

View all authors

·tian

The Fallback Model Whose System Prompt Was Tuned for Someone Else

Failover keeps your LLM app available when the primary goes down — but the fallback model reads a system prompt that was tuned for someone else, and your users notice the difference.

llm-reliability
failover
prompt-engineering
multi-model
+1
·tian

The Few-Shot Example Your Model Treated as Binding Precedent

Few-shot examples are not neutral demonstrations — they are case law. Models bind to the closest example by surface tokens and inherit its constraints, shipping confident-wrong answers eval suites cannot see.

insider
llm
prompt-engineering
evals
+1
·tian

The Fine-Tune Artifact Your Departing Engineer Took With Them

A fine-tuned model is not a file in a registry; it is the closure of a pipeline over a training set. The teams that ship only the weights discover their bus factor the day a base-model migration arrives and the original engineer is gone.

mlops
fine-tuning
reproducibility
ml-engineering
·tian

The JSON Schema Your Output Passed and Your Downstream Consumer Rejected for Semantic Drift

A JSON schema validates shape, not meaning. When an LLM upgrade shifts the distribution of values inside a still-valid schema, downstream consumers detonate while the producer's dashboard stays green.

insider
llm
structured-outputs
contract-testing
+2
·tian

The KV Cache Eviction Your Provider Called Cache Pressure and Your Bill Called a Doubled Prefix Charge

Prompt caching looks like a configured discount, but KV cache eviction on shared LLM infrastructure turns it into a probabilistic one — the same conversation can cost several times more on a busy day with no code change.

insider
prompt-caching
kv-cache
llm-cost
+2
·tian

The KV Cache Warm-Up Cron That Ran in Blue and Never in Green Because the Host Pinning Never Moved

A blue/green deployment orphaned a cron pinned to the old color, the prompt cache went cold, and the bill quietly tripled — anatomy of a silent regression and the four practices that close the seam.

insider
prompt-caching
blue-green-deployment
cron-jobs
+2
·tian

The Legal Disclaimer That Leaked From The Answer Into The Tool Call Arguments

A safety disclaimer added to your system prompt does not stop at the user-facing reply. It rides along into every tool call argument the model produces — and into the downstream systems those calls fire against.

llm-agents
prompt-engineering
compliance
tool-calling
+1
·tian

The LLM-as-Judge Ensemble That Agreed Because All Judges Were the Same Family

An LLM-as-judge ensemble drawn from one provider family measures family-internal consistency, not judgment quality, and the high agreement number is an artifact of provider selection nobody named.

llm-as-judge
evaluation
ensemble
bias
·tian

The Logprobs Field Your Provider Removed That Broke Your Confidence Router Silently

A confidence router that stopped escalating low-confidence answers, the silent provider tier change that caused it, and how response-shape contracts, population-level alerts, and a fallback written for the wrong failure mode hide together.

insider
llm
observability
api-contracts
+2
·tian

The max_tokens Default Your Provider Raised That Doubled Your Tail Response Length

An LLM provider quietly raised the default max_tokens value and your p99 output length doubled overnight. The parameter you did not send is the configuration that changed under you — here is how to stop inheriting defaults you do not control.

insider
llm
api-design
production
+2
·tian

The MCP Server Your Team Forgot Was Running with Prod Credentials

An MCP server running on a developer laptop with a CI-grade OAuth token is a production attack surface. Here is how DNS rebinding, bad bindings, and shared tokens turn one compromised tab into a deploy-key leak.

insider
mcp
security
oauth
+1
·tian

The Model Card Benchmark Whose Methodology Shifted While Your Contract Cited the Number

A benchmark number is a measurement under a protocol, and the protocol is what your vendor controls. Pin the methodology, or contract on your own eval suite.

insider
llm
benchmarks
procurement
+2
Showing 181–192 of 2313 posts
Prev16 / 193Next