Skip to main content

990 posts tagged with "insider"

View all tags

Your Context Has Mass: Data Gravity and the Return of Move-Compute-to-Data

· 9 min read
Tian Pan
Software Engineer

The Hadoop generation learned one lesson so thoroughly it became a reflex: moving data is expensive, so move the computation to the data. Every MapReduce scheduler, every HDFS block placement decision, every "data locality" dashboard existed to serve that principle. Then, somewhere between the rise of managed model APIs and the agent boom, we quietly inverted it — and nobody repriced the decision.

Look at what a modern agent loop actually does. It retrieves a stack of documents from a vector store, pulls a repo snapshot from object storage, collects tool results from half a dozen internal services, concatenates all of it into a context window, and ships the whole payload to a model endpoint that usually lives in a different VPC, often a different region, sometimes a different cloud. Then it does it again on the next turn. And the next. Your context has mass, and you are paying freight on every hop.

Your Context Pipeline Needs a Freshness SLA

· 9 min read
Tian Pan
Software Engineer

Your agent answered a customer's billing question with last quarter's pricing, and the postmortem will blame the model. It shouldn't. The prompt was assembled correctly, the retrieval scored well, the model reasoned soundly over everything it was given — and everything it was given was true three days ago. Somewhere between the CRM export, the docs sync, and the vector index rebuild, "current state of the world" quietly became "state of the world as of Tuesday," and nothing in your stack was measuring the difference.

Data engineers solved this class of problem years ago. A downstream dashboard consuming ten upstream tables gets lineage, freshness checks, and an on-call rotation that pages when the nightly job slips. The context window your agent consumes is the same thing — a materialized view joined from docs, tickets, code, CRM, and memory — except nobody owns the join, nothing measures its staleness, and when it serves yesterday's truth the failure gets filed as "the model hallucinated."

Your Data Agent Needs One Definition of Revenue

· 8 min read
Tian Pan
Software Engineer

Text-to-SQL demos never die on syntax. The model writes fluent SQL — better than most junior analysts, honestly — and the query runs, and a number comes back. The demo dies three weeks later, in production, when the CFO notices that the agent's "Q2 revenue" doesn't match the board deck. Not because the SQL was malformed, but because the warehouse contains three defensible definitions of revenue — bookings, recognized, and net-of-refunds — and the model confidently picked one. Just not the one finance uses.

This is the failure mode that matters, and it's invisible to every benchmark you've seen. The fix isn't a better model or a longer prompt. It's a piece of infrastructure most data teams already half-built and then abandoned: the semantic layer. The metrics definitions you wrote for BI dashboards — dbt metrics, LookML, cube definitions — turn out to be the missing tool contract for data agents. The teams shipping reliable agents figured out that the build order is inverted from what everyone assumed: semantic layer first, agent second.

Your Fine-Tune Is a Fork You Have to Maintain

· 10 min read
Tian Pan
Software Engineer

The budget meeting for a fine-tuning project always prices the wrong thing. Teams estimate the data pipeline, the training runs, the eval passes — a one-time investment with a clear finish line. Then the model ships, the accuracy chart goes up and to the right, and everyone moves on. Six months later an email arrives: the base model your adapter is welded to has a retirement date. Nothing about your system changed. Everything about its foundation did.

This is the part nobody prices in: a fine-tune is not a product you finished. It is a fork of someone else's codebase, and every base-model release is an upstream rebase you didn't schedule. Anyone who has carried private patches against a fast-moving open-source project knows exactly how this story goes — the fork is cheap to create and expensive to keep.

Your Internal Framework Is a Low-Resource Language

· 9 min read
Tian Pan
Software Engineer

Ask a coding agent to build a React component and it writes idiomatic, hook-shaped, accessibility-annotated code on the first try. Ask the same agent to use your in-house ORM — the one your platform team has maintained for six years, the one with excellent docs and a hundred internal consumers — and it hallucinates methods that don't exist, invents configuration options from some other library, and confidently ships code that compiles against an API it made up.

The difference isn't quality. Your ORM might be better-designed than half the open-source libraries the model handles flawlessly. The difference is training data. React has millions of public repositories behind it; your framework has zero. In the vocabulary of natural language processing, your internal framework is a low-resource language — and every consequence NLP researchers documented for low-resource languages now applies to your codebase.

When the Clock Is a Tool: Agents, Time Zones, and the Bug That Only Happens at Midnight

· 9 min read
Tian Pan
Software Engineer

Ask a large language model what time it is and you will get a confident answer that is almost certainly wrong. Not because the model is broken, but because there is no clock inside it. A transformer is a stateless text-completion engine: it maps tokens to tokens. Nowhere in that pipeline does a signal arrive that says "it is now 14:32 UTC." The current moment is not something the model perceives — it is something you have to hand it, every single turn, or it will invent one from the stale sediment of its training data.

This is the quiet failure that surfaces at the worst possible moments. Your agent believes it is Monday because the session opened on Monday, and it keeps believing that on Tuesday, on Wednesday, right up until it schedules a "tomorrow morning" reminder for a day that has already passed. It reasons about "the last 24 hours" of logs using a now that froze hours ago. It converts a meeting time across time zones and lands an hour off because it assumed the wrong side of a daylight-saving boundary. None of these look like hallucinations in the classic sense. The output is fluent, plausible, and internally consistent. It is just anchored to a moment that no longer exists.

Conway's Law Comes for Your Agent Fleet

· 9 min read
Tian Pan
Software Engineer

Pull up the architecture diagram for your multi-agent system. Now pull up your org chart. If you squint, they're the same picture. The "research agent" maps to the team that owns search. The "billing agent" has a hard boundary exactly where Finance stops talking to Product. The orchestrator that fans work out to five specialists looks suspiciously like an engineering manager with five direct reports. You didn't decide this on purpose. Conway's Law decided it for you.

Melvin Conway's 1967 observation is that any system you design will mirror the communication structure of the organization that built it. For sixty years this was a story about microservices and monoliths. But agent fleets are the most literal demonstration of the law I've ever seen: the agents are communication structures. An agent boundary is a place where one process hands a message to another and waits. When you draw those boundaries to match your teams instead of your problem, you don't just inherit your org chart's shape — you inherit its dysfunction, and you run it at machine speed.

The Indemnification Gap: When Your Agent Takes an Irreversible Action, Whose Budget Eats It?

· 9 min read
Tian Pan
Software Engineer

Your agent just issued a $40,000 refund to the wrong account, re-routed a freight order that triggered expedited shipping fees, or pushed a config change that took down a customer's production environment for six hours. The action is done. It is irreversible, or close enough that reversing it costs real money. Now the only question that matters is the one nobody asked before you shipped the thing: whose budget eats the loss?

Most teams discover the answer the hard way, in a conference room three days later, with the vendor's account manager on speakerphone reading a liability cap back to them. The cap is the annual subscription fee. The loss is forty times that. The conversation is short.

The Standup Is Lying: Coordinating Work When Agent Fleets Run Overnight

· 10 min read
Tian Pan
Software Engineer

"What did you do yesterday?" is the first question of every standup, and on a team that runs agent fleets overnight it has become impossible to answer honestly. The literal answer is: I wrote three prompts, went home, and woke up to eleven pull requests, four of which I have not read yet. The person reciting their update is not lying on purpose. The ritual is lying for them, because it was built around an assumption that no longer holds — that the unit of work is a human doing one thing at a time, serially, during business hours.

That assumption is load-bearing. It holds up the burndown chart, the sprint commitment, the velocity number, the "blocked / in progress / done" columns, and the whole choreography of who-tells-whom-what-when. Pull the assumption and the artifacts don't gracefully degrade. They keep producing numbers that look authoritative and mean nothing. A team can have a beautiful burndown and a green sprint while half its actual throughput happened between midnight and 6 a.m., attributed to no one, reviewed by no one, and reflected in no ceremony.

Approval Fatigue: How Human-in-the-Loop Gates Decay Into Rubber Stamps

· 9 min read
Tian Pan
Software Engineer

Put a human in the loop and you have a control. Put a hundred approval requests a day in front of that human and you have a rubber stamp wearing a control's badge. The two look identical on an architecture diagram. They behave nothing alike under load, and the gap between them is where most "responsible AI" deployments quietly fail.

The pattern is familiar to anyone who has watched a security operations center drown. Industry surveys put the share of alerts that go uninvestigated somewhere between a quarter and two-thirds — one frequently cited figure is that 62% of alerts are simply ignored, and 55% of teams admit to regularly missing alerts they would classify as critical. The analysts aren't lazy. They are processing a few thousand alerts a day with a false-positive rate that often exceeds 50%, and the human brain responds to that ratio exactly the way you'd expect: it stops looking. Agentic AI is now recreating this exact failure mode, one confirmation dialog at a time.

The Documentation Renaissance: Your README Is the Agent's Primary Context Surface

· 10 min read
Tian Pan
Software Engineer

For two decades, documentation was where good intentions went to die. You wrote the README during the first sprint, when the architecture was clean and your enthusiasm was high. Nobody read it. By the third sprint it was lying about the build command, and by the sixth it described a service that had been deleted. Documentation was a tax everyone agreed to pay and nobody actually paid — a moral imperative with no feedback loop. Write bad docs and nothing happened. Write no docs and nothing happened either, because the senior engineers carried the architecture in their heads.

Then we pointed coding agents at our repositories, and the feedback loop arrived overnight. The README is now the single highest-leverage file you own — not because anyone gave a motivational talk about documentation hygiene, but because the quality of that file now visibly determines whether your agent ships correct code or confidently hallucinates an architecture that no longer exists.

This is the documentation renaissance, and it has almost nothing to do with the documentation we used to write.

Comprehension Debt: The 2 A.M. System No Human Understands

· 9 min read
Tian Pan
Software Engineer

The pager goes off at 2:14 a.m. A checkout service is throwing 500s, revenue is bleeding, and you are the on-call engineer. You pull up the failing module and start reading. The code is clean — well-named functions, sensible structure, even a few helpful comments. And you have no idea what it does. You didn't write it. Nobody on your team really wrote it. An agent generated it four months ago, it passed review, the tests went green, and it has been running in production ever since. Now it's on fire, and the person who is supposed to fix it is meeting it for the first time.

This is comprehension debt: the widening gap between the amount of code your organization runs and the amount any human actually understands. It doesn't show up on a dashboard. It accrues silently while everything looks healthy, and it comes due at the worst possible moment — during an incident, when the cost of not understanding your own system is measured in downtime.