Where Your Tokens Sleep at Night: Data Residency for LLM API Calls
A customer support transcript leaves a server in Frankfurt, gets concatenated into a prompt, and crosses the Atlantic to a GPU in Virginia. The model thinks for 800 milliseconds. A response comes back. From the user's perspective, nothing happened — the chat just worked. From your regulator's perspective, you transferred personal data to a third country, and you may not be able to name the legal basis for it.
This is the part of LLM adoption that demos hide. Prototypes call api.openai.com and ship. Then a procurement questionnaire from a German bank, a French hospital, or your own legal team asks a question that the prototype never had to answer: where does the inference happen, and who can compel access to it? "The provider is SOC 2 compliant" is the reflexive answer, and it is the wrong one — it answers a question about the provider's internal controls, not about which jurisdiction's courts can reach into your prompts.
