Sovereign inference is the practice of serving AI models on hardware the operator controls, such that business data and model tokens never leave the environment — and anything that does leave, leaves explicitly, with consent, on the record.
What it requires
Locality. Weights on machines you control — owned, leased, or contractually dedicated. The defining property is that inference happens inside your boundary.
Explicit escalation. Local models have limits. A sovereign posture does not pretend otherwise; it makes crossing the boundary a decision — declared, consent-based, and recorded — rather than a default.
An evidentiary record. Sovereignty without evidence is a promise. Every inference boundary-crossing, every escalation, every model version in production belongs in an append-only record.
What it is not
It is not self-hosting as a hobby, and it is not air-gap absolutism. Most organizations do not need zero external calls; they need to know exactly which calls happen, and to hold the power to stop them. Sovereignty is a posture with a dial, not a bunker.
Turbine's dedicated and on-premises deployment models implement this posture: local inference by default, logged escalation by exception. See deployment models →