Provenance & supply chain

Where Your AI Comes From: Why Model Provenance Belongs in Your Procurement Checklist

You would never deploy a critical software component without asking who built it and where. AI models deserve the same question, and most buyers aren't asking it.

Open-weight AI models have made something remarkable possible: you can download a state-of-the-art model, run it entirely on your own hardware, and never send a byte of data to anyone. That's a genuine win for privacy and control. But "open-weight" is often mistaken for "neutral," and it isn't. Every model arrives with a history (who trained it, in what country, on what data, under what license) and for organizations with real security obligations, that history matters.

Open-weight is not origin-neutral

When you adopt an open model, you inherit decisions you didn't make. The model's behavior, its built-in guardrails, the data it was trained on, and the values baked into it were all chosen by its creator. The leading open models today come from a handful of organizations across different countries: some from US companies (Meta, OpenAI, Microsoft, IBM, Google), some from elsewhere. They are all technically excellent. They are not all the same choice for every buyer.

Two different questions people confuse

The most important thing to get straight is that "is my data safe?" and "is this model appropriate for us?" are different questions.

  • Runtime data safety is largely solved by ownership. A model you run on your own infrastructure, even fully air-gapped, does not phone home. The weights are just static numbers; they can't exfiltrate your data. On this axis, a model's country of origin is almost irrelevant: run it offline and your data stays put, full stop.
  • Provenance and supply-chain trust is a separate axis, and ownership doesn't erase it. Who trained the model? On what data? Could its behavior have been shaped, subtly or deliberately, in ways that matter to you? Does your procurement policy, or your customer's, restrict technology by country of origin? These questions don't go away just because the model runs locally.

Conflating the two leads to bad decisions in both directions: dismissing a perfectly safe model over runtime fears, or waving through a model whose provenance your own rules should have flagged.

Why regulated and defense buyers weigh origin

For most commercial use, country of origin is a minor consideration. For defense, government, and critical-infrastructure buyers, it can be decisive, for reasons that are about policy and assurance, not xenophobia:

  • Procurement rules. Many public-sector and defense contracts restrict foreign-origin technology outright. A model's origin can determine eligibility before any technical merit is considered.
  • Training-data provenance. You generally cannot fully audit what a large model was trained on. The less you can verify, the more the identity and jurisdiction of the trainer matters as a proxy for trust.
  • Behavioral assurance. A model's guardrails and biases reflect its makers' choices and legal environment. For sensitive missions, "we can't fully explain why it behaves this way" is a real risk to weigh.
  • Optics and accountability. Sometimes the requirement is simply defensible sourcing: being able to tell an auditor, a regulator, or a customer exactly where your AI came from.

None of this means foreign-origin models are unsafe or low-quality. Many are outstanding and cleanly licensed. It means origin is a legitimate axis of a procurement decision, to be weighed alongside capability, cost, and license, not ignored.

License and origin are not the same thing

A common mistake is to treat a permissive license as an all-clear. They're independent:

  • License governs what you're legally allowed to do with the weights (use commercially, modify, redistribute). Permissive licenses like Apache-2.0 and MIT are the cleanest for ownership; some strong models ship under custom terms that permit commercial use but deserve a legal read.
  • Origin is about where the model and its training came from: a trust and policy question that a license says nothing about.

You have to check both. A model can have an impeccable license and an origin your policy disallows, or vice versa. Serious model selection looks at capability, hardware fit, license, and provenance together.

What good practice looks like

  1. Record provenance as metadata. For every candidate model, capture creator, country of origin, release date, and license, and keep it visible, not buried.
  2. Match the model to the buyer. For an unrestricted commercial client, pick the strongest model that fits the budget. For a defense/government/regulated client, start from well-licensed, trusted-origin options and only step outside that set with eyes open.
  3. Put it in the deliverable. The model your organization ends up running should come with a plain statement of its provenance ("this model was built by X, in country Y, released Z, under license L") so anyone who later asks "where did this come from?" has an answer on paper.
  4. Own the choice. When you own a private, fine-tuned model, you decide its provenance up front, instead of inheriting whatever a cloud provider happens to run this quarter.

The bottom line

Open-weight AI gives you control you never had with rented cloud models, but control includes the responsibility to choose deliberately. Where your AI comes from is not a paranoid question; it's basic supply-chain hygiene, and for security-conscious organizations it belongs on the checklist next to capability, cost, and license. Ask it early. Write the answer down. And pick the provenance that fits your mission.

Want a model whose origin you can put on paper?

We help you choose the right base model for your requirements, provenance and license included and documented on every delivery, then fine-tune it into a private model you own. Start free with a readiness scorecard, or book a short call.