September 28, 2026
TheVARA Cybersecurity Requirements Checklist rules every VASP must meet: 19 policy criteria, CISO duties, the 72-hour incident deadline, rule by rule.

September 22, 2026
Sovereign LLMs and sovereign AI agents explained: what makes them sovereign, how they differ, and where enterprise risk actually sits.

September 21, 2026
Sovereign AI vs private AI: understand data residency, security risk, and compliance differences before choosing your enterprise AI deployment model.
AI supply chain security is the practice of protecting every external component that goes into building and running an AI system, including training data, pre-trained models, ML libraries, plugins, and third-party APIs, from tampering, poisoning, and compromise. It ensures the models and agents you deploy are exactly what you think they are.
The risk is now formally recognized. OWASP's Top 10 for LLM Applications lists supply chain vulnerabilities as one of its ten core risk categories (LLM03:2025). Most organizations do not build models from scratch. They download weights from public hubs, fine-tune on outside data, and connect agents to third-party tools. Each of those inputs is a place where an attacker can hide. Some model file formats can execute code the moment they are loaded, and a poisoned dataset can change a model's behavior without a single line of application code being touched.
This guide is about securing AI systems, not about using AI to manage logistics or procurement.
You'll learn how the AI development supply chain works stage by stage, where models, data, and agent tools get compromised, how real attacks have played out, and which controls reduce the risk.
AI supply chain security is the discipline of verifying and protecting everything an AI system inherits from outside your own team: the data it learns from, the models it starts with, the code that runs it, and the services it calls. The goal is to make sure nothing in that chain has been tampered with, poisoned, or swapped before it reaches production.
The threat is documented. JFrog's security team scanned models on Hugging Face and found roughly a hundred models with malicious functionality, some able to execute code on the victim's machine and give attackers a persistent backdoor. Those files were uploaded to a mainstream public hub, not a fringe site.
A traditional software supply chain consists of source code, open-source packages, build systems, and container images. You can inspect code, diff versions, and trace a package back to a commit. AI supply chain security includes all of that, because AI applications still depend on packages and pipelines the same discipline covered in our guide to what DevSecOps actually involves. It then adds artifacts that behave differently: model weights, training datasets, and fine-tuning adapters.
The practical difference is auditability. A compromised library usually leaves a readable trail in a repository, and a compromised model file does not. A poisoned model can pass every functional test and still behave badly on a specific trigger. Tools built for software supply chains, such as dependency scanners and SBOMs, cover only part of the AI problem. They work best when paired with model-specific controls like provenance checks and safe-format enforcement, and with a broader vulnerability assessment of the systems around them.
Any external input that shapes what an AI system does or where it runs is part of the supply chain:
Model weights and checkpoints, whether downloaded from a public hub or received from a vendor.
Datasets, including training data, fine-tuning data, and evaluation sets.
Embeddings and vector databases, especially those built from third-party content.
Frameworks and libraries, such as PyTorch, TensorFlow, and the packages around them.
Adapters, such as LoRA files layered on top of a base model.
Plugins, tools, and MCP servers that give agents access to other systems.
Prompts and system instructions, when copied from external templates or shared libraries.
Hosted model APIs, where you depend on a provider's security and behavior.
GPUs and cloud infrastructure, including the images and drivers your training and inference workloads run on.
Teams often track only the first item and miss the rest, which is how blind spots form. A company might vet a base model carefully and then load an unreviewed adapter or connect an agent to a community-built tool.
Three properties make AI systems harder to secure than conventional software.
The first is opacity. Model weights are large arrays of numbers that no reviewer can read for intent. You can test outputs, but you cannot inspect a model the way you would review source code, so hidden behavior can survive normal quality checks.
The second is executable model formats. Some common formats can run code when they are loaded, not only store data. Pickle files, a common serialization format for Python objects, can contain arbitrary code that runs when the file is opened. Loading an untrusted model can therefore be equivalent to running an untrusted program. Safer formats such as Safetensors exist because they do not allow arbitrary code execution, but adoption is uneven.
The third is third-party data you cannot fully audit. Training sets are often scraped or assembled from many sources, far too large to review by hand. An attacker who can influence even a small slice of that data may be able to shape model behavior without ever touching your infrastructure. This widened surface is also why the concept of an enterprise cybersecurity platform increasingly needs to account for AI-specific assets, not just conventional endpoints and applications.
Together, these properties mean a supply chain compromise in AI can be both easy to introduce and difficult to detect afterward. That is why the rest of this guide treats provenance, format safety, and isolation as core controls, not optional extras.
The AI development supply chain has five stages: data collection, training and fine-tuning, model distribution through hubs and registries, packaging and deployment, and the agents and tools that act on a model's output. Each stage brings in components you don't fully control, and each has its own way of being compromised.
Security teams often treat "the model" as one asset. In practice it is the end product of a pipeline, and an attacker only needs to compromise one stage for the damage to carry forward into everything downstream. The table below maps the pipeline. The sections after it explain each stage.
Stage | What enters | Who usually controls it | How it gets compromised |
|---|---|---|---|
Data collection and labeling | Scraped web data, licensed datasets, crowd or vendor labels | Data vendors, annotators, open-source dataset maintainers | Poisoned or manipulated samples, mislabeled data, tampered dataset files |
Training and fine-tuning | Base models, datasets, adapters, training code, GPU environments | Your ML team, plus every upstream source it pulls from | Backdoored base model, malicious adapter, vulnerable training dependency, exposed training credentials |
Model hubs and registries | Downloaded and uploaded weights, model cards, configs | Hub operators and individual publishers | Malicious files, impersonated repositories, leaked publisher tokens, unsafe serialization formats |
Packaging, deployment, and inference | Containers, ML frameworks, serving libraries, cloud images, hosted model APIs | Platform and DevOps teams, plus cloud and API providers | Vulnerable dependencies, poisoned packages, compromised images, provider-side incidents |
Agents, tools, and integrations | Plugins, MCP servers, tool descriptions, API credentials | Third-party developers and your integration team | Malicious or hijacked tools, over-permissioned access, manipulated tool instructions |

The first stage is where a model's future behavior begins. Data comes from web scraping, licensed datasets, open-source collections, and human or vendor labeling, and most of it is too large to review by hand. That makes it a low-cost target. An attacker who can influence a small share of the samples, or the people labeling them, can shift how the model behaves without ever touching your systems.
The practical risk is that the damage stays hidden. A poisoned dataset produces a model that passes normal accuracy tests and misbehaves only on specific inputs. Teams should record where every dataset came from, keep integrity hashes of the versions they trained on, and treat externally sourced or crowd-labeled data as untrusted until sampled and checked.
Training is where most outside components meet. Teams rarely train from scratch. They start from a pre-trained base model, fine-tune it on additional data, and often layer on adapters such as LoRA files. Every one of those inputs was created by someone else, and a compromised base model or adapter carries its hidden behavior straight into your version.
The training environment also matters. Training jobs run on GPU clusters with broad access to data stores and credentials, so a vulnerable dependency or a malicious training script can expose far more than the model itself. Keep training environments isolated from production, pin dependency versions, and give training jobs only the access they need.
Model hubs are the distribution layer. Teams download weights from public hubs the way developers pull packages from a registry, which makes the hub a central point of trust and a central point of failure. Weaknesses here include malicious model files, repositories that impersonate popular ones, and the credentials publishers use to update their models.
That last risk is documented. Researchers at Lasso Security reported that exposed API tokens found on GitHub and Hugging Face would have given them access to repositories for Meta's Bloom, Meta-Llama, and Pythia models, which could have let an attacker silently tamper with widely used models. The lesson is that trusting a well-known model name is not enough. You also need to verify the specific file you downloaded, its publisher, and its format.
Once a model is chosen, it must be packaged and served. This stage looks the most like traditional software supply chain security. Containers, ML frameworks, serving libraries, and cloud images all come from outside sources, and each can carry a known vulnerability or a poisoned package. Dependency confusion and typosquatting attacks work here just as they do for any other software.
Hosted model APIs add another dependency. If your application calls a third-party model, you inherit that provider's security practices and any change they make to the model's behavior. This is also the point where organizations with strict data-residency or control requirements start weighing self-hosted, sovereign AI infrastructure against a purely hosted approach see our comparison of sovereign AI vs private AI for how that trade-off plays out. Standard controls apply regardless of the path you choose: software bills of materials, image scanning, signed artifacts, and version pinning. For hosted APIs, add contractual security requirements and monitoring for unexpected changes in model output.
The final stage is where AI stops generating text and starts taking actions. Agents connect to plugins, tools, and MCP servers, and those connections are third-party code with real access to email, files, databases, and internal APIs. A malicious or compromised tool can misuse that access directly, or manipulate the agent through instructions hidden in its own description.
This is the stage where over-permissioning does the most damage, because an agent inherits every permission you grant it. Treat each tool as untrusted third-party software: approve tools individually, grant the narrowest access that works, and require human confirmation for high-impact actions.
The main AI supply chain risks come from eight places: poisoned training data, backdoored pre-trained models, unsafe model file formats, vulnerable ML frameworks, typosquatted or confused packages, compromised hubs and adapters, unclear licensing and provenance, and weak vendor or API dependencies. What they share is that each lets an attacker or a flaw enter your system through something you didn't build.
Data poisoning is the deliberate insertion of bad samples into a training or fine-tuning dataset so the resulting model behaves the way an attacker wants. The manipulation can be subtle: a small set of mislabeled or trigger-carrying examples can leave overall accuracy intact while changing the model's response to specific inputs.
That makes poisoning difficult to catch with standard testing, because the model still performs well on ordinary evaluations. The risk is highest when data is scraped from open sources, sourced from crowd workers, or assembled from datasets nobody on your team has reviewed. Recording each dataset's origin, hashing the exact versions you train on, and sampling external data before it enters a pipeline are the practical first defenses.
A backdoored pre-trained model behaves normally until it receives a specific trigger, at which point it produces output the attacker chose. Because most teams start from a downloaded base model rather than training from scratch, a compromised model passes its hidden behavior to every application built on it.
Fine-tuning does not reliably remove this, and you cannot detect it by reading the weights. What you can do is narrow the exposure: prefer models from publishers you can verify, record exactly which model version each application uses, and test behavior against adversarial inputs before deployment, not only against accuracy benchmarks.
Some model file formats can run code the moment they are loaded. Pickle, the default serialization format behind many PyTorch models, can embed arbitrary code that executes on load, so opening an untrusted model file can be the same as running an untrusted program. Safetensors was designed to avoid this: it is a safer serialization format that does not allow arbitrary code execution.
Scanning is helpful but not sufficient. JFrog analyzed the Hugging Face models that pickle scanning had flagged as unsafe and found that more than 96% of them were false positives, because basic scanners look for suspicious function names rather than analyzing what the code does. Scanners themselves also have flaws: JFrog later disclosed three critical zero-day vulnerabilities in PickleScan, a tool many platforms rely on. The dependable control is to remove the risk at the source by requiring Safetensors or other non-executable formats wherever they are available, and loading anything else only inside an isolated environment.
ML frameworks, serving libraries, and their dependencies are ordinary software, and they carry ordinary vulnerabilities. A typical AI project pulls in a deep tree of packages for training, data handling, and inference, and any one of them can contain a flaw or be compromised upstream. Identifying and prioritizing these flaws is the same discipline covered in a standard vulnerability assessment, just applied to ML-specific dependencies.
The added problem is that AI environments are often built quickly by research and data science teams, outside the patching and review routines that production software gets. Notebooks, training servers, and inference containers can run outdated packages for months. Applying the same practices you use for other software helps: maintain a software bill of materials, scan dependencies and container images, pin versions, and put training and experimentation environments under the same patch management as production. This is the kind of gap a structured vulnerability management program is built to catch.
Typosquatting and dependency confusion trick a package manager into installing an attacker's package in place of the legitimate one. Typosquatting relies on a near-identical name, while dependency confusion relies on the installer preferring a public index over a private or secondary one.
The clearest AI-related example involved PyTorch. Between December 25 and December 30, 2022, the PyTorch nightly build pulled a dependency called torchtriton, and a malicious package with the same name on PyPI was installed instead of the official one, because the public index took precedence. The malicious version was designed to send system information and files off the machine. Users of the stable PyTorch packages were not affected, which shows why pinning versions and using stable releases matters. Practical defenses include configuring installers to trust only specified indexes, reserving your internal package names on public registries, and installing from pinned, hash-verified requirements.
Public model hubs work like package registries for AI, so they inherit the same risks: malicious uploads, repositories that impersonate popular ones, and publisher accounts whose credentials can be stolen and used to modify a trusted model. A hub's scanning helps, but it does not guarantee that every file is safe.
Third-party adapters deserve special attention. Small add-on files such as LoRA adapters are easy to share and easy to load, and teams often apply them with less scrutiny than a full base model even though they change the model's behavior. Treat adapters as you would any other executable dependency: confirm the publisher, review the file format, and test the combined model before it reaches production.
Not every supply chain risk is an attack. A model or dataset with unclear licensing or unknown origin can create legal and compliance exposure, and it also makes security harder, because you cannot verify what you cannot trace. If you don't know where a model's training data came from, you cannot assess whether it was tampered with, or whether you are permitted to use it commercially. This overlaps directly with broader governance, risk, and compliance obligations most enterprises already track for other software.
The practical fix is documentation. Record the source, license terms, and version of every model and dataset you use, and review those terms before adopting a component, especially for commercial products. This inventory is also the foundation for an AI bill of materials, which the later best-practices section covers.
When an application calls a hosted model or depends on an AI vendor, that provider's security practices become part of your own. A breach at the vendor, a change in how the model behaves, or a service outage can all affect you, and you often have limited visibility into what changed or why.
The controls here are largely contractual and operational. Ask vendors about their security practices, data handling, and incident notification commitments before adopting them, limit what data you send to third-party APIs, and monitor model output for unexpected changes. For critical workflows, plan an alternative provider so a single vendor problem doesn't stop the business.

LLM and AI model supply chain attacks work by compromising something a team trusts before it reaches production: a dataset, a model file, a package, or the pipeline that builds them. The attacker never needs to break into the deployed application. They only need the victim to download and run the tampered component.
Most AI supply chain attacks fall into four patterns, and each maps onto broader categories covered in our overview of common types of cyber security attacks.
Poison. The attacker manipulates the data a model learns from so it behaves differently on chosen inputs. This is the hardest pattern to observe in the wild. Poisoning is well understood in research, but public, confirmed cases against production models are scarce, so treat it as a real risk rather than a well-documented incident type.
Tamper. The attacker modifies an artifact the victim already trusts, such as a model file, a package release, or a repository. This is often done by stealing the credentials of a legitimate publisher or by compromising the build system that produces the release.
Impersonate. The attacker publishes something that looks legitimate: a model that mimics a popular one, or a package with a near-identical name. Victims install it believing it is the real thing.
Hijack a dependency. The attacker targets a component deeper in the dependency tree, so the malicious code arrives through a library the victim never chose directly. Dependency confusion and compromised build pipelines both belong here.
These patterns combine. A stolen publisher token, for example, enables tampering, and the tampered package then spreads as a hijacked dependency.
The cases below are documented by researchers or by the affected projects themselves. Each one teaches a different control.
Incident | Vector | What the defender should learn |
|---|---|---|
Malicious models on Hugging Face (JFrog, 2024) | Model files using pickle serialization that run code when loaded, giving attackers a reverse shell. JFrog reported at least 100 malicious models on the platform. | Loading a model file can be equivalent to running a program. Prefer non-executable formats like Safetensors and load anything else only in isolation. |
Exposed publisher tokens (Lasso Security, 2024) | Researchers reported that unsecured API tokens found on GitHub and Hugging Face gave access to repositories for Meta's Bloom, Meta-Llama, and Pythia models, which could have allowed silent tampering. This was demonstrated access, not a confirmed malicious modification. | Publisher credentials are part of the supply chain. Rotate and scope tokens, and verify the file you download, not only the model's name. |
PyTorch nightly "torchtriton" (December 2022) | Dependency confusion. A malicious package with the same name as the project's own dependency was uploaded to PyPI, and because the public index took precedence, it was installed instead of the official one. The affected window ran from December 25 to 30, 2022, and stable releases were not affected. | Configure installers to trust specific indexes, reserve internal package names publicly, and pin dependencies with hash verification. |
Ultralytics on PyPI (December 2024) | Build pipeline compromise. Attackers exploited a known GitHub Actions script injection to reach the build environment and published versions carrying a cryptocurrency miner. | Your build pipeline is part of your attack surface. Restrict CI permissions, and verify that published packages match reviewed source. |
The Ultralytics case shows how much reach one build compromise has. The library had almost 60 million downloads, and the attackers injected the malicious code after the code review step was finished, so the source repository looked clean while the published package was not. The first fix also failed: version 8.3.42, released as the safe upgrade, contained the same malicious code because the maintainers had not located the compromise, and two further malicious versions followed on December 7, likely using a PyPI token stolen during the initial breach. The payload was a miner, but researchers noted the same vector could have delivered more harmful malware, such as backdoors or remote access trojans.
The compromised component usually works. The malicious PyTorch dependency, for example, kept the legitimate package's functionality while adding its own code, so users had no visible reason to suspect it. A tampered model passes ordinary accuracy tests. A poisoned dataset produces a model that looks fine until a trigger appears.
Three factors make detection harder than for conventional intrusions. First, the malicious code often runs inside a process you already trust, such as a notebook or training job, where it looks like normal activity. Second, code review does not help when the tampering happens after review, as in the Ultralytics build. Third, model files can't be read for intent, so inspection can only catch known-bad patterns, not unknown ones.
In several incidents above, the first signal was an operational symptom, such as a spike in CPU usage, not a security alert. Teams that log provenance, compare published artifacts against reviewed source, and monitor training and inference environments for unexpected network or resource behavior have a better chance of noticing something earlier the same monitoring discipline covered in our guide to what is incident response.
A supply chain attack compromises a component before deployment, while prompt injection manipulates a running model through its input. The first changes what the system is, and the second changes what the system is told.
That distinction matters for defense. If an attacker hides a backdoor in a model file, no amount of input filtering will remove it, because the behavior is built into the artifact you loaded. If an attacker sends a malicious instruction through a user message or a retrieved document, verifying your model's provenance will not stop it, because the model itself is genuine.
The two can also work together. A compromised plugin or tool can deliver prompt injection by carrying manipulated instructions in its description, and a poisoned model can be more susceptible to certain injected prompts. Treat them as separate controls with overlapping consequences: provenance, format safety, and isolation for the supply chain, and input handling, output validation, and permission limits for prompt injection.

AI agent supply chain security is the protection of everything an agent connects to and depends on: its tools, plugins, MCP servers, and the credentials it uses to act. Agents don't just generate text, they take actions with real access, so a compromised tool can misuse that access directly.
Every tool, plugin, or MCP server an agent uses is third-party software running with the agent's trust. An MCP server is a small program that lets an AI assistant reach another system, such as email, a database, or a file store. Installing one is functionally the same as installing a package from a public registry, and it carries the same risks of impersonation, tampering, and unreviewed updates.
The first publicly documented case shows how this plays out. In September 2025, Koi Security reported that an npm package called postmark-mcp, an MCP server for sending email through Postmark, had been quietly backdooring users. Researchers believe it to be the first publicly documented malicious MCP server. The package copied the real Postmark server, behaved correctly for its first 15 versions, and then added a single line in version 1.0.16 that blind-copied every outgoing email to the publisher. Reports put its usage at about 1,500 downloads per week. Postmark itself confirmed the package was not an official tool and had no connection to the company.
The lesson is that a clean version history does not prove a package is safe. The attack needed no exploit, only trust and a single added line, in a tool that agents could call hundreds of times a day.
Agents decide which tool to call by reading each tool's description, and that text is instruction the model treats as trustworthy. A poisoned description can hide extra directions, such as telling the agent to pass data to another tool or to act outside the user's request. Because most users never read these descriptions, the manipulation can go unnoticed. This is where a supply chain attack overlaps with prompt injection: the malicious instruction arrives inside a component you installed, not in a user message.
Over-permissioned credentials raise the damage. Security researchers quoted in coverage of the postmark case noted that MCP servers typically run with high trust and broad permissions inside agent toolchains, so any data they handle can be sensitive. An email tool that can read and send across an entire mailbox, or an agent holding a token with admin rights, turns one compromised component into a company-wide exposure.
Four controls cover most of the risk, and they work best together.
Least privilege means giving each tool and each agent only the access its task needs: read-only where reading is enough, scoped tokens instead of admin keys, and separate credentials per tool so one compromise doesn't spread the same principle behind a broader zero trust security approach.
Allowlists mean agents can use only tools that have been reviewed and approved. Pin each tool to a specific version and re-review before updating, since the postmark package turned malicious after fifteen trusted releases.
Sandboxing means running tools in isolated environments with restricted network access, so a malicious tool cannot freely reach internal systems or send data to arbitrary destinations.
Human approval on high-impact actions means requiring a person to confirm anything hard to reverse, such as sending external email, moving money, deleting data, or changing permissions. This is the last line of defense when a tool has already been compromised.
None of these controls replaces testing. Teams should also verify how an agent behaves when a tool misbehaves, which is the focus of AI agent security testing.
Several established frameworks address AI supply chain security directly, each aimed at a different audience: developers, government agencies, and organizations pursuing formal certification. Knowing which one applies to your situation prevents wasted effort chasing controls a client or regulator was never going to ask for.
OWASP's Top 10 for LLM Applications lists Supply Chain Vulnerabilities as LLM03, one of ten core risk categories for LLM-based systems. The 2025 revision expanded this entry specifically because agents now pull tools, frameworks, and other agents from external sources, which widens what "supply chain" means for an LLM application beyond the base model itself. The framework also folded the 2024 list's separate "Insecure Plugin Design" and "Model Theft" categories into LLM03 and LLM06 (Excessive Agency), reflecting how closely agent tooling and supply chain risk have converged.
For a development team, this is the most immediately actionable reference: it's free, developer-focused, and written around concrete failure patterns rather than governance processes.
NIST's AI Risk Management Framework provides the broad structure for identifying and managing AI-related risk across an organization, and NIST SP 800-218A applies that structure specifically to secure development. SP 800-218A augments the existing Secure Software Development Framework (SP 800-218) with practices, tasks, and recommendations specific to AI model development throughout the software development lifecycle, and it was created to support a 2023 US Executive Order tasking NIST with exactly this kind of AI-specific companion resource. It's written to be useful to producers of AI models, producers of AI systems built on those models, and the organizations acquiring them, which makes it relevant whether you're building a model in-house or evaluating one from a vendor the same due-diligence lens covered in our governance, risk, and compliance guide.
CISA, the NSA's AI Security Center, the FBI, and international partners released joint guidance in May 2025 focused specifically on data used to train and operate AI systems, identifying data supply chain vulnerabilities, maliciously modified data, and data drift as the three primary risk areas. Its recommendations, such as provenance checks, digital signatures, and source authentication for third-party datasets, translate directly into the controls covered earlier in this guide.
MITRE ATLAS takes a different approach: rather than prescribing controls, it catalogs real adversary tactics and techniques against AI systems, modeled after the well-known MITRE ATT&CK framework for conventional IT. As of its most recent update, ATLAS documents tactics spanning reconnaissance through exfiltration and impact, and its 2025 expansion added AI Supply Chain Compromise as a named technique alongside dozens of case studies of real, documented incidents. Where OWASP and NIST tell you what to build, ATLAS tells you how an attacker actually thinks, which makes it a strong reference for red team scoping and threat modeling specifically.
These three standards work together to answer one question: can you prove what's in your AI system, and where it came from? SLSA (Supply-chain Levels for Software Artifacts), maintained by the OpenSSF, defines graduated levels of build integrity, from basic documented provenance at Level 1 up to cryptographically signed, tamper-resistant provenance at Level 3. An SBOM inventories conventional software components, while an ML-BOM or AIBOM extends that inventory to models, datasets, and their AI-specific dependencies. CycloneDX, an OWASP project, added ML-BOM support specifically for this purpose, and SPDX offers a comparable AI profile as an alternative format.
None of these three is a security control on its own. They're the documentation layer that makes every other control in this guide verifiable and auditable, which is exactly why they underpin the inventory and provenance practices covered in the best-practices section above.
ISO/IEC 42001, published in December 2023, is the first international, certifiable standard for AI management systems. It follows the same clause structure as ISO 27001, which makes it compatible with an organization's existing information security management system, but its focus is broader than security alone. It governs how an organization develops, provides, or uses AI responsibly, covering governance, risk management, and the AI lifecycle, with formal third-party compliance services available for organizations that want to demonstrate compliance to customers or regulators.
The EU AI Act operates at a different level entirely: it's binding law, not a voluntary standard. Its General-Purpose AI provider obligations under Articles 50 to 55 have applied since August 2025, while its high-risk system obligations, which include documentation and risk-management requirements that touch supply chain practices, were deferred by the EU's Digital Omnibus to December 2027 for standalone systems. Organizations pursuing ISO 42001 certification often find it maps usefully onto AI Act documentation requirements, but certification itself is not legal conformity with the Act.
Framework | What it covers | Who it's for |
|---|---|---|
OWASP Top 10 for LLM Applications | Concrete LLM application risks, including supply chain (LLM03) | Developers and AppSec teams building LLM applications |
NIST AI RMF / SP 800-218A | Risk management structure plus AI-specific secure development practices | Model producers, AI system builders, and buyers evaluating vendors |
CISA/NSA/FBI AI Data Security Guidance | Data supply chain, poisoning, and drift risks across the AI lifecycle | Organizations handling AI training and operational data, especially critical infrastructure |
MITRE ATLAS | Real-world adversary tactics and techniques against AI systems | Red teams, threat modelers, and security researchers |
SLSA / SBOM / ML-BOM (AIBOM) | Build provenance and component inventory for software and AI artifacts | Platform, DevOps, and ML engineering teams |
ISO/IEC 42001 | Certifiable AI management system covering governance and lifecycle risk | Organizations seeking formal, auditable proof of responsible AI governance |
EU AI Act | Binding legal obligations for AI providers and deployers, including documentation | Organizations building, selling, or deploying AI in or into the EU |

Managing AI supply chain risk comes down to four habits: know what you're running, verify it before you trust it, isolate it while it works, and watch it after it's live. Everything below expands one of those four.
You cannot secure a model you don't know you're running. An AI inventory lists every model, dataset, adapter, and AI-related dependency in use, along with its source, version, and owner. An AI Bill of Materials, or AIBOM, formalizes that inventory into a structured, machine-readable document.
The CycloneDX standard, maintained under the OWASP CycloneDX project, added ML-BOM support specifically to represent models, datasets, and their dependencies alongside ordinary software components. SPDX offers a comparable AI profile. Either format works. What matters is having one document that says, for every AI component in production, where it came from and what it depends on, so that when a new vulnerability or malicious package surfaces, you can find every system it touches in minutes, not weeks.
Provenance means being able to prove where a model, dataset, or package actually came from, not just trusting the name on the file. Verification means checking that proof before you use the artifact, not after something goes wrong.
The SLSA framework, maintained by the OpenSSF, defines graduated levels of build integrity for exactly this purpose. At its most basic level, SLSA requires a build platform to generate provenance describing how an artifact was built; at higher levels, that provenance is cryptographically signed so consumers can verify it hasn't been altered. Apply the same logic to models and datasets: prefer publishers who sign their releases, check hashes against a known-good value before loading a file, and treat an unsigned or unverifiable artifact as unproven, regardless of how popular its source looks.
Format choice is a security control, not a technical preference. Non-executable formats such as Safetensors cannot run code on load, while pickle-based formats can. Require Safetensors or an equivalent safe format wherever a model publisher offers one, and treat a pickle-only release as a signal to look for an alternative source.
Scanning still has a role, but only as a second layer. It catches known-bad patterns, not novel ones, and it isn't a substitute for format safety or provenance checks. Run a model scanner before deployment, understand what it does and doesn't detect, and never treat a clean scan as proof that an executable-format file is safe to load outside an isolated environment.
Every package your AI stack pulls in should come from a source you control, at a version you chose deliberately. Pin dependency versions with hash verification so an installer can't silently substitute a different file for the one you tested. Where possible, mirror critical packages into a private registry instead of pulling directly from public indexes at build time, and configure installers to trust only the indexes you specify.
This single control would have stopped the 2022 PyTorch nightly incident, where a public package index was allowed to override an internal one. It also blunts typosquatting, since a private mirror only contains packages your team explicitly approved.
Treat every stage that loads an untrusted artifact as if it could be compromised, because it might be. Run training jobs, model loading, and early testing in isolated environments with restricted network access and no direct path to production credentials or data stores. If a poisoned model or a malicious package does execute code, isolation limits what it can reach this is zero trust security applied specifically to AI pipelines.
The same principle applies to inference. An agent or model that only needs to read a database should never hold write access, and a sandboxed inference environment should not have an open path to internal systems it doesn't need for its task.
Before adopting a hosted model, an API, or a third-party dataset, get answers in writing: Where did the training data come from, and under what license? What security practices govern the vendor's own build and release process? How and when will they notify you of a security incident or a material change in model behavior? Can they provide signed provenance or an AIBOM for what you're buying?
A vendor that cannot answer these questions is not necessarily unsafe, but the gap is now your risk to manage, and it belongs in your own governance, risk, and compliance program rather than being assumed away.
Detection after deployment matters as much as prevention before it, because some compromises only surface once a model is running against real traffic. Log which model version served which request, monitor for unexpected changes in output patterns or resource usage, and watch training and inference environments for network activity that doesn't match their normal profile.
Several documented AI supply chain incidents were first noticed through an operational symptom, such as unusual CPU usage, rather than a security alert. Build an incident response path specifically for a compromised model or dataset: how you would identify every system using it, from your inventory, and how you would roll back to a verified version without extended downtime.
No control here is complete on its own. Format safety stops code execution but does nothing about a subtly poisoned dataset that produces a working, well-behaved model with a hidden bias. Provenance verification confirms an artifact came from where it claims to, but not that the original publisher's own pipeline was never compromised. Scanning catches known patterns and misses novel ones by design.
Teams tend to over-invest in scanning, treating a clean result as clearance rather than one input among several, and under-invest in isolation and inventory, which are less visible but catch different failure modes. The realistic goal is layered defense: provenance and format safety to reduce what gets in, isolation to limit the damage of what slips through, and monitoring to catch what isolation didn't stop. No single control here replaces the others.
An AI supply chain security checklist is a control-by-control reference for what to verify before trusting a model, dataset, or agent tool in production, covering inventory, provenance, format safety, dependency control, isolation, and monitoring. Teams that skip even one of these controls create the exact gap that incidents like JFrog's discovery of roughly a hundred malicious models on Hugging Face have shown attackers already exploit.
Use the table below as a working reference, not a one-time audit. Each row maps to a section covered earlier in this guide, so where a control needs more context, the corresponding best-practice section explains the reasoning behind it.
Control | Why it matters | Owner | Priority |
|---|---|---|---|
Maintain an AI inventory / AIBOM | You can't secure or respond to a compromise in a model you don't know you're running | ML/platform team | Critical |
Verify provenance before use | Confirms a model or dataset actually came from its claimed source | ML/platform team | Critical |
Require safe, non-executable model formats | Prevents a loaded model file from running arbitrary code | ML/platform team | Critical |
Scan models before deployment (as a second layer, not the only check) | Catches known-bad patterns that provenance and format checks miss | Security team | High |
Pin and hash-verify dependencies | Stops dependency confusion and typosquatting from silently swapping packages | DevOps/platform team | High |
Mirror critical packages internally | Removes reliance on public index precedence at build time | DevOps team | Medium |
Isolate training and inference environments | Limits blast radius if a loaded artifact is malicious | Platform/security team | Critical |
Apply least privilege to agent tools and credentials | Prevents one compromised tool from reaching everything the agent can touch | Security/engineering team | Critical |
Allowlist and version-pin agent tools and MCP servers | Stops an unreviewed or updated tool from silently gaining new behavior | Security team | High |
Require human approval on high-impact agent actions | Provides a last line of defense when a tool has already been compromised | Product/engineering team | High |
Run vendor and model due diligence before adoption | Surfaces gaps in a third party's own security and provenance practices | Procurement/security team | Medium |
Monitor deployed models for behavioral drift | Catches compromises that only surface after deployment | Security/ML ops team | High |
Maintain a compromised-model incident response path | Shortens response time when a rollback is genuinely needed | Security team | Medium |
Treat the "Critical" rows as the floor, not the ceiling. A team that implements only inventory, provenance, safe formats, isolation, and least privilege has closed most of the paths documented in real AI supply chain incidents so far, but the "High" and "Medium" rows are what keep that coverage from decaying as models, dependencies, and agent tools get added over time.

Testing AI supply chain exposure means finding out, before an attacker does, which of your models, datasets, and agent tools can be traced, verified, and trusted, and which can't. It takes more than a standard penetration test, because the attack surface includes artifacts a conventional test never touches, such as model files and training pipelines.
The first step is inventory, viewed through an attacker's eyes. List every model in production and where it was sourced, every dataset used for training or fine-tuning, every adapter layered onto a base model, and every tool, plugin, or MCP server an agent can call. For each one, ask who controls it and what would happen if it were replaced with something malicious. This is the same discipline behind attack surface management, extended to cover AI-specific assets alongside conventional infrastructure see our broader guide to what attack surface management involves and, for internet-facing exposure specifically, external attack surface management.
This map should extend beyond the deployed model to the pipeline around it: the build system that packages and deploys models, the credentials that pipeline holds, and the registries or hubs it pulls from. An attack surface map built only around the model in production misses exactly the stages, such as training and packaging, where several documented AI supply chain incidents have actually occurred.
AI-focused penetration testing and red teaming go beyond testing an application's inputs and outputs. They examine whether a model's provenance can actually be verified, whether unsafe file formats or unpinned dependencies are present in the pipeline, whether agent tools are over-permissioned relative to what they need, and whether a compromised or malicious component could reach production undetected.
Red teaming adds an adversarial layer on top of that assessment: simulating how an attacker would attempt to poison training data, smuggle a backdoored model into a pipeline, or abuse an over-permissioned agent tool, rather than only checking whether controls exist on paper. Femto Security's AI agentic pentesting and red teaming services are built around this distinction between testing whether controls exist and testing whether those controls hold under a real attempt to defeat them the same distinction covered in our guide to red teaming vs penetration testing.
For teams still deciding where to start, our overview of penetration testing methods, types, and tools explains how these approaches compare, and Femto's penetration testing services and vulnerability assessments cover the conventional layers an AI stack still depends on.
Much of the AI supply chain lives in code that never looks like "the model": training scripts, CI/CD pipelines that package and publish artifacts, and the integration code that connects agents to tools and credentials. A secure code review of these components checks for the issues that enable supply chain compromise directly, such as unpinned dependencies, overly broad build permissions, hardcoded credentials, and unsafe deserialization calls that load model files without restriction the same practices covered in what DevSecOps involves.
This work is most effective paired with the inventory and testing steps above: code review finds the flaws in how a pipeline is built, while testing confirms whether those flaws are actually exploitable. A supply chain compromise with real reach illustrates the stakes. When attackers compromised the Ultralytics Python package's build pipeline in December 2024 through a known GitHub Actions script injection, the library had nearly 60 million downloads, and the malicious code was inserted after the code review step had already passed, so the source repository looked clean while the published package did not. Femto Security's secure code review service covers this layer specifically for ML pipelines and AI integrations, alongside conventional application code.
AI supply chain security carries different regulatory weight depending on where a business operates: GCC regulators are actively building AI-specific guidance and data protection enforcement, while the EU and US already have binding obligations that touch supply chain transparency directly. A multinational buyer needs to satisfy the strictest applicable regime, not the most convenient one.
The UAE governs data used in AI systems primarily through Federal Decree-Law No. 45 of 2021 on the Protection of Personal Data, in force since 2 January 2022, which covers the processing of personal data belonging to anyone in the UAE regardless of where the controller or processor is located see our full breakdown of the UAE Data Protection Law. National AI guidance, including the UAE AI Ethics Principles and Guidelines, remains advisory rather than a standalone binding AI law. DIFC's Data Protection Law adds a specific provision, Regulation 10, addressing the processing of personal data through autonomous and semi-autonomous systems, which directly touches how AI systems in that free zone must handle data.
Government and public-sector entities exploring national or on-premises AI deployments face an additional layer of consideration: our guides to what is AI sovereignty and securing sovereign AI infrastructure cover the security implications specific to government and national AI programs, an area Femto also supports directly through our government practice. VARA-regulated entities in Dubai's virtual asset sector layer additional, sector-specific cybersecurity requirements on top of this, covered in Femto's VARA compliance guidance and our broader look at how VARA compliance is redefining cybersecurity standards for UAE virtual asset businesses. For a full picture of the wider regulatory landscape, see our overview of UAE cybersecurity regulations.
Saudi Arabia's Personal Data Protection Law is now in full force and covers automated processing, while the Saudi Data and Artificial Intelligence Authority (SDAIA) has issued a Generative AI Guideline for Government and a separate one for public use. Neither guideline is binding regulation in the way the PDPL is, but SDAIA is the central authority shaping how AI governance in the Kingdom develops, and its guidance is a reasonable signal of where enforcement expectations are heading.
For both jurisdictions, the practical implication for AI supply chain security is the same: data protection law already applies to whatever personal data flows through training pipelines and models, regardless of whether a dedicated AI law exists yet, so provenance and data-handling controls double as compliance controls today.
For GCC organizations in regulated sectors, most notably Web3 and fintech, AI supply chain risk intersects directly with existing compliance obligations. A VARA-regulated exchange using an AI model to screen transactions, for instance, needs to be able to demonstrate the provenance and integrity of that model to satisfy the same operational resilience expectations that apply to its other systems, including the VASP compliance roadmap these businesses already follow. Where an AI system touches smart contract logic or on-chain infrastructure directly, this also overlaps with smart contract security auditing. This is not a separate compliance track; it's an extension of controls these organizations already maintain, applied to a newer category of dependency.
For organizations selling into or operating in the EU, the AI Act's General-Purpose AI provider obligations under Articles 50 to 55 have applied since 2 August 2025 and are unaffected by the Digital Omnibus on AI, which entered into force on 27 July 2026 and deferred the Act's separate high-risk system obligations, originally due 2 August 2026, to 2 December 2027 for standalone Annex III systems and 2 August 2028 for AI embedded in regulated Annex I products. Article 50 transparency obligations for deployers also remain on their original August 2026 schedule, with only the generative-AI watermarking requirement under Article 50(2) given a short grace period to December 2026. Supply chain transparency expectations tied to model documentation and provenance sit within the obligations that are already in force, not the ones that were deferred, so this is not ground multinational buyers can treat as settled or delayed.
In the US, CISA, the NSA's AI Security Center, and the FBI jointly released best-practice guidance in May 2025 identifying data supply chain vulnerabilities as one of three primary AI data security risks, alongside maliciously modified data and data drift, and recommending controls such as provenance checks, digital signatures, and source authentication for third-party datasets. This guidance is non-binding but reflects the baseline US federal agencies now expect from organizations handling AI data, particularly in critical infrastructure or government-adjacent sectors the kind of environment covered under Femto's enterprise cybersecurity and government practices.

Most AI supply chain failures trace back to a handful of avoidable habits, not sophisticated attacks. Four stand out across the incidents and controls covered in this guide.
Trusting a model because it's popular. Download counts and star ratings measure adoption, not security. A widely used model on a mainstream hub can still be malicious: JFrog's security researchers scanned models on Hugging Face and found roughly a hundred with malicious functionality, some able to execute code and give attackers a persistent backdoor, hosted on a mainstream platform rather than a fringe site. Popularity is not provenance, and it should never substitute for verifying a specific file's source and integrity.
Scanning code but never the model file. Teams with mature application security programs often extend that discipline to source code and dependencies while treating the model file itself as opaque and unreviewable. That's backwards. The model file is frequently the component most capable of running code on load, particularly when it uses an executable format like pickle, and it deserves the same scrutiny as any other executable artifact, not less.
Treating agent tools as trusted. An MCP server or plugin is third-party code the moment it's installed, regardless of how it's described or how long it's been in use. The postmark-mcp package on npm behaved correctly for its first 15 versions before a single added line began quietly copying every outgoing email to its publisher, which shows that a clean version history is not the same as a tool that stays safe.
Having no inventory of which models are running where. Without an inventory, a newly disclosed vulnerability or a report of a malicious model means asking every team individually whether they're affected, which is slow and unreliable under time pressure. An AIBOM or equivalent inventory turns that into a lookup that takes minutes instead of days, and it's the single control that makes every other one in this guide actionable at speed. This same gap is a recurring theme in broader cyber security threats research: teams are consistently better at securing what they've inventoried than what they haven't.
AI supply chain security comes down to four habits: inventory what you're running, verify it before you trust it, isolate it while it operates, and test it before an attacker does. None of these controls is exotic, and none requires rebuilding your stack from scratch. Together, they close the gap that let researchers find roughly a hundred malicious models sitting on a mainstream public hub, waiting for someone to download and run them.
If you want a clear picture of where your own AI stack stands against these controls, AI-focused penetration testing and security testing is the direct next step.
No. A supply chain attack compromises a component, such as a model, dataset, or package, before it's deployed. Prompt injection manipulates a model's behavior through its input after it's already running. Verifying provenance stops the first; it does nothing against the second, because the model itself may be entirely genuine.
Yes. Some model file formats, particularly those based on Python's pickle serialization, can execute code the moment they're loaded, so a malicious file can behave like a program rather than passive data. JFrog found roughly a hundred models with malicious functionality on Hugging Face, some capable of giving attackers a reverse shell, which is why security teams recommend safe, non-executable formats such as Safetensors wherever available.
An AI Bill of Materials is a structured record of the models, datasets, and dependencies an AI system relies on, including their source and version. It's the foundation for knowing what you're running and responding quickly if one of those components is later found to be compromised. Any organization running more than a handful of AI components in production benefits from maintaining one, and it's increasingly expected as a baseline control rather than an advanced practice.
Not inherently. Open-source models can be inspected and self-hosted, which gives you more control, but that control is only a security benefit if you actually verify provenance and format safety yourself. Commercial APIs shift that verification burden to the vendor, which is convenient but means your security now depends on practices you can't directly inspect. Neither option is automatically safer; each shifts the responsibility differently.
Confirm the model came from a publisher you can identify and trust, check any cryptographic signature or hash the publisher provides against the file you downloaded, and prefer sources that publish build provenance, such as SLSA-style attestations, rather than an unsigned file with no traceable origin. Where none of this is available, treat the model as unverified and isolate it accordingly.
The Act's high-risk system obligations include requirements around data governance, documentation, and risk management that touch supply chain practices, and its General-Purpose AI provider obligations, in force since August 2025, include transparency requirements relevant to model provenance. It is not a supply-chain-specific security regulation, but its documentation and governance requirements overlap meaningfully with the controls described in this guide.