Data and Model Poisoning (LLM04)
An attacker injects malicious data into training, fine-tuning, or RAG-corpus content to alter model behavior in their favor, often subtly and often persistently.
- Rank
- LLM04 of 10
- In the check
- Cited by 1 of the 16 questions
What it looks like in practice
Three shapes this risk takes in real deployments.
Example 1
Poisoning a public web corpus that the target model later trains on.
Example 2
Inserting backdoor-trigger content into a fine-tuning dataset.
Example 3
Poisoning a RAG corpus with content designed to bias outputs on specific queries.
Controls that close it
These count toward the Model dimension of the check.
Provenance tracking for training data
Adversarial testing for backdoors
RAG corpus content review
Anomaly detection on training-data ingest
Where the check cites it
The AI Posture Check cites OWASP LLM Top 10, including this entry, when placing you at Crawl, Walk, Run or Sprint.
| Question the check may ask | Dimension | Citation |
|---|---|---|
| For models you build or fine-tune, are there controls against theft and extraction? | Model | OWASP LLM04, MITRE ATLAS |
Other frameworks the check cites
Score yourself against this framework.
Five questions, each citing its source. You get your stage, your place on the chart and the one move that matters next.
- A few minutes for most people
- Free, from CWS
- Your stage, the chart and the next move, by email or live with an engineer