LLMSecTest¶
These pages describe v0.3.0
The site is built from main, so it can describe a version newer than the one
pip install llmsectest gives you. Run llmsectest --version to see what you have.
The changelog says what arrived when.
Pytest-native security testing for LLM applications, mapped to the OWASP LLM Top 10 (2025).
LLMSecTest tests applications that use an LLM (a system prompt, guardrails, RAG, and tools around a model) rather than bare models, for the security risks in the OWASP Top 10 for LLM Applications, and emits SARIF / HTML / JSON / Markdown reports that drop straight into CI/CD.
pip install llmsectest
# point it at your running app and test it black-box
llmsectest --target app:http://localhost:8000/chat
A failing probe is a finding, so a non-zero exit fails your pipeline when your app is vulnerable.
Why¶
LLM applications fail in ways an ordinary test suite never looks for. Prompt injection,
sensitive-data disclosure and unsafe output handling are common, well-documented failure modes. A
team shipping an assistant in finance or healthcare has no open, CI-ready way to check one against a
recognized standard. LLMSecTest is that check: MIT-licensed, fully open-source, built on pytest so
it fits the tools developers already use.
What it tests¶
The OWASP LLM Top 10 spans two testing modalities. LLMSecTest is honest about which apply to a given
target, run llmsectest --check to see the live map.
| Category | How it's tested | |
|---|---|---|
| LLM01 | Prompt Injection | black-box, your app endpoint |
| LLM02 | Sensitive Information Disclosure | black-box (or white-box) |
| LLM03 | Supply Chain | white-box, your deps (--repo) |
| LLM04 | Data and Model Poisoning | white-box, your model files (--model-scan) |
| LLM05 | Improper Output Handling | black-box (or white-box) |
| LLM06 | Excessive Agency | black-box (or white-box) |
| LLM07 | System Prompt Leakage | black-box, prompt extraction |
| LLM08 | Vector and Embedding Weaknesses | black-box, your RAG app (--app-canary / --app-rag-poison); white-box, your vector store (--vector-store) |
| LLM09 | Misinformation | black-box, nonexistent-entity confabulation |
| LLM10 | Unbounded Consumption | black-box, flood / output amplification |
→ Getting started · Test your running app · Red-team your defense · OWASP coverage · API reference
Status
Pre-alpha (active grant development). All 10/10 OWASP LLM Top 10 (2025) categories ship today:
black-box probes for LLM01/02/05/06/07/09/10, white-box scanners for LLM03 (supply chain)
(--repo) and LLM04 (data and model poisoning) (--model-scan), and LLM08 (vector &
embedding weaknesses) both ways, black-box RAG probes plus an offline embedding-inversion
exposure scan of a persisted store (--vector-store). What remains to do is depth rather than breadth. Coverage claims here always
match what the tool does, see llmsectest --check.
Which edition and why not the newer one
A 2026 edition of the list came out on 3 August 2026. This tool implements 2025 and says so on every surface. Read side by side, the two lists hold the same ten risks: nothing was added, dropped, merged or split. Eight categories change number and LLM07 System Prompt Leakage becomes LLM08 Hidden Context Exposure with a wider remit that now covers retrieved policy text and tool schemas. Every report and baseline published here carries 2025 numbers, so a silent renumber would change what already-published records mean, and OWASP's own per-category pages still carry 2025. Both change before this does.
Funding¶
LLMSecTest is funded by the German Federal Ministry of Research, Technology and Space (BMFTR) through the Prototype Fund, funding code (Förderkennzeichen) 16IS26S10. The funding guideline is implemented by the Open Knowledge Foundation Deutschland; the project agency is VDI/VDE-IT.