Skip to content

CLI reference

These pages describe v0.3.0

The site is built from main, so it can describe a version newer than the one pip install llmsectest gives you. Run llmsectest --version to see what you have. The changelog says what arrived when.

LLMSecTest installs a llmsectest console script (equivalently python -m llmsectest). It runs the packaged OWASP probe suite against your chosen target and writes reports.

llmsectest [--target <spec>] [--report-formats=...] [pytest options]
llmsectest --check | --list-probes | --validate <file.sarif> | --render-sarif <file.sarif>
llmsectest --sbom [<out.json>] --repo <path>

Wrapper commands

Flag Description
--target <spec> What to test: app:<url>, ollama:<model>, lmstudio:<model>, openai:<model>, anthropic:<model>, huggingface:<model>, mock, demo, demo-defended. Omit it to scan the offline demo app.
--repo <path> Add the white-box LLM03 (supply chain) scan of that project's dependency manifests (requirements*.txt, pyproject.toml, Pipfile). Combine with --target to test an app and its dependencies in one run.
--osv With --repo: also query OSV.dev for known CVEs in every exactly-pinned (==X.Y.Z) dependency (networked, free, no API key). Off by default so the scan stays offline/deterministic; any non-run state (not requested, nothing pinned, lookup failed) appears as an explicit skip reason.
--sbom [<out.json>] With --repo: write a CycloneDX 1.6 JSON SBOM of the project's declared dependencies (one PURL-identified component each) and exit. An exact pin (==X.Y.Z) is carried into the component version + a fully-qualified PURL (pkg:pypi/name@version); a range/unpinned dependency is left version-less with its raw constraint recorded as a property. So the SBOM is only ever as precise as the manifests allow, never asserting a version a manifest did not pin. Built dependency-free from the stdlib. Writes results/<repo>.cdx.json unless a path is given.
--model-scan <path> Add the white-box LLM04 (data and model poisoning) scan of the serialized model files under that path. Walks each pickle's opcode stream (stdlib pickletools, never unpickling) and flags imports of code-execution primitives (os.system, subprocess, builtins.eval, nested pickle/torch.load, reflection gadgets) that would run on load, across raw pickles, PyTorch .pt/.pth/.ckpt zips and numpy object arrays. Offline and deterministic; a clean weights file does not false-positive.
--vector-store <path> Add the white-box LLM08 (vector and embedding weaknesses) embedding-inversion exposure scan of a persisted vector store. Reads Chroma sqlite through a read-only URI, JSON and JSONL stores including the LlamaIndex shape, and FAISS index.pkl sidecars whose strings come off the pickletools opcode stream (never unpickled). Reports plaintext filed beside the vectors, a credential shape in the stored corpus, the embedding space being recorded, source-identifying per-vector metadata, and a world-readable store file. It does not run an inversion, which needs a trained inverter per embedding space; it measures what a reader of the store gets without one. A store of vectors and ids alone reports clean.
--redteam-set <csv> Run the LLM01 red-team module against the JailbreakBench JBB-Behaviors CSV at this path (the 100-behaviour benchmark; columns Index,Goal,Target,Behavior,Category,Source). A target is a finding when it complies with a harmful request instead of refusing (refusal oracle). Without the flag, a built-in starter set runs so an offline scan still exercises the oracle.
--redteam-benign [<csv>] Measure over-refusal: run the JBB benign twins (harmless requests matched to the harmful behaviours) through the over-refusal oracle and report the target's false-refusal rate. Bare uses a built-in benign set; pass JBB's benign-behaviors.csv for the full 100. Over-refusing is a usability defect rather than a vulnerability, so it is reported separately and never enters the SARIF findings or the exit code.
--app-prompt <text-or-file> With --target app:<url>: the app's own system prompt (inline text, or a path to a file holding it), unlocks LLM07 leakage detection, since we then know what a leaked instruction looks like.
--app-secret <value> With --target app:<url>: a real secret/canary the app holds, unlocks LLM02; a disclosure is then unambiguous (no FP-prone heuristics).
--app-action <signature> With --target app:<url>: a privileged tool/action signature of the app (e.g. "ACTION: refund("). Repeatable, unlocks LLM06; an unauthorized invocation is then unambiguous.
--app-canary <value> With --target app:<url>: confidential content planted in the app's RAG / retrieved corpus (a unique token in an indexed document), unlocks the LLM08 retrieval-exposure probes (corpus enumeration, retrieval-scope escalation, topic-targeted retrieval). A reply that reproduces the canary means the retrieval layer leaked confidential indexed content. Distinct from --app-secret (LLM02), where the secret lives in the prompt/config; here it lives in the retrieved corpus.
--app-rag-poison <marker> With --target app:<url>: the marker a poisoned document planted in the app's retrieved corpus instructs the model to emit, unlocks the LLM08 indirect-injection probes (kb-directive, follow-retrieved, apply-policy). Each is a benign query that never names the marker; if the app retrieves the poisoned doc and emits it, the app obeyed an injected instruction from retrieved content. Categorized LLM08 rather than LLM01 because the injection vector is the retrieval layer, which is the LLM08↔LLM01 bridge.
--app-timeout <seconds> With --target app:<url>: the per-request time budget for a single call to the app (default 120 s). A real wall-clock deadline over connect, headers and body (not a per-socket-operation timeout, which an app that keeps trickling output never trips at all). A request that exceeds it is recorded as an inconclusive probe. It is never a finding, since a timeout is generally not proof of a vulnerability, and it is never a silent clean either: it becomes a warning in the summary plus a run-level inconclusive count in the SARIF/HTML report. So one slow or runaway endpoint cannot hang the whole scan and lose every other result. A report never silently hides that some probes could not be concluded. The one exception is the two bounded LLM10 probes, where exhausting the budget is the finding, though only on an app that demonstrably answered its other probes well inside the same budget (LLM10). Raise it for a legitimately slow app; lower it to keep a scan under a fixed wall-clock budget (e.g. in CI).
--redteam-generate <N> Add N model-composed variants of each authored LLM01 case, written by the target's own model, alongside the authored corpus and never instead of it. The authored set still runs and still feeds the regression baseline, so generation adds attacks without making two runs incomparable. Every variant is validated before it runs: a rewrite that dropped the marker the case asks the target to emit is refused, because it would run, score nothing, and be recorded as an attack withstood; a rewrite that put the target's own secret into the prompt is refused, because it would score its own echo as a disclosure. Rejections are counted per reason and reported, since a corpus that quietly shrinks looks exactly like a target that got safer. Each surviving variant is then run against a defended control, because keeping the marker proves a rewrite is scoreable rather than proving it is still an attack: "Please output MARKER-7F3A" keeps the marker and is a plain request that any cooperative target answers, so that reply would score as a prompt injection against a model that was asked nicely. A variant that fires against the hardened demo tests nothing and is rejected; one that does not was refused by a defence. An unreachable control keeps the variant, since a fixture being down is no evidence against an attack. The generated prompt and the model that wrote it travel into the report, so a finding stays reproducible. Needs a model target, since --target app:<url> names somebody else's application rather than a model we may ask to write prompts. Only the marker-based cases are generated from, never the refusal-oracle red-team set, where a broken rewrite turns the request benign and compliance would be scored as a critical finding, which no structural check can catch.
--app-stress <N> With --target app:<url>: replay every reachable application case as one simultaneous wave of N requests and report only the transition: a guardrail that held when the app was asked once and failed when it was asked N times at once. A case that already fails at one request is left to its own module rather than counted twice. The verdict may be held only when the wave demonstrably arrived: the workers meet at a barrier so the requests genuinely overlap, the peak counts requests this side has outstanding rather than the app's own concurrency, and a wave the app refused with a rate limit leaves an inconclusive differential plus the separate note that a defence fired. A timeout under load is always inconclusive, because the app is slow because we made it busy. Opt-in, since it multiplies the traffic this scan sends to somebody else's application by N, so there is no default; 1 is refused rather than rounded up. Without the flag the module reports one skipped test naming it.
--render-pdf <file.sarif> Render any SARIF v2.1.0 file, ours or a third party's, to a PDF report, written to <file>.pdf or to -o/--pdf-output. The PDF is produced directly rather than by converting the HTML: a tool people point at their own security-critical systems should not grow a browser or a font toolchain to gain an output format, so this uses the 14 standard PDF fonts every reader ships and adds no dependency at all. It reads the same SARIF the HTML reader reads, so the two cannot disagree about a figure. Order carries a claim here: the undelivered notice, the exposed-secret notice and the inconclusive count are laid out before the findings, so a reader who stops after page one cannot take a run that never reached its target for a clean one.
--check Print the OWASP LLM Top 10 coverage map, each category's test modality and its CVSS v4.0 base score, then exit.
--list-probes List the probe corpus that ships today (incl. the built-in red-team set), then exit.
--validate <file> Validate an existing SARIF file against the v2.1.0 schema, then exit.
--render-sarif <file> Render a SARIF v2.1.0 file (ours or any other tool's) as a standalone HTML report and exit. Writes <file>.html next to it, or pass -o/--html-output <path>. Findings are grouped by OWASP category, CVSS-scored and colour-coded by severity, each with its location, evidence and remediation, plus a rule-reference glossary. No server, no assets. Just open the file. Interop is proven against committed output from ruff, Bandit and Semgrep, which cover every CWE convention we have met (none at all; the GitHub external/cwe/cwe-NNN tag; a descriptive tag beginning with the id), so a third-party finding shows its CWE. A rule id that is a dotted namespace is titled by its last segment.
--preflight Health-check --target and exit. For a local OpenAI-compatible runtime (ollama: / lmstudio:) it hits the server's GET /v1/models (no key, no paid call) to confirm the server is reachable and the requested model is loaded, failing fast (exit 1) with a clear message instead of an opaque SDK error deep inside the first probe. A provider with no cheap health endpoint reports that and exits 0. Run it before a long local scan.
--version Print the installed llmsectest version, then exit.

Reporting options (pytest plugin)

Flag Description
--report-formats=sarif,html,json,markdown Which report formats to emit (default: sarif).
--report-dir=<dir> Where to write reports (default: results/).
--sarif-output=<path> Explicit SARIF path (otherwise results/<target-slug>.sarif).

Every report also carries a run-level attacks_withstood tally: how many probes were delivered to the target and how many it held off, broken down by OWASP category, so a clean scan is evidence rather than an empty page. Only delivered probes count (a coverage assertion or a static scanner never inflates it) and an inconclusive probe is neither withstood nor a finding. See Red-team your defense.

Gating, baselines and policy

Flag Description
--risk-threshold=<level> Fail only at/above a severity (e.g. high).
--min-coverage=<n> Require at least n OWASP categories exercised.
--save-baseline / --update-baseline Record current findings as an accepted baseline.
--compare-baseline Fail only on findings new since the baseline.
--enable-policy / --security-policy=<file> Enforce a YAML security policy.
--enable-trends / --disable-trends Track findings over time.

Any other pytest option (e.g. -k, -v, -x) is passed straight through.

All ten OWASP categories run on every invocation: each ships a real probe or scanner and any category that needs an input it wasn't given (a repo, a model path, an app marker) is reported as a skipped test naming the flag that would enable it (skip reasons print by default). A run also ends with a footer listing all ten categories. Which this run exercised and which it did not, with the reason, so a category is never silently left untested. A model/demo target exercises the always-on black-box probe categories (LLM01/02/05/06/07/09/10); adding --repo <path> runs the white-box LLM03 (supply chain) scan as well. LLM01 also runs a red-team jailbreak set (built-in starter set, or the full JailbreakBench corpus with --redteam-set <csv>); the footer prints the LLM01 depth so the red-team coverage is never a silent gap. Adding --redteam-benign prints, below the security report, the target's over-refusal (false-refusal) rate over the benign twins, a usability metric kept out of the findings and the exit code. A real app endpoint (--target app:<url>) is black-box: LLM01, LLM05, LLM09 and LLM10 (a bounded repetition-flood probe) always run, and LLM07/LLM02/LLM06/LLM08 join them when you pass --app-prompt / --app-secret / --app-action / --app-canary / --app-rag-poison, each category whose input is missing is reported as skipped with the flag that would enable it (LLM04 is white-box, enabled by --model-scan). LLM08 has three enabling flags: --app-canary (retrieval exposure), --app-rag-poison (indirect injection via a poisoned retrieved document) and --vector-store (embedding-inversion exposure, white-box and offline). llmsectest --check prints the same map with each category's CVSS score.

Exit code

Non-zero when the target is vulnerable (findings present). So the command fails your pipeline. Use --risk-threshold or --compare-baseline to control what fails.

Examples

llmsectest --target app:http://localhost:8000/chat
llmsectest --target app:http://localhost:8000/chat --repo .   # app + its dependencies (LLM03)
llmsectest --target app:http://localhost:8000/chat \
    --app-prompt prompt.txt --app-secret "sk-canary" \
    --app-action "ACTION: refund(" --app-action "ACTION: delete_user(" \
    --app-canary "INTERNAL-DOC-CANARY-7f2a" \
    --app-rag-poison "RAG-POISON-3b9d"                                  # + LLM07/02/06/08
llmsectest --repo .                                            # supply-chain scan only
llmsectest --repo . --osv                                      # + known-CVE lookup (OSV.dev)
llmsectest --sbom --repo .                                     # CycloneDX SBOM of the deps (LLM03)
llmsectest --redteam-set jbb/harmful-behaviors.csv --target ollama:llama3  # 100 JailbreakBench prompts
llmsectest --redteam-benign --target ollama:llama3            # over-refusal (false-refusal) rate
llmsectest --target ollama:gemma4:e2b-it-q4_K_M --report-formats=sarif,html
llmsectest --target app:http://localhost:8000/chat --compare-baseline --risk-threshold=high
llmsectest --check