Skip to main content

Research and methodology

Reproducible methods for evaluating LLM security controls. Each page states the corpus, the metric definitions and the procedure in enough detail to be re-run against your own traffic and your own model. Where we have not measured something ourselves, the page says so.


  • How to benchmark an LLM guardrail on your own traffic

    A reproducible method for measuring an LLM guardrail: corpus construction, metric definitions, latency measurement, and the six mistakes that void the result.

    · 9 min read
    benchmark
    methodology
    evaluation
    prompt injection
    false positives