Research and methodology
Reproducible methods for evaluating LLM security controls. Each page states the corpus, the metric definitions and the procedure in enough detail to be re-run against your own traffic and your own model. Where we have not measured something ourselves, the page says so.
How to benchmark an LLM guardrail on your own traffic
A reproducible method for measuring an LLM guardrail: corpus construction, metric definitions, latency measurement, and the six mistakes that void the result.
· 9 min readbenchmarkmethodologyevaluationprompt injectionfalse positives