dashboard_server not running.
Start it to run live evals: python dashboard_server.py
Metrics below are from the last committed benchmark run.
Run evaluation
~5–15s · offline, no API calls
Evaluation results
● Live result
0.97
F1 — grouped held-out
deepset/prompt-injections · committed benchmark
0.7188
F1 — novel-phrasing OOD
Out-of-distribution · committed benchmark
Precision / Recall / F1 by dataset
Confidence score distribution