Interactive Evidence Dashboards

ML Security
Interactive Evidence

Interactive evidence dashboards for 13 ML security repositories. Selected benchmark results, detection metrics, and gap analysis from committed artifacts and test runs.

Honesty note:  These are interactive evidence dashboards, not live security monitoring. Most data is static benchmark output embedded in HTML. The MCP gateway is the only one with real-time capability when a server is running. Weak results are shown as-is — limitations are explained, not hidden.
All 13 repositories

MCP Security Gateway Monitor

5-layer MCP tool-call firewall. 37/37 blocked on the fixed bundled self-test catalog; this is a regression signal, not a real-world detection rate.

HF Model Provenance Scanner

Supply chain attack scanner for HuggingFace. Internal fixture suite: 12/12 detected. Not a general detection rate.

LLM Redteam Framework

Prompt-injection detector. F1=0.9714 grouped held-out / 0.7188 novel-phrasing OOD. Gap documented.

Adversarial ML Lab

FGSM/PGD/C&W attacks + Randomized Smoothing defense. CIFAR-10 SmallCNN measured: clean 71.82%, PGD-20 0.00% at ε=8/255.

Model Privacy Attacks

Membership inference + model extraction. Adult/OpenML DirectMIA mean AUC 0.557 across 5 seeds.

Dataset Poisoning Detector

Synthetic label-flip benchmark: spectral detector average F1 0.6142 versus 0.1531 for the feature-space ensemble across 5/10/20% contamination.

PulseNet RUL Forecasting

ICS/OT research prototype using NASA C-MAPSS. FD001 validation records F1=0.54; committed inference benchmark records p99=13.242 ms over 500 samples. No production deployment is claimed.

ATT&CK v19 Core

MITRE ATT&CK v19 data models. Current repository verification reports 156 passing tests across 108 test functions in 11 files.

Attack Detection Engine

5-source telemetry pipeline mapped to ATT&CK. The repository documents 42 rules but currently labels that count unverified; CI configuration sets a 53% minimum coverage gate.

AWS Agent Identity Guard

Static IAM policy linter for AI agent workloads. Checks wildcards, PassRole, trust boundary issues.

Unified ML Security Platform

Authenticated integration control plane. Latest verified CI: 62 unit tests, 35 integration tests, 51.76% statement coverage. Local Compose uses contract stubs; no operated production deployment is claimed.

ML Security Benchmark Suite

Benchmark runner with HMAC-SHA256 signed evidence. Smoke tests only — not full production benchmarks.

ML Security Command Center

Selected-project status summary from static evidence snapshots. Visualization only — not the full repository inventory or live monitoring.