Interactive evidence dashboards for 13 ML security repositories. Selected benchmark results, detection metrics, and gap analysis from committed artifacts and test runs.
Inline MCP/JSON-RPC security gateway with prompt-injection patterns, capability checks, exfiltration signals, audit logging, rate limiting, and SIEM-oriented telemetry.
Non-executing model-artifact scanner for pickle risk, provenance gaps, impersonation, suspicious loaders, serialization formats, and supply-chain indicators.
Offline prompt-injection detector evaluated on two datasets. F1=0.9714 on grouped held-out templates; novel-phrasing OOD F1=0.7188. Gap is documented, not papered over.
5-layer MCP tool-call firewall. 37/37 blocked on the fixed bundled self-test catalog; this is a regression signal, not a real-world detection rate.
Supply chain attack scanner for HuggingFace. Internal fixture suite: 12/12 detected. Not a general detection rate.
Prompt-injection detector. F1=0.9714 grouped held-out / 0.7188 novel-phrasing OOD. Gap documented.
FGSM/PGD/C&W attacks + Randomized Smoothing defense. CIFAR-10 SmallCNN measured: clean 71.82%, PGD-20 0.00% at ε=8/255.
Membership inference + model extraction. Adult/OpenML DirectMIA mean AUC 0.557 across 5 seeds.
Synthetic label-flip benchmark: spectral detector average F1 0.6142 versus 0.1531 for the feature-space ensemble across 5/10/20% contamination.
ICS/OT research prototype using NASA C-MAPSS. FD001 validation records F1=0.54; committed inference benchmark records p99=13.242 ms over 500 samples. No production deployment is claimed.
MITRE ATT&CK v19 data models. Current repository verification reports 156 passing tests across 108 test functions in 11 files.
5-source telemetry pipeline mapped to ATT&CK. The repository documents 42 rules but currently labels that count unverified; CI configuration sets a 53% minimum coverage gate.
Static IAM policy linter for AI agent workloads. Checks wildcards, PassRole, trust boundary issues.
Authenticated integration control plane. Latest verified CI: 62 unit tests, 35 integration tests, 51.76% statement coverage. Local Compose uses contract stubs; no operated production deployment is claimed.
Benchmark runner with HMAC-SHA256 signed evidence. Smoke tests only — not full production benchmarks.
Selected-project status summary from static evidence snapshots. Visualization only — not the full repository inventory or live monitoring.