# KNOWN_LIMITATIONS — v1 live

- **Small validation run:** 390 rows total; only **10** SchemaBench prompts per (model, mode), and **8** rows per seed benchmark file for XSTest/OR-Bench seeds.
- **Seed benchmarks only:** `xstest_seed` / `orbench_seed` are bundled **subsets**, not full official corpora unless separately downloaded.
- **Deterministic scoring:** Metrics reflect rule-based judges; edge cases and nuanced safety may be mis-scored.
- **Local models:** `llama3.1:8b`, `mistral`, `gemma3:4b` are not frontier models; behavior may not generalize.
- **Hardware/runtime:** Host may be GPU-limited (e.g. RTX-class consumer cards); latency and timeouts can bias which runs succeed (here: no transport failures).
- **No strong claims:** Do not treat aggregate deltas as proven safety improvements; **human review** remains required (`REVIEW_QUEUE.csv`).
