RL-PWB-1 · the packs
Benchmark packs
The confident-wrong failures that ship, one industry at a time.
The RavnLab Plausible-Wrong Benchmark (RL-PWB-1) tests for the failure that actually reaches customers: the answer that reads perfectly and is false. Below are its four vertical cuts, each anchored by domain practitioners and each independently citable. We ran the full benchmark on five open models in public - see the run.
The four packs
RL-PWB-1-LEGAL · anchored by vetted attorneys
Five ways legal AI is confidently, expensively wrong. 5 expert-anchored cases.
RL-PWB-1-CLIN · anchored by practicing clinicians
Five ways clinical AI reads safe and is dangerous. 5 expert-anchored cases.
RL-PWB-1-FIN · anchored by CPAs and CFPs
Five ways financial AI is fluent and false. 5 expert-anchored cases.
RL-PWB-1-ENG · anchored by PEs and senior builders
Five ways field-decision AI cracks walls. 5 expert-anchored cases.
We build a pack from your domain's real cases and grade your AI against it, before you ship.
Contact sales