Researchers Have Ranked AI Models Based on Risk, and Found a Wild Range
AIR-Bench applies a regulation- and policy-based risk taxonomy to model evaluation.
AIR-Bench applies a regulation- and policy-based risk taxonomy to model evaluation.
Our study found that even benign fine-tuning can weaken safety alignment.
Results and backdoor-mitigation benchmarks from the IEEE Trojan Removal Competition, which I chaired.
An inaugural Amazon fellowship at Virginia Tech for research on AI risks and safeguards.
Citations to coauthored research, not institutional endorsements.
The initial public draft cites our fine-tuning safety study in its discussion of model safeguards and access.
Cites our work on removing learned backdoors (I-BAU).