AI Integrity: Defending against backdoors and secret loyalities
2026
Open publication workspace · Sign in to read the full PDF.
AI-generated summary
1) This report is crucial for understanding and defending against the growing threat of AI integrity attacks, which could compromise national security and critical infrastructure.
2) * Explores the evolving landscape of AI integrity threats, including model sabotage and subversion.
* Details four complementary defense approaches: AI infrastructure security, data auditing, model auditing and evaluation, and AI control.
* Proposes four policy recommendations for the US government to strengthen AI integrity, focusing on red team exercises, NIST frameworks, information sharing, and ARPA programs.
3) AI integrity, backdoors, secret loyalties, AI security, national security
Check the original publication for accuracy and context.