Skip to content
dotdock

AI Integrity: Defending against backdoors and secret loyalities

2026

Publication cover

Open publication workspace · Sign in to read the full PDF.

AI-generated summary

1) This report is crucial for understanding and defending against the growing threat of AI integrity attacks, which could compromise national security and critical infrastructure.
2) * Explores the evolving landscape of AI integrity threats, including model sabotage and subversion.
* Details four complementary defense approaches: AI infrastructure security, data auditing, model auditing and evaluation, and AI control.
* Proposes four policy recommendations for the US government to strengthen AI integrity, focusing on red team exercises, NIST frameworks, information sharing, and ARPA programs.
3) AI integrity, backdoors, secret loyalties, AI security, national security

Check the original publication for accuracy and context.