www.darkreading.com 7/28/2026, 8:28:53 PM · external

GAVEL framework improves AI safety by analysing model activations

GAVEL framework improves AI safety by analysing model activations
CyberSIXT Evidence Panel Source marked as original reporting
Primary Source arxiv.org

RESEARCHERS propose a model-agnostic approach to AI safety that emphasizes activation analysis of large language models (LLMs) to identify potential harmful actions, moving beyond traditional token analysis. This method involves categorizing cognitive elements (CEs) that can signal risky behaviors, such as phishing attempts.

Their system, termed GAVEL (Governance via Activation-based Verification and Extensible Logic), aims to improve detection accuracy of malicious activities by focusing on neural activation patterns rather than just input or output content. The researchers plan to introduce this framework at Black Hat USA 2026, encouraging the community to adopt and contribute to the development of these detection rules, enhancing overall AI security.

View full article

Article by CyberSIXT