Trojan Model Detection Using Activation Optimization

Hussein, Mohamed E., Janakiraman, Sudharshan Subramaniam, AbdAlmageed, Wael

Jun-7-2023–arXiv.org Artificial Intelligence

Due to data's unavailability or large size, and the high computational and human labor costs of training machine learning models, it is a common practice to rely on open source pre-trained models whenever possible. However, this practice is worry some from the security perspective. Pre-trained models can be infected with Trojan attacks, in which the attacker embeds a trigger in the model such that the model's behavior can be controlled by the attacker when the trigger is present in the input. In this paper, we present our preliminary work on a novel method for Trojan model detection. Our method creates a signature for a model based on activation optimization. A classifier is then trained to detect a Trojan model given its signature. Our method achieves state of the art performance on two public datasets.

artificial intelligence, deep learning, machine learning, (18 more...)

arXiv.org Artificial Intelligence

Jun-7-2023

arXiv.org PDF

Add feedback

Country:
- North America
  - United States > Virginia
    - Arlington County > Arlington (0.04)
  - Canada > Quebec
    - Montreal (0.04)
- Africa > Middle East
  - Egypt (0.04)

Genre:
- Research Report (0.70)

Industry:
- Information Technology > Security & Privacy (1.00)
- Government (0.69)

Technology:
- Information Technology
  - Security & Privacy (1.00)
  - Artificial Intelligence > Machine Learning
    - Neural Networks > Deep Learning (0.69)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found