Hackers Outsmart AI Safety with Clever Tricks

New reports reveal how criminals are exploiting AI vulnerabilities to launch sophisticated attacks.

AI2Day Newsdesk3 min read
A digital representation of cybersecurity shields and email icons, illustrating the concept of email security
Share

Key points

  • Cisco Talos discovered criminals bypass AI filters by claiming 'authorised testing' across multiple platforms.
  • CrowdStrike found 88% of attacks on new software flaws occur within 48 hours of disclosure.
  • North Korean group Stardust Chollima poisoned the Axios npm package in March 2026, spreading malware to developers.
  • Cloud-related cybercrime rose 171% in the first half of 2026, according to CrowdStrike.
  • Hackers hide instructions in files to trick AI assistants into executing malicious acts.

Criminals are getting creative with AI, as first reported by ThreatVectr. New studies from Cisco Talos and CrowdStrike show hackers are using AI tools to generate malicious software and speed up attacks. Cisco Talos, which is part of Cisco, found hackers trick AI by claiming their requests are for 'authorised testing' or legal hacking contests. This tactic works across various platforms, including Claude Code and Gemini. When a model refuses, hackers switch to versions without safety measures.

How fast is AI changing the hacking game?

AI is shrinking the time between software flaws becoming public and criminals exploiting them. CrowdStrike's research shows 88% of attacks on newly disclosed software gaps happen within 48 hours. AI is helping hackers close this window faster, making it harder for defenders to react in time. In March 2026, the hacking group Stardust Chollima, linked to North Korea, tainted a software package called Axios. This led to malware spreading to any developer using it. Such tactics put everyday apps and services at risk, potentially making them carriers of harmful code.

What new tricks are hackers using?

Hackers are embedding secret instructions in ordinary files to control AI assistants. When AI processes these files, it follows the hidden directives instead of user commands. Imagine a hidden note in a letter that misdirects a postal worker, and they follow it without question. This method bypasses traditional defenses since it requires no software installation. It’s a growing concern among security experts, as it can easily spread through widely-used platforms like Microsoft Word.

Event Date Impact
Axios npm package poisoned March 2026 Widespread developer exposure
Mastra AI framework affected June 2026 At least 131 packages compromised
Cloud crime increase 1H 2026 Up 171% year on year

For the average person, this means apps they use daily could unknowingly contain harmful code. If a service reports a data breach, check if it involves supply chain attacks or compromised code. Change passwords immediately if it does. Using a password manager can greatly reduce the risk by ensuring each site has a unique password.

Common questions

How do hackers trick AI into helping them?

Hackers tell AI models they are performing 'authorised testing' or are in hacking competitions. This often bypasses safety filters and allows them to generate harmful software.

Why is this a problem for regular users?

Everyday applications could be carriers of hidden malware due to these exploits. Users should be vigilant about data breach notifications and change passwords if supply chain attacks are mentioned.

How can I protect myself from these risks?

Use a password manager to create unique passwords for each service, reducing the risk of widespread harm if one account is compromised.

© 2026 AI2Day