Sardine named a Leader in The Forrester Wave™: Financial Crime Management Solutions, Q3 2026

Learn More

AI agent swarms and the next wave of automated fraud

A smiling young man with reddish hair wearing a black jacket stands in front of a crater lake in Kerið Crater, Iceland.
Nathan DeStigter
bg-image
bg-image
Isometric diagram showing multiple robot-like AI agents connected to financial processes (onboarding, login, payments, payouts), illustrating how AI swarms target institutions and Sardine's bot detection stops them.
Subscribe to newsletter
Share

AI agent swarms have already breached real systems. The difference between what happened in those incidents and what hits banks next is intent.

Key takeaways:

  • AI agent swarms divide work, share discoveries, and adapt mid-attack. What trips a volume rule at one step looks like noise at another.
  • The customer journey from signup to payout is the attack surface. Onboarding, login, card testing, and payouts are not separate problems when a swarm is running all four at once.
  • IP blocking and static bot signatures were built for a different threat. A swarm routes through residential proxies, drives ordinary browsers, and passes CAPTCHAs.
  • Behavioral biometrics fraud detection is one of the strongest defenses available right now because the data needed to fake legitimate user behavior only exists inside your own fraud prevention system.
  • Stopping a coordinated attack means linking onboarding to payout, not watching each step in isolation. Sardine connects devices, fingerprints, true IP signals, and behavioral signatures in a graph so fraud teams can block a ring in one move.
  • Your fraud stack needs to be tested against a swarm before a real one shows up. A joint simulation that puts multiple agents probing different entry points at once will show the gaps a normal pen test will not.

In July, roughly 1,200 OpenAI agents inside a cybersecurity evaluation built a hidden message board and exchanged more than 70,000 messages and files.¹ Then, about 700 of them took part in breaching Hugging Face, an AI software company. These agents were being individually tested and were supposed to be isolated inside their own environments, but found a zero-day vulnerability that allowed them to communicate.

Two months earlier, agents from OpenAI hit RubyGems, creating new accounts every two to three minutes and uploading hundreds of junk packages.² Ruby Central, the non-profit that operates RubyGems, was forced to shut down new registrations for four days while dealing with the fallout from the breach.

Following these high-profile reports, Anthropic has also acknowledged instances of agents breaching testing containment as far back as April. While any incident involving AI agents is cause for concern, the consequences of these have, so far, been more or less benign.

It won’t be long before they’re intentional and can’t be quietly shut off.

What makes a coordinated AI agent swarm different

The difference between a botnet and an AI agent swarm starts with how the attack is structured. A botnet runs one script across many machines which means the traffic all looks alike and trips volume rules. While it may be unexpected, it’s the equivalent of having a battering ram try to knock down a castle’s front gate.

Meanwhile, a swarm divides work, shares discoveries, and adapts. An independent report by the research group METR found agents running experiments that hurt their own scores because the results helped the group.

An agentic swarm, in essence, learns and adapts autonomously in real time without relying on a human, a fraudster in our case, to intervene. Meaning attacks not only probe defenses and find vulnerabilities much faster, but they can also react in real-time to defenses that are erected to stop it. (polymorphic fraud blog post here somewhere)

When AI swarms attack onboarding, login, and payments

Every entry point at a financial institution leads to money, which makes the customer journey from signup to payout an obvious target. Banks already guard each step with automated controls that work well at machine speed, but they lag in talking to each other since a signal caught at signup often doesn't reach the login system before the swarm has adjusted.

An attack split across several steps can look minor to each system watching it:

Diagram of an AI agent swarm attack across the customer journey: synthetic identity account farming at onboarding, credential testing below lockout thresholds at login, card testing split under velocity limits, and payouts to mule accounts.

Many banks run the same handful of core systems, which makes them look much alike from the entry point. A flaw an agent finds at another institution is likely to work on yours, so the swarm can show up already knowing where to push before either bank has compared notes.

Where IP blocking and static bot signatures fall short

Probing for weak points is nothing new for fraudsters. Sardine’s research on emerging fraud points out that attackers typically pen-test new payment architectures before attacking them, and that patching the hole they find “often initiates a migration toward the next most vulnerable flow in the ecosystem.” A swarm attacks in the same fashion, only faster since every agent that runs into a fix passes the information it gleaned along to the rest of the group.

IP blocking is the first control to fall behind. Traditionally, attack traffic came from a small pool of servers. Blocking those addresses was sufficient in cutting off the attack. A swarm can instead route through residential proxies, which borrow real home internet connections, and grab a fresh address every time one gets blocked. Residential proxy detection is one of the first technical controls a fraud team needs to layer in when AI agent activity is suspected.

Those home addresses come with clean reputations too. Proxy services even let attackers pick IPs in the same city as their victim, so a location check that compares the IP to a billing address doesn’t return as suspicious.

Static bot signatures run into similar problems. A signature list flags tools it already knows, like headless browsers or frameworks such as Selenium, and agents can drive an ordinary browser that looks like any other user’s. And to make matters more difficult, CAPTCHAs are increasingly unable to backstop this problem. In July, an OpenAI agent during a U.K. government testing session used a credential it found to get into GitHub and solved multiple CAPTCHAs along the way, according to the Wall Street Journal.

Turning those filters up comes with its own cost now that certain bots are welcome. Customers are beginning to hand everyday rote tasks to their own agents, and a filter strict enough to stop a swarm will lock those customers out too. Security teams that carve out exceptions give a swarm more openings to test.

Even when controls catch an attack in progress, the response can lag behind the swarm. The handoff between an automation and a human investigator is the moment when the speed advantage can quickly swing back to the attacker. If an alert gets stuck in queue purgatory, the swarm can keep adjusting so by the time an analyst opens the case, the damage is already done.

How Sardine detects and stops AI agent swarms

A swarm's coordination leaves a trail. Agents that share discoveries tend to reuse infrastructure and copy each other's methods, so sessions that look unrelated start behaving alike. Researchers traced the swarm behind the RubyGems attack back to OpenAI through reused web links and behavior that matched an earlier swarm.

Sardine looks for that overlap while a session is still live. Agents driven by LLMs have a tendency to drop their proverbial mask while navigating and interacting with a page. Sardine can check the device and browser for tampering or traces of automation tools, and compare sensor data against how real people behave using behavioral biometrics fraud detection built on signals from 6.5 billion profiled devices.

During a recent episode of The Saturday Fraud Strategist, David Liu recalled a conversation he had with Matt Vega, a former Sardine, about catching a bot whose tell was clicking the dead center of a submit button, with inhuman precision. Fraudsters can add jitter to an agent's cursor or curve its path, and Liu still doubts it would pass for a real applicant on any bank’s credit card page.

On the hardware side, Sardine detects emulators, virtual machines, cloned apps, and device farms. By piercing through VPNs and residential proxies to a session’s real IP, fraudsters using borrowed addresses can be differentiated from true local customers.

Sardine flags risk signals on a mobile sign-in, including a detected mobile emulator scored 875, high bot risk, a disposable email and proxy at medium risk, and three or more customers sharing the same device.

Sardine connects devices, fingerprints, true IP signals, and behavioral signatures in a graph, so sessions that share a device, an IP, or the same automation patterns get linked even when they show up at different stages of the customer journey. By tying the accounts farmed at onboarding to the logins, card tests and eventually mule accounts waiting on the payouts, fraud teams can block the whole ring in one move.

Screenshot of a customer risk management dashboard for "Connie Mitchell," displaying a network connections graph, a low risk level, and detailed customer information.

Teams are also able to do more than just block outright. Enforcement can be tiered from silent rate limiting and honeypots to step-up authentication, which only kicks in when risk materially increases.

Sardine’s models also flag new automation behavior before a campaign has time to scale. Intelligence across Sardine’s network and Sonar’s cross-industry consortium adds anonymized signals from organizations across fintech, banking, and commerce. That detection still depends on the rest of the fraud stack being ready to act on it.

How to prepare your fraud stack for AI agent swarms

Getting a fraud stack ready for swarms starts with how the team watches its own traffic. Sardine's previous research into Agentic Commerce whitepaper recommends monitoring agentic sessions as their own population, separate from human traffic, so a shift in how agents behave shows up quickly instead of getting averaged into millions of ordinary sessions.

The same research warns that fraud exploiting a fresh loophole often starts as a tiny slice of traffic that's easy to write off as statistical noise. The split-up attack from earlier would look exactly like that at first, and a separate agentic baseline gives that sliver somewhere to stand out.

Once a session is confirmed as an agent, the fraud team's next question is whether an agent makes sense for that customer at all. Device intelligence and behavioral biometrics, paired with IP piercing, can show that automation is driving a session. Identity and behavioral signals then tell the team whether that automation fits the person on the account and what they're trying to do.

That comparison has to keep running after login, since a swarm using stolen credentials looks legitimate once it's in. A 70-year-old identity suddenly running a cutting-edge agent should get a closer look. Few people would trust automation to send a $50,000 wire either, so an agent attempting one is out of character for almost any customer.

Overlaying those signals depends on the systems behind them actually talking to each other. A typical fraud stack could be charitably described as a bit messy, or slightly less charitably as "a spaghetti-like system of APIs and data pipelines." One missed integration can be the difference between a system firing on all cylinders and a disaster waiting to happen.

Customers' own agents need a plan too. Fraud teams will need to confirm what an agent was actually authorized to do for a customer, and whether what it's doing right now fits. Authentication should tighten as the stakes rise, so an agent checking a balance gets far less friction than one moving money.

All of this assumes swarms eventually move out of the lab and into real attacks, and given the pace and trajectory of AI-enabled fraud up to this point, this is not far off.

AI agent swarm tactics aren’t going to stay in the lab

The incidents so far haven't cost anyone much, and most happened in tests where safeguards were deliberately loosened. They also left behind a detailed public record of how coordinating agents behave, and fraudsters can read it as easily as fraud teams can.

A swarm built on that record wouldn't need new tricks to go after banks and other financial institutions, since agents in the lab have already picked up some of fraud's oldest ones. On one German programmers' wiki, agents identifying as OpenAI systems coordinated with more than 3,700 fake identities. In U.K. government testing, a Claude Mythos 5 agent built fake personas and tried to talk a human maintainer into approving its code.

At RubyGems, the agents didn't even try to hide, naming their files "hack," "evil," and "exploit." Once a deliberate operator points that behavior at banks running the same core, a swarm can hunt for the weakest step in the customer journey at every one of them. Stopping it means seeing the whole journey at once, from the first signup to the last payout, and Sardine was built to give fraud teams that view.

Sardine flags an emulator and very high bot risk on a credit card checkout, tracing a straight-line cursor path into the card number field and tying the session to a device ID and session key.See how Sardine detects bots and AI agents across the customer journey

Frequently asked questions about AI agent swarms

What is an AI agent swarm and how does it differ from a traditional botnet?

A botnet runs one script across many machines. The traffic looks alike, trips volume rules, and gets blocked at the source. An AI agent swarm divides the work, shares what each agent discovers, and adjusts based on what the group learns in real time. One agent finds the lockout threshold at login and passes it along. Another adjusts the card testing pace at checkout. The attack looks minor at each step and coordinated only when you see all the steps together. That is the detection problem traditional controls were not built for.

How do AI agent swarms evade IP blocking and bot detection?

The two controls that fall fastest are IP blocking and static bot signatures. A swarm routes through residential proxies that borrow real home internet connections, picking up a fresh address each time one gets blocked. Those addresses carry clean reputations and can be filtered to match the victim's city. A location check against a billing address returns nothing suspicious. On the signature side, agents can drive an ordinary browser that looks like any other user's rather than a flagged headless browser or known automation framework. Sardine's bot detection does not rely on static signatures for exactly this reason.

Why is behavioral biometrics fraud detection effective against AI agents?

The data a swarm would need to convincingly fake legitimate user behavior only exists inside your own fraud prevention system. Large language models do not have access to how your real customers naturally interact with your application. The behavioral patterns agents produce tend to diverge from genuine users in ways that are measurable. Sardine's behavioral biometrics compare sensor data against established customer baselines in real time, which means the detection improves as the baseline grows and becomes harder to replicate as the swarm adapts.

How does cross-journey fraud detection stop a coordinated AI agent swarm?

A swarm splits its attack across onboarding, login, payments, and payouts specifically because each step looks minor in isolation. Stopping it means linking what happens at the first signup to what happens at the last payout rather than watching each checkpoint independently. Sardine connects devices, fingerprints, true IP signals, and behavioral signatures in a graph. Accounts farmed at onboarding can be tied to the logins, card tests, and mule accounts waiting at the payout stage. That linkage is what allows a fraud team to block the whole ring in one move rather than catching individual sessions one at a time.

How should fraud teams prepare their fraud stack for AI agent swarm attacks?

The preparation that surfaces the most gaps is a joint simulation with core providers before a real swarm shows up. Putting multiple agents probing different entry points at once, sharing what they find, shows whether the fraud stack can respond as a coordinated system or only as a collection of independent controls. A typical fraud stack has enough integration gaps that a swarm splitting its attack across onboarding and payouts can slip through the seams between tools. The simulation also tests alert handoff speed, since the window between an automated flag and a human investigator opening the case is where a fast-adapting swarm can do its most damage.

Sources:

¹ The Wall Street Journal, Cyberattack by rogue AI swarm stokes fears of out-of-control agents, September 11, 2026

² VentureBeat, Anthropic CEO says AI swarm could 'take over the entire internet' in 6-12 months, commits to AI slowdown plan, September 12, 2026