SardineCon SF/2026

Learn More
The Saturday Fraud Strategist

The Rise of Agentic Fraud Ops, Pt. 2: 5 Steps for Adopting AI Agents in Fraud Ops

12 min

Most fraud teams that have started adopting AI agents in fraud operations started in the right place. They piloted inside investigation. They enriched alerts, structured cases, and recommended resolutions for investigators to validate. And for the most part, they are seeing real efficiency gains.

The problem is that almost nobody goes further.

In part one of this series, I made the argument that the master KPI in fraud is not precision or accuracy. It is the reaction cycle. The time it takes your system to detect a gap, whether that is a new fraud attack or a misbehaving control, and ship a fix for it. Automating your investigation process is the first step in that journey. It is not the journey.

Even if you automate investigations completely, the rest of your links in the chain are still running at human speed. The rules you deploy to flag events are still degrading. The labels feeding your models are still arriving weeks late. You are running faster investigations inside a broken loop.

This episode is about closing that loop. All five steps of it.

Before we get into the notes, if you landed here first, I'd recommend going back and listening to part one. We covered quite a lot that will make this one easier to follow. Link is below.

What you’ll hear in this episode:

  • Why automation of fraud investigation is only the first step in adopting AI agents in fraud, not the destination
  • How the fraud reaction cycle breaks down into five distinct steps, each producing the input the next one needs
  • Why fraud alert clustering is the step that turns a pile of unrelated alerts into a curated set of assembled ring-level cases
  • Why automated fraud labeling is the single most important bottleneck in the entire reaction cycle and how to close it
  • How continuous fraud labeling at scale changes what your models and rules can do
  • Why fraud risk segmentation is the most underestimated layer in fraud strategy and why it has to come before detection automation
  • How the fraud rule recommendation engine in step five only works if the four steps before it are already in place
  • Why fraud ops AI transformation is not a one-quarter project and what teams further along actually look like
  • The organizational and governance capabilities your team needs to build at each stage before the next stage makes sense
  • Why the goal is not just lower cost but a fundamentally different and better fraud organization

You should listen to this episode if you:

  • Are working through the question of where to start when adopting AI agents in fraud and want a concrete, sequenced answer
  • Have already deployed agents inside investigation and are trying to figure out what comes next
  • Manage fraud analytics, rule writing, or model governance and want to understand where agentic AI fits into your work specifically
  • Are responsible for fraud ops reaction time and want to understand how to measure and improve it at scale
  • Have felt the pain of delayed chargebacks slowing down model retraining and want to understand how automated fraud labeling solves it
  • Are building the business case for agentic AI in fraud and need a framework that goes beyond efficiency gains in investigations
  • Lead a fraud team that is feeling pressure to adopt AI quickly and want clarity on how to do it without creating the governance problems that cause these projects to fail
Episode notes & key takeaways

Why investigation automation is the right starting point but not the full picture

Most teams start with the automation of fraud investigation and there is a good reason for that. It is where agentic AI is most mature. It is where expertise is already concentrated. And it is where the savings show up first. The three areas to cover in step one are data enrichment, which moves from investigator to agent before the case is even opened, case output structure, which standardizes documentation and makes it machine readable for downstream use, and reasoning explanation, which shifts the investigator's role from making decisions to validating them.

That last point matters more than it sounds. Validating an AI recommendation is a different skill than making a judgment call from scratch. Step one is the lowest-stakes place to start building it.

What step one produces:

  • Individual investigations that are faster and better documented
  • Case output that is standardized, machine readable, and ready for downstream agentic flows
  • A team that is beginning to develop the skill of governing AI recommendations rather than replacing them

Why fraud alert clustering is where real efficiency begins

Once investigations are structured, step two is having AI agents cluster them. New alerts get matched against known cases or fraud rings by comparing device fingerprints, funding patterns, and behavioral signatures. If a match is found, the new event gets attributed to an existing ring without any human time spent on it. If no match is found, it becomes a candidate for a new case by comparing it against other alerts in the queue.

The investigator's queue stops being a pile of unrelated alerts and becomes a curated set of assembled cases. Ring expansion happens automatically, and one ruling now labels thousands of events at once. This is also where the structural work from step one pays off. Teams that skipped it cannot build clustering at all.

What step two produces:

  • A queue of assembled ring-level cases instead of individual unrelated alerts
  • Automated ring attribution that generates labels at scale
  • The dependency that makes step three possible

Why automated fraud labeling is the most critical bottleneck in the reaction cycle

Step three is where the reaction cycle actually changes speed. Chargebacks arrive weeks after the transaction, sometimes months. By the time they do, the attack pattern has either changed or already generated significant losses. Confirmed fraud cases take time to appear because a human is the one who usually triggers them. Whether it is a customer filing a dispute or an investigator ruling on a case, humans are the bottleneck. Agentic AI removes them from that loop.

By step three, you already have an AI agent proposing investigation outcomes from step one and clustering generating ring-level scale from step two. Now you start automating the labeling. The important thing to understand is that labels for training do not have to be perfect. They feed model retrains and rule backtests, not customer-facing decisions. A modest error rate from an agent is no worse than the noise already present in chargebacks and human investigation rulings. That tolerance for imperfection is what makes continuous fraud labeling at scale possible.

What step three produces:

  • The slowest part of the reaction cycle goes from weeks to minutes
  • Fresh labels that immediately feed model retraining and rule refresh
  • The fuel that makes steps four and five useful rather than broken

Why fraud risk segmentation is the most overlooked layer in fraud strategy

Step four is where most fraud teams have a blind spot. Segmentation is the layer that decides how each event gets processed and assessed. Which rules fire. Which vendors get called. Whether step-up authentication is triggered. Most teams either treat it as an afterthought or manage it as a set-and-forget configuration. But this is the layer where fraud strategy actually gets implemented. Ignoring it slows the ability to react to emerging threats.

There is also a governance problem worth naming. Segmentation likely does not have a defined owner, clear goals, or a process for moving populations between risk segments in most organizations. You cannot automate something nobody is managing manually first. The team has to know what good looks like before it can govern an agent that proposes changes.

Step four also introduces a new motion. The earlier steps had agents observing the live data pipeline and processing what flows through it. Recommending changes to segmentation is different. The agent analyzes historical performance across populations and proposes policy changes rather than case rulings. That changes how you test, validate, and observe the decisions. It is more data driven and less expert driven. And it is managed by fraud analytics, not investigations. This is where agentic AI starts expanding to more teams.

What step four produces:

  • A team that can govern data-driven AI recommendations, not just expert-driven ones
  • A process for continuously optimizing risk populations rather than setting and forgetting them
  • The governance muscle that step five requires

Why the fraud rule recommendation engine comes last

Step five is detection. Rules and models. The reason it comes last is not because it is less important. It is because it has the strictest dependencies. Detection automation needs fresh, trustworthy labels to be useful. Without automated labeling already in place from step three, an agent proposing new rules is just proposing wrong rules faster. And governing agents that create and modify rules is complex, especially if the team has not yet built the experience of managing data-driven agents from step four.

Unlike segmentation where the agent chooses between options that already exist, detection is generative. It proposes things that did not exist before. New rule patterns. New features. Possibly new machine learning models. The governance experience from step four is what makes that manageable rather than chaotic.

When step five is in place, the loop closes. Alerts come in, get clustered into ring-level cases, get labeled automatically, and fuel agents running continuous analysis for segment changes and new rule recommendations. The fraud reaction cycle runs at machine speed.

What step five produces:

  • Rules and models that refresh continuously rather than degrading between manual update cycles
  • A detection layer that responds to emerging threats rather than reacting to them after the fact
  • A fraud ops organization that operates differently in kind, not just faster at the same things

What fraud ops AI transformation actually looks like

The teams furthest along are six to twelve months in and they would tell you they have meaningful work still ahead. Each step requires tooling, organizational change, and time to absorb a new way of working before the next step makes sense.

The sequence is not only about where teams can reduce costs. It is about clearing dependencies and building new capabilities at each stage. Skip steps and you end up with rule automation running on stale labels or policy automation that nobody knows how to govern. That is exactly how AI adoption projects fail and exactly why teams end up rehiring the people they thought they no longer needed.

The goal is not a leaner version of the same organization. It is a fundamentally better one.

  • Adopting AI agents in fraud operations is a five-step sequence where each step produces the input or capability the next one requires
  • Automation of fraud investigation is the right starting point but only addresses one link in the reaction cycle
  • Fraud alert clustering transforms the investigator queue from unrelated alerts into curated ring-level cases and generates labels at scale
  • Automated fraud labeling is the most critical bottleneck in the reaction cycle, reducing the delay from weeks to minutes
  • Continuous fraud labeling at scale is what makes model retraining and rule refresh genuinely responsive to emerging threats
  • Fraud risk segmentation is the most overlooked layer in fraud strategy and must have a defined owner and process before it can be automated
  • The fraud rule recommendation engine in step five is generative, not just selective, which requires the governance experience built in step four
  • Fraud ops AI transformation is not a one-quarter project, teams six to twelve months in report meaningful work still ahead
  • Human in the loop fraud controls do not disappear in this model, they shift from making decisions to validating AI recommendations
  • The goal is not a cheaper version of the current organization but a fundamentally different one that runs the reaction cycle at machine speed
Final takeaway

The reaction cycle is the number that matters. And right now, most fraud teams are running it at human speed while the threats they are defending against adapt at machine speed.

The five steps I laid out in this episode are sequential for a reason. Each one clears a dependency the next step cannot work around. Start where most teams start, in investigation, and follow the sequence. Do not skip ahead.

The teams that try to jump to rule automation before they have labels, or to labels before they have structure, are the ones that end up reversing course. Go in order, build the capability at each stage, and the end state is an organization that does not just run faster. It operates differently.

This is part of a series. If you landed here first, you may want to go back and listen to the previous episode. We’ve already covered quite a lot that will make this one much easier to follow.

Catch up on part 1

Not ready to stop the conversation about my, and hopefully your, favorite subject? Subscribe to The Saturday Fraud Strategist newsletter.

Connect with Chen Zamir | LinkedIn
Host of The Saturday Fraud Strategist
Helping fintechs build smarter fraud defenses
Co-author of “The Fraud Fighter’s AI Playbook

Episode transcript
Chen Zamir
Chen Zamir
00:06
When I speak to fraud teams that have started using Agentic AI, I find that most started in the right place. They're piloting agents inside investigation, enriching alerts, structuring cases, and recommending resolutions for investigators to validate. And for the most part, these teams are seeing clear efficiency gains from using agents. The problem I see is that almost nobody's going further. In a previous video, link is down below, I argued that the master KPI in fraud isn't precision or accuracy. It's the reaction cycle. The time it takes your system to detect a gap, be a new fraud attack or misbehaving in fraud controls, and ship a fix for it. And while fraudsters are deploying AI to adapt faster to new defenses, fraud teams should also use agentic AI to match that speed. Now, automating parts of your investigation process is just the first step in that journey. It is not the journey. Why? Because investigations are just part of the entire process that makes your reaction cycle complete. Even if you automate that one step completely. Which, admittedly is easier said than done, the rest of your links in the chain are still run at human speed. So yeah, your investigations are going to run faster, but you're still running a broken loop. The rules you deploy to flag events in the first place are still degrading. The labels you're feeding your models are still arriving weeks late. Now imagine for a second what it would mean if the entire cycle ran at machine speed. Your system then becomes something entirely different. You make decisions on entire rings instead of single individual events. You're not running faster investigations anymore. You're making decisions at scale. And each of these decisions not only reduce your workload, but they also generate hundreds, if not sometimes thousands of new labels you can use immediately. And if you can generate labels fast and at scale, you can then use those to refresh your rules, your policies, and your models.
Chen Zamir
Chen Zamir
02:08
And if those are refreshed more frequently, then you can mitigate how fast they degrade. Then in essence, you've built a reaction cycle that runs at machine speed. Sounds great, right? But also pretty complex. And one of the questions I hear most from teams who are making their first steps with AI is where do we start? In this video, I'd like to share the five steps it takes to transform your reaction cycle from running at human speed to running at machine speed. Each step produces an input the next one needs. Or builds the capability the team needs to govern it. So, you want to make sure you first identify where you are in this journey and then follow through the next steps in the right order. Skip any and you'd likely get rule automation running on state labels or policy automation nobody knows how to govern. And that's exactly how AI adoption projects fail.
Chen Zamir
Chen Zamir
02:59
Step one, the first place we start at is the same one we already mentioned, fraud investigations. After all, there's a reason why most teams start there. It's the most straightforward and natural place to pilot AI agents. This is also where Agentic AI is most mature today, where your team's expertise is likely concentrated and where the savings show up first. Specifically, you want to cover three areas. The first one is handing over data enrichment from the investigator to the agent. The 15 minutes of pulling device data, identity context, account history, and external lookups now gets done in parallel and attached to the case before the investigator even opens it. The second one is to structure case output. The documentation, reporting, suggested next steps are now uniformed across all your human investigators and match perfectly your SOPs. You also get the added benefit that your cases are now logged in a machine readable form, which is what you'll need if you plan on using it further downstream for other agentic flows. And finally, the third area you want to hand over to agents is explaining their reasoning. Not only recommended decisions. A because you anyways need it for the regulator if and when they'll come knocking, but more importantly B because that's when the investigator's role shifts from making a decision to validating it. For many on your team, especially the less senior, that's a new skill and critical skill to learn. And this is the lowest stakes way to start practicing it. When you've done all of that, you should notice individual investigations become not noticeably faster, but the real gains are still invisible at this point. It's making the investigation output better structured, standardized, and machine readable.
Chen Zamir
Chen Zamir
04:49
Once investigations are structured, we can move on to step two, having AI agents cluster them. As they come in, new alerts get matched against known cases or fraud rings. We do that by matching device fingerprints, funding patterns, or behavioral signature. If we find a match, we can attribute this new event to an existing ring without the human needing to spend time on that. And if we don't find a match, the new alert becomes a candidate for a new case by comparing it to other existing alerts in our queue. In this way, the investigator's queue stops being a pile of unrelated alerts and becomes a curated set of assembled cases. Render expansion happens automatically and unlocks your team's true efficiency as one ruling now labels thousands of events at once. And this is exactly where the structural work we mentioned in the previous step pays off. Teams that skipped it can't build clustering at all.
Chen Zamir
Chen Zamir
05:48
Okay, so far you may have gained in efficiency, but your reaction cycle as a whole is not meaningfully faster. Step three is where we change that. And I can't stress how important it is. You see, the biggest bottleneck in your reaction cycle is labels and specifically getting fresh labels fast. We all know that chargebacks arrive weeks after the transaction, sometimes months. And by the time they do, the attack pattern either already changed or created big losses. Many times it's both. And that's not only relevant to chargebacks. Confirmed fraud cases take time to appear. Why? Because a human is the one who usually triggers them. Whether this is a customer who files any kind of dispute or an investigator that looked at it and ruled it as fraud. And that's exactly the bottleneck we want to remove from our reaction cycle. We don't want humans to slow it down. But with agentic AI, you can close this gap with relative ease. You already have an AI agent proposing investigation outcomes and by now you should have already gained the confidence needed to start removing the humans from the loop entirely. Maybe not for all the cases, but definitely for most of them. At this point, you are taking the AI recommendations you introduced in step one, combine them with the scale you established in step two, and you start automating it. And the thing to remember is that labels for training don't have to be perfect. These labels shouldn't affect anything other than your training data set. They only fit model retrains and rule back test, not customer facing decisions. Now, can wrong labels potentially undermine your decisions? Theoretically, yes. But in practice, that isn't such a big concern. Rules and machine learning models already work with noisy data. A modest error rate from an agent is no worse than the noise you already have in chargebacks and investigation rulings. After all, humans make mistakes, too. And this tolerance for imperfection is what makes continuous labeling at scale possible. So by the end of this step, the slowest part of your reaction cycle just went from weeks to minutes. And that's when the real possibilities open up.
Chen Zamir
Chen Zamir
07:58
Okay. So where do we go from here? In step four, we're starting to introduce AI agents into your risk segmentation. What's that? Segmentation is the layer that decides how you'll process and assess the risk of each event. Which rules fire, which vendors get called, or whether step up authentication is triggered. To be frank, many fraud teams tend to treat this step as an afterthought. And even the ones that don't usually resort to set and forget policies. But this is the layer where you actually implement your fraud strategy and ignoring it slows your ability to react to emerging threats. But here also lies the problem of AI adoption. Segmentation likely doesn't have a defined owner, goals, KPIs, or even a process for moving populations between different risk segments. And the thing is that you simply can't automate something nobody's managing manually first. Your team has to learn what good looks like before it can govern an agent that proposes changes. There's also another difference worth flagging. This is a new motion for your team. The earlier steps all follow the same pattern where the agent observes the live data pipeline and processes what goes through it. But recommending a change to your risk segmentation works differently. The agent analyzes historical performance across populations and proposes policy changes, not case rulings. This isn't just semantics. It changes how you test, validate, and observe agent decisions. It's much more data driven and much less expert-driven. This means that the team which will manage it would likely not be the team that runs investigations. But the team that manages rules, models, and other analytical components of your system. Meaning fraud analytics. This is where we're starting to expand the use of AI agents to more teams and processes across your organization. Having said all of that, I also want to stress that this isn't as intimidating as it sounds. Let me give you an example. An agent recommendation can be to take a segment of accounts aged 25 to 30 day old and move them out of the high-risk bucket and into the medium risk because their behavior looks like established users. So, you see, this isn't about finding new fraud patterns or emerging threats. It's more about optimizing existing segments and processes. That makes the problem much more bounded than fraud rewriting for example which makes it exactly what you want to do to build that muscle.
Chen Zamir
Chen Zamir
10:33
And then lastly, step five touches exactly that. Handing over your detection layer to AI agents. What is the detection layer? It's where you run your signals and the algorithms that use them to detect fraud. Or in simple words your rules and your models. When we are able to refresh these in an automated way, we expedite the last phase in the reaction cycle. Which is actually reacting to emerging threats. Not only detecting them. So why is it the last step if it's so important? Well, there are two reasons for that. The first reason we've already covered. Fuel detection automation needs fresh trustworthy labels to be useful. And without automated early labeling already in place, an agent proposing new rules is just proposing wrong rules faster. We don't want that. The second reason is that governing agents that create and tweak rules is complex. And that is especially true if your team hasn't built the muscle for writing rules themselves or managing agents that run deep data analysis like we've done in the previous step. Because unlike in segmentation where the agent pick between options that already existed, detection is generative. It proposes things that didn't exist before, new rule patterns, new features, and possibly even new machine learning models. Mastering agentic segmentation gives your team the governance experience agentic detection requires. And once you've mastered this step, you're running your reaction cycle at machine speed. Alerts come in, get clustered into ring level cases, get label automatically, and fuel agents running continuous analysis for segment changes and new rule recommendations. And that, ladies and gentlemen, is the future of fraud ops.
Chen Zamir
Chen Zamir
12:19
Now, let's be honest. This isn't a one quarter project. The teams further along are 6 to 12 months in, and they'd say they have meaningful work still ahead. Each step needs tooling, organizational change, and time to absorb a new way of working before the next step makes sense. And fraud teams are facing a real risk here. On one hand, the pressure to adopt AI and scale back headcount is growing by the day. But on the other hand, there is very little clarity on how to do it safely. In such an uncertain environment, it's crucial to understand your end goal. Not just lower cost, but a better organization. The sequence I just shared isn't only about where you can cut more costs. It's also about clearing dependencies and growing new team capabilities and skills. Ignore these and you might need to hastily rehire the team you thought was redundant. And we all know it's not that simple.