Sardine named a Leader in The Forrester Wave™: Financial Crime Management Solutions, Q3 2026

Learn More
The Saturday Fraud Strategist

The Rise of Agentic Ops, part 4: How to Monitor AI Agents

If one of your AI agents had been degrading for six weeks would you know? Most people tell me no, and that’s the problem I’m digging into.

Drifting agents can generate outputs that look fine on the outside while quietly getting worse underneath. Earlier in this series I talked about the reaction cycle as the master KPI of fraud effectiveness, and how agentic AI can make that cycle dramatically faster. This time I want to answer the question that matters once you’ve deployed those agents. How do you know they’re still working.

Most dashboards answer the wrong questions and only answer whether an agent is running. AI agent monitoring means tracking an agent deliberately, and I will walk you through exactly how to do it.

What you’ll hear in this episode:

  • Why a degrading agent is genuinely more dangerous than no agent at all.
  • How to measure fraud reaction cycle speed at each individual stage rather than just watching one lagging number.
  • Why AI agent performance metrics fraud teams should track don’t need to be perfectly automated to be useful.
  • The difference between human-in-the-loop agent monitoring and autonomous agent monitoring for agents making decisions at scale.
  • What a rising rejection rate actually tells you.
  • How to catch a silently failing autonomous agent before real damage compounds.
  • A practical three-layer AI agent monitoring dashboard fraud teams can build.

You should listen to this episode if you:

  • Are running any agentic fraud ops monitoring program and want a real framework for catching a degrading agent before it shows up in your losses.
  • Are responsible for AI agent governance fraud policies and need language that connects technical monitoring to leadership reporting.
  • Have deployed human-in-the-loop tools like investigation copilots or rule recommendation agents and want to know what to actually track.
  • Are running autonomous agents, like auto-labeling or alert clustering, with no human reviewing every decision, and worry about silent failure.
  • Want to build a genuine business case for AI agent ROI fraud investment using the reaction cycle instead of just automation hours saved.
Episode notes & key takeaways

Uptime is the wrong thing to monitor

Most dashboards only tell you whether an agent is live, and that's the least useful question you can ask. A drifting agent can look perfectly healthy while it quietly gets worse at its actual job. Real AI agent monitoring means tracking whether an agent is still contributing value and behaving as expected, not just whether it's still running.

Break the reaction cycle into stages you can actually measure

Detection speed, scoping speed, and fix design speed are each worth tracking on their own instead of watching one lagging average. None of this needs to be perfectly automated to be useful. A rough manual log kept once a week is enough to show whether AI agents are meaningfully shaving time off each of these stages, and that's a better use of your effort than waiting to build a perfect measurement system first.

Cycle-level metrics lag, and that's exactly the danger

Reaction cycle numbers are rolling averages, which means fraud agent degradation signals can hide inside them for weeks before anyone notices. By the time a slowdown finally shows up in your overall cycle time, the damage has usually already been done. You want signals that catch the drift earlier than that.

Human-in-the-loop agents leave a paper trail, but autonomous agents don't

For agents that a human reviews before anything gets actioned, like investigation copilots or rule recommendation tools, agreement rate and rejection rate are your best leading indicators. For autonomous agents making decisions with no review step at all, like auto-labeling or alert clustering, failure is silent by default. Cross-source agreement rate and distribution shift are the signals that let you catch that kind of drift before it compounds into a real problem.

Build your monitoring in three layers

Real time alerts should catch threshold breaches the same day they happen. A weekly review should look at trend lines and overall impact, not just whether an alert fired. And a monthly report should connect agent performance to cost and ROI for leadership. None of this requires new tooling, and the responsibility for it sits with your fraud analytics function.

Final takeaway

I opened this episode with a simple scenario. One of your agents has been degrading for six weeks, and you don't know it yet. If you're tracking agreement rate for your human-in-the-loop agents, you'd catch that in days. If you're tracking cross-source agreement and distribution shift for your autonomous agents, you'd catch it within a week. At the same time, you'd be able to show exactly how much value those same agents are creating, not just in hours saved, but in dollars saved by reacting to fraud faster. That's the real difference between running your agents and letting your agents run you.

This is part of a series. If you landed here first, you may want to go back and listen to the previous episodes. We’ve already covered quite a lot that will make this one much easier to follow.

Catch up on The Rise of Agentic Fraud Ops, part 1
Catch up on The Rise of Agentic Fraud Ops, part 2
Catch up on The Rise of Agentic Fraud Ops, part 3

Not ready to stop the conversation about my, and hopefully your, favorite subject? Subscribe to The Saturday Fraud Strategist newsletter.

Connect with Chen Zamir | LinkedIn
Host of The Saturday Fraud Strategist
Helping fintechs build smarter fraud defenses
Co-author of “The Fraud Fighter’s AI Playbook

Episode transcript
Chen Zamir
Chen Zamir
00:00
Here's a question I've been asking fraud leaders lately. If I told you one of your AI agents had been degrading for 6 weeks, would you know? Most would say no. And it's a problem because drifting agents are actually worse than no agents. They generate outputs that look fine from the outside while they degrade. And it's especially problematic when these outputs are being used as inputs for downstream processes like other agents. And so you can imagine how easy it is to encounter cascading events that very quickly get out of control. But the truth is that if you take a look at your average dashboard, it's designed to answer something else entirely than these questions. Is this agent even running? That is the least useful question to ask. I mean, of course, system health is important, but focus only on that and you'll miss the really important things to monitor. Whether your agents are improving your outcomes and not just whether they are live. So, how do you keep track of your agents? There are two things you want to watch for. First, is the agent actually contributing value? And the second, is it behaving as expected? Let's talk about it. In the first video of this series, I mentioned that in my view, the master KPI of fraud effectiveness is the reaction cycle. The time it takes your system to detect a gap and deploy a fix for it. And in the second video of the series, I outline how you can streamline your reaction cycle with Agentic AI to make it significantly faster. But how do you actually measure it? How do you connect AI agent performance metrics to how fast you're stopping fraud? Fraud AI agents can definitely help with that, but not necessarily in an even manner across the board. Here's what I would suggest to track and how. First, detecting coordinated fraud patterns faster. The first step in the fraud reaction cycle is detecting that there's a system gap. Without agents, this is slow by default. Alerts arrive and analysts eventually notices a pattern and then someone needs to pull related cases by hand to validate there's an issue. A new attack can run for days or even weeks before anyone connects a dots easily. But if you follow the steps in the second video of this series and you implemented alert clustering, a new ring can now surface in hours. Every alert triggers an agent that searches against known cases and matches events to existing rigs. So how do you track this? You track it by measuring the time from the first event of a new attack entering your system to when your team actually recognizing it. Now, I know what you're thinking. This is super tricky. So, I let you in on a little secret. Not every metric you measure needs to happen automatically. Just make sure that every week someone in your team goes through the alerts they've actually picked up and note down how long it took since the issue first started. Do it enough times and you'll be able to see a trend, especially if you start doing that before you implement agentic AI, as you should. The second thing you want to do is to scope fraud rings with AI agents. Once a pattern is flagged, the next question to ask is how big is it? How many accounts? What time period? What's the actual loss exposure? This is where Agentic AI moves from spotting a signal to mapping the full population at risk. Without AI, the scoping is manual and analysts will have to pull related cases, cross reference them with device data and build a picture account by account. But with AI agents, the same work can be done in a much shorter time window. How do you track this? Measure the time between when an issue got noticed to when you had it scoped in dollar or account exposure terms. Again, this doesn't have to be super sophisticated. The point isn't to say it takes 5 minutes and 12 seconds to scope an attack instead of 5 minutes and 36 seconds. It's being able to show it takes less than 30 minutes versus the day or two it took before. And that is pretty easy to do in the same manner I mentioned before by recording it manually once a week. By the way, I know that what I just said may sound like sacrilege to some of you. What? Do something manually when we talk about AI monitoring? How can I say that while I preach for more automation? But this is exactly where I see teams get distracted by investing their attention and tokens into productivity hacks instead of focusing on the right things. Don't get me wrong, if you can automate these metrics with or without AI, you should absolutely do so. But don't let it stop you from measuring them if you can't. Finally, the last thing you want to track is how fast you design and test fixes.
Chen Zamir
Chen Zamir
04:42
Once you understand the issue and what causes it, you need a fix. And AI agent can help with the two steps that used to consume the most time. Proposing a solution and proving it works. When fed a specific data set, agents can identify patterns, learn from labels, and propose a fix like a new rule. Then they can run the back test all before human even reviews it. So the analyst's job shifts from building the rule to stress testing and approving it. Now the reason why I would measure these two phases together is because in some cases like rule writing, it's a bit difficult to know when design ends and testing begins. And in other cases like SOP or policy changes, it might be that there would be no testing phase. So when you think about how to measure it, this one is a bit tricky. Supposedly you need to start measuring it when you have the root cause figured out, but this can prove to be quite a fuzzy definition. So again, you don't have to overengineer it and get a super accurate measure. It's enough to have a ballpark range for how much time it took since you started working on a solution and until it was approved for deployment. You want to see if when using AI agents, you can consistently shave a meaningful chunk of time. And if you don't see through rough measurements, it's likely not doing enough work anyway. Here's the problem with cycle level metrics. They lag. Meaning, you're likely going to measure some sort of a roll in average. And as we just discussed, a rough one at that. So if for example your rules agents start to degrade, it'll be quite hard to catch. Analyst would spend more time on prompting or tasks that it would oneshot before would now take several prompts to achieve. Or a labeling agent that started misclassifying 3 weeks ago will eventually cause your rules to drift and elevate fraud rates, but until you notice it and understand where it's coming from, it might be going on for weeks. And that's why you want to make sure your agents are not only live are not only noticeably making you react faster, but also that they do not degrade with time. And you want to know something is wrong before it shows up in your reaction cycle metrics as by then it's too late. But these signals look different depending on whether the agent has a human-in-the-loop or not. Human-in-the-loop agents are the first type of AI agents you'd monitor. These include investigation co-pilots or rule recommendation agents that have humans reviewing every output before anything gets actioned. And because you have a human in the loop, these agents are easier to monitor because the review itself creates a record. Did the human agree or reject the agents decision?
Chen Zamir
Chen Zamir
07:26
This creates a natural trail you can record and monitor. Agreement rate is the best leading indicator and one of the most useful agent metrics for human reviewed workflows. If investigators are ruling differently than the agent more often than they used to, then the agent is degrading. For rule recommendation agents, rejection rate is the equivalent. It's another AI agent performance metric that shows whether the agent is still producing useful outputs. It's a bit trickier because rejected rules don't reach production, so they don't necessarily impact your general KPIs. But a rise in rejection rate means the agent is generating faulty proposals that consume human time and resources without moving the needle. And by the way, there can be many reasons for agent degradation. Maybe there are underlying data issues. Maybe fraud has shifted. Maybe moving to a newer model didn't work well. Whatever the reason is, you could identify issues earlier by tracking these metrics. Meaning you could also start solving it faster. Exactly the concept behind the reaction cycle. Autonomous agents such as auto labeling and alert clustering don't have humans reviewing every decision they make. So when they fail, the failure is silent until something downstream breaks. Let me give an example. When an auto labeling agent starts misclassifying legitimate accounts could get labeled as fraud without anyone noticing. Obviously, no one wants that. This is why AI governance has to account for autonomous agents differently than human-in-the-loop flows. But the problem here is scale. The reason humans are not in the loop is because these flows simply make too many decisions or they make decisions in near real time. So not only you cannot track human agreement rates, the danger here is substantial. Any slight degradation can have a severe impact, but you can track other agreement rates coming from other sources or cross- source agreement rates. Let me give an example. Say you want to monitor your labeling agent.
Chen Zamir
Chen Zamir
09:31
It's not your only source of labels, right? You probably have at the very least the chargebacks you're getting as well. Now, if two sources for the same decision that were aligned 90% of the time drop to 70% of the time, you know something changed. You don't know which source is wrong yet, but you know where and when to look before the damage compounds. Another signal worth tracking is distribution shift. If your labeling agent starts flagging 40% more events as fraud in a segment that hasn't seen elevated fraud rates, that's worth investigating. Again, there can be many reasons for why an autonomous agent degrades. Just like with human in the flow agents, the point is that you now have the metrics and the alerts so you could catch it early. In the previous video of this series, I outlined how managing AI agents, including monitoring them, falls under the responsibility of your fraud analytics function. And from a tooling perspective, it's nothing new. What do we actually have here? Real-time alerts for threshold breaches. For example, when the agreement rate of a human-in-the-loop agent falls down below a certain value or when an autonomous agent's distribution skews more than a set rate. These alerts form the foundation layer of AI governance and should fire automatically on Slack or email or wherever your team catches alerts and get looked at the same day. Then you have weekly reviews of trend lines. Here you're not only making sure everything looks good regardless of whether an alert was fired or not. You also take a look at impact. Measure your reaction cycle speed and make sure that your agents provide actual value. And lastly, you have monthly reports on cost and ROI that you can share with your leadership team. And especially for the reaction cycle, it's pretty straightforward. If you can demonstrate an attack was blocked 4 days faster, you just save 4 days of losses. That's part of the ROI calculation, not just how many work hours you saved with automation. I started this video with a simple scenario. One of your fraud agents has been degrading for six weeks.
Chen Zamir
Chen Zamir
11:34
What do you know? If you're tracking agreement rates for your human-in-the-loop agents, you'd know in days. And if you're tracking cross agreements and distribution shifts for your autonomous agents, you'd know in a week. At the same time, you can now also track the value agents create. Not only how many hours you saved, but also how many dollars you saved by detecting and reacting to fraud quicker. That's how you connect your reaction cycle to your top KPIs. And this is the difference between you running agents and the agents running you.