Sardine named a Leader in The Forrester Wave™: Financial Crime Management Solutions, Q3 2026

Learn More

What it actually looks like to run an AI-native fraud team

Brittany Geronimo
Brittany Geronimo
bg-image
bg-image
Diagram of an AI-native fraud team's feedback loop, featuring steps for Data, Evaluate, and Human Review.
Subscribe to newsletter
Share

Key takeaways:

  • The fraud analyst role changed before the org chart did. At Imprint, tedious monitoring and governance work is now automated, freeing analysts to spend more time finding new fraud patterns and feeding them back into models.
  • AI fraud agent calibration is an ongoing requirement, not a one-time setup. Agents left unchecked flag everything as suspicious and false positives eat the gains.
  • The real product of Intuit's invoice agent turned out to be the fraud agent feedback loop, not the model itself. Fewer, more decisive signals outperformed volume.
  • Fraud label quality defines your ceiling. If ground truth comes from less rigorous reviews, the agent can never surpass them. Imprint measures agent performance against trained fraud specialists, not average reviewers.
  • Nobody running an AI native fraud team is betting on one large general-purpose model. All three companies converged on specialized agents built on top of a shared harness, with domain teams owning the agents relevant to their work.

At SardineCon, we put three risk leaders on stage who are living this every day: Karthik Thillai, who leads payments risk at Intuit; Lawrence Lin Murata, CEO at Slope; and Lalitha Rao, CRO at Imprint. Between payments risk, SMB credit underwriting, and consumer credit and fraud, the panel covered a lot of ground. But the same themes kept surfacing: workflows built for humans don't work for agents, feedback loops matter more than model architecture, and nobody up there is trying to build one model to rule everything.

Here's what stood out.

The fraud analyst role changed before the org chart did

Every panelist described the same shift from less time on manual monitoring and SLA chasing to more time on the work that actually requires judgment.

At Imprint, Lalitha said the day in the life of a risk analyst is "entirely different" from four months ago, because a lot of the tedious monitoring and governance work is now automated. That freed the team to spend more time on strategic and growth-focused work. At Intuit, Karthik described something similar at the organizational level with the old lines between strategy, forecasting, and analytics becoming blurred, and analysts now own outcomes end to end instead of handing off between functions.

Four panelists sit on a stage, two holding microphones, with a screen displaying names and titles behind them.

None of the panelists framed this as headcount replacement. Lalitha stated directly that the goal is to let the team scale without scaling headcount linearly. Ideally, analysts can spend their time finding new fraud patterns this way and feed them back into the models instead of doing repetitive review work. Karthik made a related point about Intuit's AI agents handling account review around the clock, without the ramp-up time or skill variance you get with a growing human team.

Slope's SMB fraud underwriting stack: Transactions as tokens

Lawrence gave the clearest technical walkthrough of the day. Slope's underwriting used to rely on SMB owners submitting financials for manual review, a slow and tedious process. Slope Transformer, their first model, changes that by ingesting raw open banking transaction data directly and converting it into financial metrics. Because it's pulled from actual bank transactions rather than self-reported documents, it's much harder to falsify.

The more interesting layer sits on top: a foundation model that treats each bank transaction like a token and predicts what comes next, similar in spirit to how language models predict the next word. That lets Slope forecast whether a given revenue or expense stream is likely to recur, and when, and at what size, which turns out to outperform the older rules-based approach Slope used to run in shadow mode before it graduated to production.

Lawrence was careful to separate the modeling work from the compliance work. The probability-of-default model that actually makes lending decisions has to stay fully explainable, and Slope ensures adverse action notices and regulatory disclosures comply with governing law on behalf of its partner bank, Lead Bank.

Imprint's real-time fraud agents, by the numbers

Lalitha shared some of the sharpest numbers of the panel. Imprint runs real-time agents that watch for emerging fraud patterns and propose fixes, whether that's tightening a rule or engineering a new feature. Each agent run covers about six hours and does what would otherwise take an analyst nine days, for a token cost of a few hundred dollars. Lalitha said that translates into hundreds of thousands of dollars in prevented fraud losses every month.

The catch is agents left unchecked tend to flag everything as suspicious. Karthik and Lalitha both raised this as a pattern across their teams, agents need to be actively calibrated back toward what normal transaction behavior actually looks like, or false positives eat the gains.

The AI invoice agent nobody expected

Karthik shared how Intuit's analysts built an agent that runs hourly against every invoice moving through QuickBooks' invoicing platform, evaluating whether an invoice looks like an advance payment, a completed service, an inflated amount, or something that looks fabricated. It doesn't make a fraud call outright. It assigns a risk rating that feeds into the broader fraud and credit decisioning agents downstream, as one signal among several.

Lalitha Rao, Chief Risk Officer, and Karthik Thillai, Head of Payments Risk, speak during a panel discussion.

The lesson Karthik came back to is that more data isn't automatically better. Intuit initially fed the model everything available and got a signal-to-noise problem in return. Fewer, more decisive signals outperformed volume, and the real product turned out to be the feedback loop, not the model itself.

Two kinds of labels, and a higher bar than the "average analyst"

Intuit's feedback approach runs on two label types. Eventual labels come from real-world outcomes, chargebacks, account closures, the kind of ground truth that takes time to materialize. Synthetic labels come from experienced risk professionals manually annotating sample transactions in the meantime. Karthik said combining both gives the model faster signal while still producing an audit trail that governance and compliance partners can review.

Imprint follows a similar philosophy, and sets the bar intentionally high by measuring agent performance against trained fraud specialists. Lalitha's insight is that label quality defines your ceiling so if your ground truth comes from less rigorous reviews, the agent can never surpass them.

Nobody's betting on one model to do everything

Karthik summed up where the group has landed on architecture. We are in a moment where the industry's early enthusiasm for one large, general-purpose model handling every use case has cooled. Feed a model too much context and you get noise, not precision. The pattern all three companies converged on is a shared agent harness that works with any frontier model, with domain teams, credit, fraud, ops, building their own specialized agents on top using their own expertise.

That preference for specialization showed up again in the proprietary data conversation. All three panelists said they'd rather build in-house than hand data to a frontier lab. Lawrence's reasoning is that a smaller model trained on Slope's own data can match frontier-model accuracy while beating it on latency and cost. Lalitha pointed to Imprint's position between merchant and customer, which gives them behavioral data that's genuinely hard to fabricate, unlike a stolen identity, a fabricated relationship history with a specific partner doesn't hold up. Karthik's team buys external signals but keeps decisioning in-house, partly because past vendor relationships came with a one-way data flow that never actually improved Intuit's own models.

Where it went wrong, and how they caught it

Asked directly about failure points, the panel didn't dodge it. Lalitha described early results that were interesting but not always accurate, and said "Claude got it wrong" stopped being an acceptable answer internally. Every agent now has a named person accountable for what it does.

An audience views a panel discussion on a stage with a large screen displaying names and titles of panelists including CEO and Chief Risk Officer.

Karthik's team tracks real-time agreement rates between AI and human agents, and a significant deviation triggers ramping the agent down or suspending it outright. Lawrence's approach leans on visibility with Slope's internal agent platform running in public Slack channels. Doing so allows engineers to catch a wrong answer in the open and allows the team to build shared judgment about what's actually correct, together.

The takeaway

Across payments risk, SMB underwriting, and consumer credit and fraud, the same operating principles kept showing up. Workflows should automate the tedious parts so people can do the judgment work, build tight feedback loops before you trust a model with a real decision, hold agents to a higher bar than an average human reviewer, and resist the pull toward one model that tries to do everything. It's already live at three companies, not as a proof of concept, but in actual production. And the teams running this playbook aren't just keeping up. Their analysts are handling more cases without adding more people.

FAQs about an AI native fraud team

What does an AI native fraud team look like in production?

An AI native fraud team uses specialized agents for distinct tasks like anomaly detection, rule writing, invoice risk scoring, and case investigation, rather than one general-purpose model. The teams at Intuit, Slope, and Imprint landed on the same idea. Instead of one model trying to handle everything, teams build purpose-built agents for specific jobs. Different parts of the business own the agents relevant to their work.

How do fraud agent feedback loops improve model performance over time?

The feedback loop is what makes the whole thing improve over time. The tricky part is that you often do not know right away whether a fraud decision was correct. Chargebacks and account closures can take weeks to show up. So Intuit's team does two things in parallel. They wait for those real outcomes to come in, and in the meantime they have experienced risk professionals manually review sample transactions and mark them. That way the model keeps learning without having to wait weeks every time it needs new training data. Karthik's team found that getting this process right mattered more than anything else about the model itself.

How do fraud teams keep AI fraud agents from going off the rails?

The short answer is that someone has to own each agent. At Imprint, every agent has a named person accountable for what it does. At Intuit, Karthik's team tracks real-time agreement rates between AI and human reviewers. If the gap gets too wide, the agent gets ramped down or suspended until the team figures out what went wrong. At Slope, Lawrence runs the internal agent platform through public Slack channels on purpose. Engineers can catch a wrong answer in the open and the whole team builds shared judgment about what correct actually looks like. Fraud agent oversight is not a set-it-and-forget-it problem. It is an ongoing operational responsibility.

Why are AI fraud agents prone to false positives and how do teams fix it?

Left to their own devices, agents err toward flagging everything as suspicious. Both Karthik and Lalitha raised this unprompted. The fix is ongoing calibration back toward what normal transaction behavior actually looks like for your specific business. It is not a one-time setup. And the quality of your ground truth sets a hard ceiling. If your labels come from less rigorous reviews, the agent simply cannot outperform them no matter how good the model is.

How does SMB fraud underwriting AI work differently from traditional credit underwriting?

The old way was asking business owners to submit financial documents and having someone review them manually. Slow, and easy to game. Slope's Transformer model pulls raw transaction data directly from open banking and converts it into financial metrics automatically. On top of that sits a foundation model that treats each bank transaction like a token and predicts what comes next, the same way a language model predicts the next word. Because it is based on actual transaction history rather than self-reported documents, it is significantly harder to falsify.