Sardine named a Leader in The Forrester Wave™: Financial Crime Management Solutions, Q3 2026

Learn More

Introducing Sardine AI Labs

Soups Ranjan
Soups Ranjan
bg-image
bg-image
Isometric illustration of an AI lab with text introducing Sardine AI Labs, a new research group for financial crime detection, highlighting improved fraud detection and research grants.
Subscribe to newsletter
Share

Can AI learn the structure of financial behavior well enough to catch attacks it was never explicitly trained to find?

That's the question behind Sardine AI Labs, our applied research group studying how AI can learn the language of financial behavior. We're exploring how models can learn from sequences of financial activity, transfer what they learn across institutions, anticipate attacks they weren't explicitly trained to find, and help risk teams make faster, better decisions.

But better model performance is only part of the problem. We test each advance against the demands of production risk systems, including detection performance, cross-institution generalization, and low-latency inference. Model accuracy is measured alongside latency, governance, explainability, and the requirements for human oversight.

We’re launching Sardine AI Labs with findings from our first foundation model for issuing risk, the research areas we’re pursuing next, and up to $375,000 in grants for independent researchers working on the same problems.

Solving the hardest problems in financial crime

The next wave of foundation model progress is likely to come from domain-specific AI neolabs with access to proprietary, real-world data. Fraud and financial crime prevention is a good example: even strong academic researchers often lack the large-scale datasets needed to push the field forward.

Fincrime prevention is also expensive, contributing to higher payment-processing expenses across card, ACH, and wire transactions. Some of the hardest problems here are still unsolved. A consumer whose name closely resembles someone on a sanctions list can sit in an onboarding queue for weeks, because matching two names to the same person is still genuinely hard. Spotting money laundering, terrorist financing, or drug trafficking inside billions of transactions is just as hard: existing systems can generate false positives that delay legitimate transactions while still missing illicit activity, exposing institutions to significant regulatory and financial consequences.

Learning the language of financial behavior

Much like language, payment activity and access behavior contain distinct patterns. Transaction histories, devices, and IP addresses form behavioral profiles that foundation models can analyze to identify anomalies and detect potential fraud.

Most risk models start with an individual transaction and ask whether it looks suspicious. We’re also interested in what happens when the model learns from the history around that transaction too. A payment can look ordinary in isolation. Its meaning changes when you know what happened before it, which device initiated it, how the customer usually behaves, and where that payment sits in a longer sequence of activity.

That’s the premise behind much of our work in Sardine AI Labs: treat financial activity as a sequence, not a collection of isolated events.

Our research draws on Sardine’s network of:

  • 6.5B+ devices profiled
  • 441M consumer identities
  • 3.4M businesses screened
  • 6.6B transactions analyzed
  • $1.8T in transactions protected

That network gives us a large base for studying how financial behavior evolves over time across sessions, transactions, and customer profiles. It's also why we're positioned to build a model purpose-built for risk, rather than one built for something else and retrofitted to it.

What we learned from our first foundation model for issuing risk

Our first payments foundation model tested a simple idea:

Can a model learn useful representations of cardholder behavior before it ever sees a fraud label?

Instead of training only on labeled fraud outcomes, the model first learns from unlabeled transaction histories. Those learned representations are then supplied as features to the issuing fraud model already in production.

The approach is learn first, classify second:

  1. Tokenize merchant, amount, time, geography, channel, and terminal information.
  2. Pretrain on more than 10 months of unlabeled transaction histories.
  3. Represent each cardholder’s history with an eight-layer transformer.
  4. Combine those learned features with the existing fraud model.

The model was trained on about one billion transactions over the last two years from 13 card issuers, learning from complete cardholder transaction histories rather than reducing that behavior to a static set of features.

The initial results:

Held-out issuer

Improvement in fraud detection accuracy (AUC-PR)

Consumer card issuer

68%

Business card issuer

41%

The held-out issuer results matter because they test one of the hardest problems in fraud detection: can a model transfer what it learned elsewhere to an institution it has never seen before?

New card issuers pose a particular cold-start challenge. At launch, they lack the training data needed to detect fraud reliably, even as fraud rings seek to exploit newly released financial products. In testing with issuers entirely excluded from training, the model showed gains across consumer, business, and global cross-border card issuers.

One of the more important findings was what the foundation model didn’t do: it didn’t replace the existing fraud model. Our Head of Data Science, Niranjan Shetty, said: "The most important finding is that foundation models can learn directly from a user’s transaction behavior, allowing us to identify fraud more accurately than approaches that reduce that behavior to a set of tabular features. Because these patterns transfer across financial institutions, the model is not limited to a single card program. The next step is to make this intelligence fast, explainable, and reliable enough to support real-world risk decisions.”

The learned embeddings were weaker on their own. The strongest performance came from combining them with the existing model.

That points to a more useful role for foundation models in fraud: give existing risk models a richer representation of customer behavior rather than replace the entire risk stack.

The data behind the model

The initial research uses card-issuing histories from 13 issuers across consumer, SMB, crypto-card, and neobank programs.

Coverage runs from January 2024 through July 2026, with up to 31 months per issuer.

Isometric illustration of a Sardine AI Labs research machine in coral and teal, with a spiral column inside a glass cylinder and data cubes moving along a track, representing foundation models learning financial behavior to detect fraud.

Each cardholder’s complete history becomes a single sequence, reaching approximately 2,300 transactions. Each transaction contains 14 fixed fields spanning merchant, MCC, amount, timing, geography, entry mode, presence context, terminal information, and card balance.

The model learns a representation of each cardholder's behavior: what happened, when it happened, and what tends to happen next. It doesn't generate financial data.

From research performance to production performance

A model that performs well offline isn't automatically useful in a live risk system.

Sardine AI Labs is focused on turning advances in AI into measurable protection against financial crime. For us, that means production-grade requirements from the start: real-world latency, governance, and explainability, not just accuracy on a benchmark. Our research tackles open problems in sequential modeling, transfer learning, adversarial robustness, and explainability, while testing each advance against production requirements including detection performance, cross-institution generalization, and low-latency inference.

What we’re researching next

The first model gave us evidence that sequence-based representations can improve issuing fraud detection. The next questions are about how far that approach can go.

Cross-issuer generalization

How can representations learned across participating issuers improve fraud detection for an institution the model has never seen?

Multi-stream event modeling

Financial behavior extends well beyond card transactions.

Device intelligence, behavioral biometrics, onboarding and KYC, logins, card transactions, ACH and transfers, and disputes all capture different parts of the same customer history.

We’re researching how those events can be modeled as a single customer sequence, including shared tokenization, relationships between events, and the value of sparse signals relative to dense transaction activity.

Scaling laws for financial models

How does performance change as the number of issuers, length of customer histories, and model size increase?

We’re researching whether consistent scaling laws exist for financial event models.

Name matching

Can learned representations improve how we determine whether two records belong to the same person?

We’re exploring whether signals such as name, address, date of birth, and past addresses can improve how two identities are matched.

Graph-based approaches

Financial activity is both sequential and relational. Payments connect people, businesses, accounts, devices, and counterparties into networks.

We’re exploring how graph transformers can produce embeddings that represent counterparty relationships in payments, including whether those embeddings can identify money-laundering typologies such as payments chained through multiple counterparties to evade detection.

Inference at low latencies

A model can perform well offline and still be unusable for real-time decisioning.

We’re researching approaches for serving card-fraud models with more than 100 million parameters at very low latency while using each customer’s full transaction history. This may require supporting context windows of more than one million tokens.

The key question is: how much context can you use without slowing down the decision?

Cold-start learning

How much history is required before learned embeddings outperform traditional aggregation-based features?

Applications beyond fraud

Can representations learned from financial activity transfer to other predictive tasks?

Potential applications include:

  • Credit risk
  • Churn prediction
  • Spend forecasting
  • Merchant categorization

These are research areas, not product claims.

$375,000 for independent research

We also want researchers outside Sardine working on these questions.

Through the new Sardine AI Fellowship Program, Sardine AI Labs is offering up to five research fellowships, totaling $375,000, to independent researchers currently enrolled in universities in the US or Canada who are working on problems relevant to financial crime detection.

Research areas include:

  • Cross-issuer generalization
  • Multi-stream event modeling
  • Scaling laws
  • Identity matching
  • Detecting money-laundering patterns
  • Real-time inference in the low hundreds of milliseconds
  • Cold-start learning
  • Other applications beyond fraud

Applications are open through December 13, 2026. Shortlisted applicants will be notified in December, and the inaugural cohort of fellows will be announced in mid-January 2027.

Who can apply

Independent researchers currently enrolled in universities in the US or Canada

Awards

Up to five research fellowships, totaling $375,000

Applications close

December 13, 2026

Shortlisted applicants notified

December 2026

Fellows announced

Mid-January 2027

How to apply

sardine.ai/ai-labs

Learn more about eligibility requirements here, and apply for a research grant here.

Follow the research

We’re publishing the methods, results, and limitations behind this work, not just the headline numbers. Our first technical whitepaper goes deeper into the issuing foundation model, including the data preparation, tokenization, architecture, experiments, results, and open questions.

We’ll continue publishing as we test these approaches across more institutions, signals, and financial crime problems.

A fraud score only tells you what a model decided. We're working on models that understand the behavior behind it.

FAQs

What is Sardine AI Lab?

Sardine AI Labs is Sardine’s applied research group focused on advancing AI for financial crime detection. We research how models can learn from financial behavior, generalize across institutions, and combine transaction, device, identity, and behavioral signals more effectively.

What data does the foundation model research use?

The initial research uses card-issuing histories from 13 issuers across consumer, SMB, crypto-card, and neobank programs, covering January 2024 through July 2026.

Each cardholder’s history is represented as a sequence of roughly 2,300 transactions, with 14 fields per transaction spanning merchant, MCC, amount, timing, geography, entry mode, presence context, terminal information, and card balance.

How is Sardine AI Labs different from Sardine’s existing foundation model work?

The foundation model work is the research itself.

Sardine AI Labs is where we publish that research, make the research agenda visible, and open parts of the work to external researchers through our grant program.


Who can apply for a research grant?

Independent researchers and academic teams working on problems relevant to financial crime detection can apply.

Applications are reviewed based on the proposed research approach and the researcher’s track record.

Where can I read the research?

Sardine AI Labs will publish technical reports and shorter research explainers covering the methods, results, and limitations behind the work.

The first technical report goes deeper into the issuing foundation model, including data preparation, tokenization, architecture, results, and open questions.