SardineCon SF/2026

Learn More
The Saturday Fraud Strategist

誤検知マスタークラス 第2回:誤検知の発生源を特定する方法

9 min

私が不正対策チームでよく目にする最も一般的な間違いの一つは、誤検知に正面から立ち向かおうとすることです。

ある顧客が不満を訴えます。CEO は「このモデルはブロックしすぎだ」と言います。誰かがダッシュボードを開き、詐欺検知モデルのしきい値を調整し、ついでにいくつかの不正検知ルールを少しいじると、突然みんなが“前進している”ような気分になります。

正直なところ、あまり印象がよくありません。

誤検知を減らすことが間違った目標だからではありません。それ自体はまさに正しい目標です。問題は、ほとんどのチームがいきなり戦術レベルの対応に走ってしまうことです。そして、誤検知の減らし方について戦術的なことしかやらないのであれば、得られる成果もおそらく戦術レベルにとどまると考えるべきでしょう。

このエピソードでは、誤検知マスタークラスの第2部として、誤検知を生み出す不正検知モデルを、優先順位をつけて実際に改善できる「バケット」に分解する方法を取り上げます。誤検知がどこから生じているのか、どれが不正検知モデルや不正ルール、手動レビュー、上流パートナー、不正アナリスト、データ品質の問題、破損した不正シグナル、そして決済不正検知ワークフローによって引き起こされているのかを見ていきます。

誤検知を定量化することは有用です。

しかし、それは計画ではありません。

このエピソードでお届けする内容:

  • なぜ誤検知を減らすには、単なるモデル調整ではなく根本原因の分析が必要なのか
  • イベントを実際に拒否したのが誰なのかを特定する方法:ルール、不正検知モデルのしきい値、AIエージェント、不正アナリスト、手動審査チーム、イシュア、アクワイアラ、または不正対策ベンダーのいずれか
  • なぜ上流の決済パートナーが、自社の不正防止システムでは直接対処できない誤検知を生み出してしまうのか
  • 不正対策における意思決定が、決済不正検知、不正リスクスコアリング、オペレーションワークフローの各段階でどのように分解されるか
  • なぜ不正対策システムの最適化は、最悪の不正要因の特定から始まるのか
  • データ品質の問題や不正検知シグナルの破損、モデルドリフトが、実際には不正ではないのに不正リスクのように見える誤検知を生み出す仕組み
  • 不正対策チームが、十分な規模があり、改善可能で、優先して取り組む価値の高い課題のグループをどのように優先順位づけできるか

誰が聞くべきか:

  • 不正検知の精度向上に取り組む不正対策部門のリーダー
  • 手作業の審査キューを処理している不正検知アナリスト
  • 不正対策ルール、不正検知モデルのしきい値、および不正リスクスコアリングを管理するリスクチーム
  • 不正検知モデルおよびモデルドリフトに責任を持つデータサイエンスチーム
  • 発行者による支払い拒否や上流パートナーの判断に対応する決済不正検知チーム
  • 損失を増やさずに誤検知を減らそうとしている不正防止チーム
  • 誤検知のダッシュボードをじっと見つめて「で、次はどうすればいいんだ?」と思ったことがある人なら誰でも
エピソードノート

いら立たしいが、よくあるパターン

不正対策チームが誤検知に対応するのは、多くの場合、顧客からの苦情や経営層へのエスカレーション、あるいは「モデルが取引を止めすぎているのではないか」という突然の懸念が生じてからです。

まあ、そうだね。

でも、ここで自分に問いかけてみてください。本当に根本原因を解決しようとしているのか、それとも、いちばん目につきやすい部分だけを調整しているだけなのか?

その違いは重要です。

最新の不正防止システムは、チームに多くのツールを提供します

不正検知モデル、不正ルール、不正リスクスコアリング、手動審査、AIエージェント、決済不正検知システム、そして不正対策アナリストは、すべて不審な行為を阻止するうえで重要な役割を果たします。

あなたを守ってくれる同じシステムでも、各判断が実際にどこから来ているのかを理解していないと、優良なユーザーまでブロックしてしまう可能性があります。

最大の課題は可視性です

誤検知は、自社で設定した不正検知ルールから発生することがあります。慎重すぎる不正検知モデルのしきい値が原因となる場合もあります。人手による審査から生じることもありますし、イシュアー、アクワイアラー、プロセッサー、不正対策ベンダー、その他の上流パートナーが原因となることもあります。

誰が「ノー」と言ったのか分からなければ、何を直せばよいのか分からない。

偽陽性は単なるモデルの誤りではありません

優良な顧客がブロックされてしまいます。詐欺アナリストは防げたはずの案件の審査に時間を浪費します。手動審査のキューは増え続けます。不正検知チームは本来取り組むべきではない問題を追いかけてしまいます。不正対策の責任者は、なぜ顧客の摩擦が高まっているのか説明に苦労します。

そして、そのすべての混乱のただ中で、まったく正当なユーザーが「なぜ自分が不審だと判断されたのか」と首をかしげています。

あまり良くありません。

データ品質の問題

不正検知のロジック自体は問題ない場合もあります。問題なのは、システムが破損した不正検知シグナルに基づいて動いていることです。IPフィールドが間違っているのかもしれません。デバイスIDが欠落しているのかもしれません。決済メタデータがそもそも渡ってきていないのかもしれません。あるいは、モバイルSDKのバグによって、iOSトラフィックが異常に見えているのかもしれません。

突然、あなたの不正検知モデルがドリフトし始めます。不正対策ルールは誤作動を起こし、不正リスクスコアリングは実際よりも悪く見えてしまいます。

そして今、チームは本当の問題が壊れた入力であるにもかかわらず、「モデルの最適化」に取り組んでいるのです。

典型的ですね。

今後の道筋

より良い方法は、誤検知をアクター、パートナー、ユーザージャーニー、プロダクトフロー、プラットフォーム、支払い方法、地域、そしてデータに起因する問題ごとに分類することです。

次のような問題を優先してください:

  • 重要と言えるほど大きい
  • 実際に自分でコントロールできること
  • あなたが解決できないデータ上の問題が原因ではないこと
  • 重大な影響を及ぼす不正検知ルール、モデルのしきい値、または手動審査ポリシーに関連している

そうすることで、不正対策システムの最適化は、場当たり的な対応ではなく戦略的な取り組みになります。

主なポイント

誤検知を減らすことは、単に不正検知モデルのしきい値や不正ルール、目視による審査だけの問題ではありません。パートナー、プラットフォーム、ワークフロー、アナリスト、汚染された不正シグナル、そしてデータ品質の問題などを含む、不正判定システム全体を正しく理解することが重要なのです。

問題の原因がどこにあるか分かれば、ようやく何から優先して改善すべきか判断できます。

正直なところ、そのほうがずっと良い状況だと言えます。

私の、そしてできればあなたの一番好きなテーマについての会話を、まだ終わらせたくありませんか? ぜひ「The Saturday Fraud Strategist」ニュースレターを購読してください。

Episode transcript
Chen Zamir
Chen Zamir
00:08
One of the common mistakes I see fraud teams make is attacking false positives head on. They'll get a customer complaint, and they'll now trace back why this false positive happened and how to make sure it doesn't happen again, or the CEO complains that the model is blocking too much, and now they try to optimize it. The problem, of course, isn't the fact that they're trying to minimize their system's false positives. It's the fact that they're being tactical about it. And if you're tactical about how you do things, you can only expect tactical gains at best. In part one of the false positives masterclass, and if you missed that one, the link is down below, we accepted the reality that your system isn't built to report its own mistakes. We discussed several methods that can help you work around that and generate a usable picture of your false positives. Then comes the next problem. Quantifying false positives is merely an observation. It is not a plan. To reduce false positives in a meaningful way, you have to go one level deeper and ask a different set of questions. Where do they come from? Which are driven by my own system, and which my partners? Which are actually data quality problems masquerading as fraud risk? And most importantly, where should I start? That's what the second part of the series is about, breaking down your false positives into buckets that you can prioritize and act on. How do you do that? Here's my seven-step process. The first and most important step is to attach every false positive to the actor that made a decision. When a transaction is declined or an onboarding attempt is blocked, someone or something said no. That someone might be a rule in your system, machine learning model threshold, an AI agent, human analyst, manual reviewer, even a third-party partner, an issuer, acquirer, or fraud vendor. Without this categorization, you will default to optimizing the things you can see, usually your rules and your models. That's how teams spend months fine-tuning rules, only to later learn that most of their declines were coming from an issuer they never spoke to. So, your first task is to map every decline to the actor that made that decision. Sometimes it will be a single rule. Sometimes it will be a decision based on a certain score threshold. Sometimes it will be a human decision. Sometimes it will be a response coming back from a payment partner. The point is that you don't have to do this perfectly on day one. Even a rough breakdown makes an enormous difference, because it will help you get a sense for where the most value lies. Once you have that basic map, you can ask the following question: How much of this is even under my control? This is particularly important in card payments, where a single card transaction can pass through half a dozen hands before the issuer finally says yes or no. Between the cardholder and the issuing bank, you often have the merchant or platform itself, a payment service provider, an acquirer, an acquiring processor, a card network, an issuer processor, and finally the issuing bank. And that's even without counting the third-party fraud vendors that some of these actors plug into their own stack. Each of these actors can decline a transaction. Each has its own risk logic, and each contributes its own false positives. Now, if 60% of your false positives are driven by upstream partners, then your maximum sphere of influence is capped at 40%. That doesn't mean you should give up and walk away, but it does mean you should rethink your goals, how you set expectations internally, and where you spend your energy. And I think that this is an important point that I see a lot of fraud leaders miss. You cannot tune rules you don't own. You can, however, quantify the impact, make it visible, make sure everyone in the organization understands where the limits of your influence actually are. And trust me on that. This saves a lot of frustration later. And also, strategically, if you're unhappy with your partner's performance and believe it's substandard, you can always work to replace them. Once you have your decision sorted into high-level buckets, it's time to go one level deeper.
Chen Zamir
Chen Zamir
03:50
For example, you take a rule engine bucket that is responsible, say, for 25% of all false positives, and you break those 25% down to individual rules. And you do that as best you can for each one of your buckets. Now, when I say as best you can, what I mean is that realistically this can easily become time-consuming and low-value work. But let's try to break it down into the top five offenders in each category, and remember that even if a rule is relatively accurate, high volume usually means it is a primary source of false positives. High-volume decision is a good rule of thumb in terms of where to start. The point is that you want to identify the specific actors which are responsible for the most amount of false positives in absolute terms. These are your quick wins. At the same time, your gut might tell you that some offenders are flying below radar, legacy rules, outdated policies, or solutions that were never properly validated. Don't ignore that gut instinct, even if these actors have low volumes, or at least don't ignore it before you collect data on it. So, we now know what is blocking your users, but we still need to know where those users are coming from, even within the part of the stack you control. False positives are rarely evenly distributed. They tend to cluster around specific flows, such as mobile versus web, iOS versus Android, maybe different products or payment methods, or maybe even like new versus established users. If you only look at the actor that declined the event, you will see a rule or a model misbehaving. But if you look at which flow the user went through, you might find a deeper issue. Specific flow that produces corrupted or missing data, causing your entire fraud stack to misfire. Here's an example. Imagine you have an integration bug in your mobile SDK that sends incorrect IP data for iOS signups. You don't see that bug at first. What you see is that several geo-based rules suddenly have higher false positive rates on that platform, or a model that uses IP-based features also seems to drift, or maybe analysts complain that events coming from mobile look weird. If you only look at the rule level, you waste weeks recalibrating rule sets. But breaking it down by flow, you quickly realize the logic is fine on web, and the issue is isolated to iOS. Now, keep in mind that this isn't always a data bug. It might also be that a specific flow concentrates many good users who behave differently than your general population, which on its own can drive false positives up. But the point is, once you've bucketed your false positives by actor, do the same by user journey. Which product, which platform, which payment method, which geography, which specific funnel? I guarantee you'll see patterns emerge very quickly. The moment you start seeing patterns by flow, you will almost always run into the same culprit: data quality. Sometimes the underlying fraud logic is actually fine, and yet on a particular slice of traffic, your performance tanks. That's often because the system is operating on corrupted or missing inputs, default IP addresses, where the real IP failed to capture placeholder emails or garbage values, device IDs that reset to null on some OS versions, payment metadata that is never passed through for a certain method, timeout, or integration errors with third-party intelligence sources. You want to locate the exact data fields that are affected in those population segments you've identified. A reliable method I use all the time is to simply group data points by their value and look for values that have a suspiciously high count. And if you think about it, how many times have you done exactly that to uncover fraud patterns just to stumble upon a bug? Once you identify data quality issues, you have some detective work to do. First, we need to remember that often corrupted data points have cascading effects. Corrupted email field would, of course, corrupt email velocity checks that are based on it. So, your first task is to identify all the impacted data points and link those to your misbehaving rules or models. For most organizations, this exercise can prove incredibly hard, but you have a shortcut: your list of top offender solutions from step three.
Chen Zamir
Chen Zamir
07:41
All you need to do is to cross-reference your corrupted data points with the inputs used by your worst-performing rules. This will save you days of analyzing data skills by associating data issues with solutions. You will be able to complete the last link in the chain, attributing a value to each of these issues, just as you've done with offenders. You have a dollar amount, which you can put on that email bug, at least in terms of false positives. And now it's time to bring it all together. Now you have a complete map, not only of how many false positives you have, but also where and why they are generated. With that, prioritization becomes much more straightforward. You start with buckets that are large enough to matter, under your control, and not caused by data quality issues that you cannot fix. In my experience, in many organizations, these are likely to be a small number of high-volume, high-impact rules, one or two model thresholds that were set conservatively, maybe a specific case management policy that encourages overblocking, a couple flows where your system produces data issues. These are exactly the areas we'll focus on in part three, where we'll get into the mechanics of fixing your decision logic to be less trigger-happy. The work would only be effective because you've done the root cause analysis first and know how to invest smartly. For now, if you've gone from, we have false positives, to, we know where most of them come from, and which parts we can actually fix, you are already well ahead of most teams, and that's a good place to be at.
Chen Zamir
Chen Zamir
09:07
You.