The Saturday Fraud Strategist

フォルスポジティブ・マスタークラス 第3回:システム内部の誤検知を減らす方法

9 min

さて、誤検知を減らす話をしましょう。ほとんどのチームは、いきなり具体的な対策に飛びつきたがります。ルールを調整する。しきい値を変える。例外を追加する。微妙なグレーケースを手動レビューに回す。なるほど、それらはどれも役に立つかもしれません。でも正直なところ、そこから始めてしまうと、おそらくは手探りでやっているだけです。

それに、不正防止の現場で勘に頼るやり方は、正直いって私の好みの運用モデルではありません。うまくいかないからではなく、たまにうまくいってしまうからこそ厄介なんです。そうなると、みんな自信満々になってしまう。それはあまり良い状態とは言えません。

今回のエピソードでは、False Positives マスタークラスの続きとして、計測やバケット分けの話から、みんなが本当に知りたい「おかしな動きをしているシステムの部分をどう直すか」というテーマに進みます。ただ、目的は単に誤検知(フォルスポジティブ)を減らすことではありません。誤検知を減らしつつ、3週間後に損失が顕在化してから「このダッシュボード、裏切ったな」と全員が黙って見つめ始めるような、新たな不正問題を生み出さないことが本当のポイントなのです。

このエピソードのテーマは「規律」です。内容は、手動審査、不正検知ルール、不正モデルの精度、不正モデルの再現率、シャドーモードでのテスト、データ品質の問題、そして最終的にはどの不正対策チームも必ず向き合わなければならない、不快だけれど避けて通れない問い――「このルールは本当に役に立っているのか? それとも、2022年のあの不正急増のときの印象に引きずられて、ただ感情的に執着しているだけなのか?」という問いについてです。

このエピソードでお届けする内容:

  • 誤検知を減らすには、本能ではなく手動でのレビューから始めるべき理由
  • 不正検知ルールを削除・緩和・改善のどれにすべきかを判断する方法
  • ルールが不正を検知しても優良ユーザーに悪影響を与えてしまう場合に、不正検知モデルの適合率と再現率がなぜ重要になるのか
  • 不正利用者の抜け道をうっかり作ってしまわないように除外条件を設計する方法
  • リリース前にシャドウモードテストとチャレンジャールールが不可欠な理由
  • データ品質の問題が、本来は妥当な不正防止ロジックを誤作動させてしまう可能性について
  • なぜデータが壊れているとき、詐欺対策チームには洗練さより実務的な対応が求められるのか

このエピソードは次のような方におすすめです:

  • 不正対策ルール、不正検知ルール、モデル、AIエージェント、または審査フローを担当している
  • 不正被害を増やすことなく、誤検知を減らそうとしている
  • 誤検知率が高いものの、その原因となっているシステムのどの部分かが分からない
  • 手動審査のサンプルや主要な問題ケースを、より体系的にレビューできる方法が必要です
  • 破損したデータやノイズの多いシグナル、あるいは不正防止ロジックが誤作動を繰り返すフローに対処している
エピソードの概要と主なポイント

誤検知を減らすには、当てずっぽうではなく、まず実際に目で確かめることから始める

このエピソードで最初にして最も居心地の悪いポイントは、同時にいちばん単純なことでもあります。ダッシュボードを眺めているだけでは、誤検知の問題を理屈だけで解決することはできません。実際に見る必要があります。手作業で。本物のイベントを。派手さはまったくないですよね。誰もが、不正対策の仕事に就いたのは、ブロックされた100件のケースを精査して「このルールはなぜ存在するのか」と自問するためではなかったはずです。でも正直なところ、有益な仕事はまさにここから始まるのです。

ルールやモデル、あるいはAIエージェントが大量の誤検知を出している場合、最初のステップは、ブロックされたもののサンプルを手作業で確認することです。多くの場合、50〜100件ほどのイベントを見れば、おおよその方向性はつかめます。ここで論文を書く必要はありません。あなたがやるべきことは、実務的な問いに答えることです。

  • このルールは、本当に私たちが考えているほど不正確なのでしょうか?
  • この規則は、今もなお存在する価値があるのでしょうか?
  • それは今でも自動的に判断を下すべきでしょうか?
  • 代わりに手動審査に回すことはできますか?
  • 実質的に不正による損失を増やすことなく、誤検知を減らすにはどうすればよいですか?

重要なのは、すべての答えがあなたのビジネス次第だということです。「これなら十分」といえる普遍的な基準は存在しません。手動でレビューできる体制があるチームもあれば、まったくないチームもあります。アナリストに回すには件数が多すぎるため、うるさく感じるルールでも依然として必要な場合があります。うんざりしますか?そのとおりです。でも、それが現実です。

不具合のある不正検知ルールに対する3つの対処パターン

証拠を精査すると、不具合を起こしている不正検知ルールやモデルのほとんどは、3つのカテゴリーのいずれかに分類されます。ここから先は、作業がより整理しやすくなります。

最初のカテゴリーは、本来まったく存在すべきでないルールです。これは「低いところにぶら下がっている果実」のような、手をつけやすい対象です。たいていの場合、危機的な状況の中で作られ、一度なにかを検知したあと、そのまま誰も触りたがらずに本番環境に永遠に残り続けます。とても人間的で、とてもよくあることですが、決して良い状態ではありません。もし不正検知としてのカバレッジがごくわずかで、誤検知(フォルス・ポジティブ)が多いのであれば、そのルールを削除することが、最も安全で、かつすっきりとした解決策になり得ます。

2つ目のカテゴリーは、「存在すべきだが、もはや自動で意思決定すべきではない」ルールです。これは中程度の精度のルールで、意味のある不正は検知できるものの、自動拒否を正当化できるほど十分には安定していません。もし不正対策チームに余力があるなら、そのロジックをケース管理側に移すことで、不正検知のカバレッジを維持しつつ、誤検知(誤判定)を減らすことができます。人間をプロセスに介在させることで、その担当者がリアルタイムでパターンの変化を見抜けるようになります。

トレードオフはコストです。目視による審査は有用ですが、無料ではありません。誤検知は減らせますが、その分オペレーションの負荷は増えます。その妥協が本当に見合うのか、自分たちに問いかける必要があります。

3つ目のカテゴリは、最も一般的で、かつ最も興味深いものです。それは「存在すべきだが、改善が必要なルール」です。これらのルールは、リコール(検知率)は高い一方で、精度は低いことがよくあります。平たく言えば、不正はしっかり捕まえるものの、善良なユーザーまで巻き込みすぎてしまうのです。ですから、取るべき対応はルールを削除することではありません。取るべき対応は、それらを洗練させて改善することです。

誤検知の削減は、逆方向から行う不正検知である

誤検知を減らすことを考えるうえで有用な見方のひとつは、それが本質的には「逆向きの不正検知」であるということです。

不正検知ルールを作成する際は、通常、まず実際に発生した不正事案から始めます。不正事案を精査し、それらに共通するパターンを特定し、そのパターンを正常な利用者全体と比較したうえで、正当な取引をあまり巻き込まずに不正だけを検知できるロジックを構築します。

誤検知を減らすには、目的を逆にして同じことを行います。まずは誤検知であると確認されたケースから始め、それらに共通するパターンを特定し、そのパターンを不正のケースと比較したうえで、リスクを増やしすぎることなく、優良なユーザーを不正から切り分けるための除外条件を構築します。

ここでチームがトラブルに陥ることがあります。「多くの誤検知はXのように見える」と言うだけでは不十分です。なるほど、有用ではあります。しかし、不正のケースもXのように見えるのでしょうか?もしそうなら、その除外条件は除外になっていません。それは、見た目を少し整えただけの不正への招待状です。

最後のその一歩が重要です。本当に大事です。

不正の抜け道を生まない除外ルールの作り方

このエピソードの例は、あえてシンプルにしています。「IP の国がアカウントの国と一致しなければ拒否する」というものです。これは基本的な不一致ルールであり、たしかに現実のルールはもっと複雑なことが多いです。ただし、私たちが思いたがっているほど、いつも複雑というわけでもありません。

このルールでブロックされた100件のイベントを手動でレビューするとします。そのうち30件は不正で、70件は正当な取引でした。その70件のうち20件は、IPアドレスがカナダにあるように見える米国ユーザーによるものです。これは意味のある誤検知パターンだと言えます。

さて、次の問いが重要です。あなたの不正事案にも、米国アカウントでカナダのIPという同じパターンが見られますか? もし見られないのであれば、きれいに分離できている可能性があります。通勤者がいるのかもしれませんし、VPNがカナダのエンドポイント経由でルーティングしているのかもしれません。ユーザー層が国境付近に住んでいるのかもしれません。理由が何であれ、誤検知のパターンは実在しており、不正との重なりは小さいと言えます。

これにより、より安全な除外ルールの根拠が得られます。IP の国がアカウントの国と一致しない場合は拒否し、ただし IP の国がカナダでアカウントの国が米国の場合は除外対象としない、という条件です。

あるいは、あなたのビジネスの展開状況や利用可能な機能によっては、それを周辺国まで含めたロジックに広げることになるかもしれません。重要なのは、具体的なルールそのものではなく、その方法論です。除外条件は、誤検知(偽陽性)の集団から導き出し、不正の集団に対して検証されるべきものです。

そうやって誤検知を減らしつつ、新たな不正の入り口をうっかり作らないようにするのです。みんながこのステップを踏んでいると考えるのは、さすがに楽観的すぎるでしょうか。たぶんそうでしょう。

リリース前に、すべての不正検知ロジックの変更をテストする

たとえ「単なる」除外であっても、不正防止ロジックを変更するたびに、リスクの姿勢は変化します。新しいルールをリリースするのと同じものとして扱ってください。

ここではシャドウモードでのテストとチャレンジャールールが最も頼りになる存在です。現在のルールは有効にしたまま、改良版を並行して走らせてください。そして、その差分を計測しましょう。

知りたいのは次の点です。

  • 新しいバージョンでは、どれくらいのイベントを追加で利用できますか?
  • それらのイベントのうち、最終的に不正行為に発展するのはどれくらいありますか?
  • その変更によって、どれくらいの誤検知を防ぐことができますか?
  • 新しいロジックは時間が経っても一貫して動作しますか?

結果が維持されるようであれば、より自信を持って本番導入できます。そうでなければ調整してください。華やかなものではありませんし、大げさなローンチの瞬間でもありません。ただ、より良い不正対策の運用になるだけです。

モニタリング期間は、あなたの環境で不正がどれくらいの速さで顕在化・成熟するかに応じて決めるべきです。時間に余裕があるなら、チャレンジャー版と元のバージョンで不正の結果を比較できるだけ十分に長く待ちましょう。時間的なプレッシャーがある場合は、ランダムサンプルを手作業でレビューしてください。完璧ではありませんが、何も分からないまま進めるよりははるかに良い方法です。

データ品質の問題によって、優れたロジックでも悪く見えてしまうことがあります

問題は必ずしもルールそのものではなく、そのルールに入力されるデータである場合があります。

これは厄介な問題です。というのも、誰もが不正検知ロジックの変更で解決したがるからです。ルールを変える、スコアを下げる、例外を追加する、とにかくリリースする。しかし、もし基盤となるフィールドが破損していたり、欠落していたり、一貫性がなかったり、特定のフローで壊れていたりする場合、そのルールは設計どおりに動いているだけかもしれません。単に入力データが悪いのです。

なるほど。面倒だけど、知っておいて損はないね。

データに関する問題が自分たちの管理下にあるのであれば、長期的に最も良い解決策は、そのデータ品質の問題そのものを修正することです。具体的には、API ペイロードを修正したり、SDK のバージョンを更新したり、社内の連携部分を直したり、プロダクトやエンジニアリングチームと連携してギャップを埋めたりすることが考えられます。数日かかるかもしれませんし、数週間かかるかもしれません。しかし、ひとたびデータが修正されれば、不正検知ルールは再び正しく動作し始める可能性があります。

また、同時にほかの問題も解決できるかもしれません。データの問題が台無しにするのは、たいてい一つだけではありません。たいていはシステム内をひっそりとさまよいながら、いくつものものごとを悪化させていきます。最悪の意味で、非常に効率的なのです。

データをすぐに修正できない場合は、ロジックを調整する

すべてのデータ品質の問題が、短期的に解決できるわけではありません。原因が外部パートナーにあることもあれば、自分ではコントロールできないプラットフォームにあることもあります。形式上は自分たちが管理しているシステムが原因でも、修正は14個のロードマップ上の優先事項のさらに向こう側に追いやられ、「来四半期こそ」と言い続けるチームがひとつあるだけ、ということもあります。見栄えはよくありませんが、よくある話です。

すぐにデータを修正できない場合は、そのフローに対する不正検知ロジックを調整しましょう。たとえば、その問題のあるフローをルールの対象から外す、ルールの重みを下げる、ルールの影響度を変更する、あるいはそれらのケースを手動審査に回す、といった対応が考えられます。

ここでは、不正防止は洗練さよりも実用性が求められます。きれいなシステムであることは望ましいですが、きちんと機能するシステムのほうが重要です。正解が必ずしも理論的に最も満足のいくものとは限りません。ときには、不正リスクを抑えつつ、優良なユーザーの利用を止めないことこそが正しい答えになることもあります。

最終的なポイント:

誤検知を減らすことは、システムを甘くすることではなく、より精度を高めることです。

つまり、実際に起きた事象を精査し、削除すべきルールと、優先度を下げるか改善すべきルールを切り分け、根拠に基づいて除外条件を作成し、それらの除外条件を不正事案に対して検証し、シャドーモードで変更をテストし、問題がロジックそのものではない場合にはデータ品質の問題を修正する、ということです。

ともあれ、少し耳の痛い結論はこうです。根本となる事例を確認せずに誤検知を減らしているだけなら、それはチューニングではありません。ただの当てずっぽうです。

もしかしたら運よくうまくいくかもしれません。たぶんね。

しかし、不正防止において幸運は本当の意味での管理手段ではありません。それは一時的な状態に過ぎないのです。

つながる:Chen Zamir | LinkedIn

「The Saturday Fraud Strategist」のホスト

フィンテック企業がより賢い不正防止対策を構築できるよう支援します

『The Fraud Fighter’s AI Playbook』の共著者

私の、そしてできればあなたの一番好きなテーマについての会話を、まだ終わらせたくありませんか? ぜひ購読してください:The Saturday Fraud Strategistニュースレター。

Episode transcript
Chen Zamir
Chen Zamir
00:03
When people talk about reducing false positives, they usually jump straight to tactics: tuning rules, adjusting thresholds, adding exemptions, redirecting edge cases to manual review, and so on. All of that is important, but it shouldn't be where you start. In the previous parts of this masterclass, we've covered the two pieces of work that most teams skip: how to measure your false positives, and how to map them into buckets by actor, by solution, by flow, by data quality, and by what is or isn't under your control. That's the groundwork. Now you're ready to tackle the main course, fixing the parts of the system that are misbehaving. But this is where we should approach the work with discipline. Otherwise, you'll end up making changes that don't matter, overlook changes that would have mattered, or worst of all, open the door for additional fraud without realizing it. So let's get down to it.
Chen Zamir
Chen Zamir
01:07
One of the best ways to ground yourself before making any change is to go back to manual review. Pick a rule, a model, or an AI agent that creates a high number of false positives, and manually inspect a sample of the events it blocked. 50 to 100 should be enough in most cases. What you're trying to answer are often very basic questions. Is this rule actually as inaccurate as we think? Do we still want this rule to exist at all? If it exists, should it still make automated decisions, or should it fit into manual review instead? And if it should remain automated, what would it take to reduce its false positives without materially increasing our fraud losses? Each of these questions is business specific. There is no universal threshold for good enough. Sometimes your fraud ops team can take on the additional workload. Sometimes you have no manual reviews at all as part of your process. And sometimes a rule needs to remain automated, even if it's noisy, because you simply cannot afford to manually review that volume. But you need clarity before you change anything. The worst case is thinking a rule is fine because it looks logical, only to discover under review that it's blocking overwhelmingly legitimate traffic.
Chen Zamir
Chen Zamir
02:43
You can't reason your way through these situations, you have to look. But most importantly, this isn't about validating that you're fixing a real problem. It's about discovering how to do it.
Chen Zamir
Chen Zamir
03:02
Once you've reviewed the evidence, almost every misbehaving solution falls into one of three categories. The first is that the rule shouldn't exist at all. This happens more often than teams admit. Many times it would catch only a small population, would not block much fraud, but would generate a disproportionate amount of false positives. Usually it exists only because someone added it during a crisis, and no one ever bothered to revisit it since then. Removing such rules feels uncomfortable the first few times you do it, but if the fraud they catch is negligible and the false positives are significant, turning them off is one of the cleanest, safest ways to improve your system. So this category has your low-hanging fruits. The second category is of rules that should exist, but no longer make automated decisions. This is common with mid accuracy logic that still captures meaningful fraud, but not reliably enough to block without human review. If you have manual review capacity, and if this particular logic tends to produce events your analysts are comfortable judging, then flagging it to case management instead of auto decline can be the perfect compromise. You preserve fraud coverage, reduce false positives, and put a human in the loop that can observe pattern changes in real time. The downside, of course, is that it's resource intense. You have fewer false positives, but at higher operational costs. And to be honest, I try as much as possible not to resort to it. The third category is of rules that should exist, but they need to be improved first.
Chen Zamir
Chen Zamir
05:06
This is the most interesting scenario, and it's also the most common one. In this case, the logic captures fraud effectively, meaning the recall is high, but at the cost of impacting many good users, meaning the precision is low. So you need to refine it. The question is how?
Chen Zamir
Chen Zamir
05:31
The process for improving a false positive heavy rule is exactly the same as when you design a new fraud rule, just in reverse. Think about it. When you design a fraud rule, you review a set of confirmed fraud cases, identify the pattern they share, compare that fraud pattern to the general good population, and build logic that captures the fraud without catching too many good events. Now, to reduce false positives, you do exactly the same thing, but with the opposite goal in mind. You review a set of confirmed false positive cases, identify the repeating patterns, compare those patterns to the fraud cases to make sure you know how to separate them, and finally build exclusions that capture that separation without releasing too much fraud. And this last step is crucial. It's not enough to say a lot of false positives look like X. You need to prove that your fraud cases don't also look the same. Otherwise, your exclusion will undermine your fraud detection. But here's the thing: if you've worked through the steps I outlined without skipping any, you already did this manual review while validating the false positive metrics. Let's take an example. Suppose you have a very standard mismatch rule: decline if IP country mismatch account country. Obviously, given a simplified scenario, but keep in mind that in reality, many rules are not that much more sophisticated than that. Now imagine you manually review a sample of 100 events this rule blocked. After tagging them, you notice 30 were fraud, 70 were legitimate, and out of those 70 legitimate events, 20 involved US users whose IP addresses appeared in Canada. 20 out of 70 is a meaningful pattern. Maybe your business has a large US-Canadian commuting demographic. Maybe VPNs for Canadian endpoints are common. Maybe part of your user base works close to the border. For whatever reason, this specific mismatch, US account and Canadian IP, seems to produce a lot of false positives. If in your fraud cases you see little to no fraud coming from the same US-Canada pattern,
Chen Zamir
Chen Zamir
07:59
Then you have a clean separation. That means you can safely introduce an exclusion. Decline if IP country mismatch account country and not IP country equals Canada and account country
Chen Zamir
Chen Zamir
08:18
Equals US. Or better yet, broaden it slightly to include general neighboring country logic, depending on your business footprint and available features. The point is that the exclusion is derived from reviewing your false positive population, and is validated against your fraud population. This is how you refine rules without accidentally creating backdoors for fraudsters to exploit.
Chen Zamir
Chen Zamir
09:10
Anytime you change a fraud logic, even if it's just an exclusion on an existing rule, you're changing your posture. That's the same as releasing a new rule completely from scratch. So before you release the new version, test it. Shadow mode and challenger rules are your best friends here.
Chen Zamir
Chen Zamir
09:36
Keep the existing rule active, but run your refined rule in parallel. Measure the differences. How many additional events would it have allowed to go through? How many of those events matured into fraud? How many false positives would it have prevented? If the results hold over time, you can deploy with confidence. If not, adjust. How much time should you monitor these changes? Ideally, enough to gain a good measure of how fast fraud matures in the new version, and whether that's in line or not with the original version. And if you're really pressed for time, you may want to consider manually reviewing random samples.
Chen Zamir
Chen Zamir
10:25
As we explored in part two of the masterclass, sometimes the root cause of false positives is not faulty logic, but corrupted data. Best case, you've already identified it in your groundwork analysis. Worst case, you went through all of the motions just to figure it out on your detailed manual review. When this is the identified root cause, you have two options, depending on whether the data issue is fixable in the near term. Option one is to fix the data quality issue itself. If the corrupted field is under your control, be it your SDK, your internal integration, or your infrastructure, then the best long-term solution is to repair the data. That might mean fixing an API payload, adjusting an SDK version, or working with a product team to plug a hole. It may take a few days or a few weeks, depending on your organization. But once the data is fixed, the rule will often start behaving correctly again. Also, it's likely that you also solved a whole bunch of other issues that were plaguing other solutions at the same time. Your second option is to adjust the fraud logic for that flow. Not all data issues are internally fixable. Sometimes they originate from external parties that you just cannot control. Or, in the most frustrating cases, you do control these issues, but there's no fix in sight. If you cannot fix the underlying data soon, then you may need to adjust your fraud logic specifically for that flow. You have a couple of choices here. Either exclude the problematic flow from the rule altogether, lower the rule's weight or change its impact, or move those cases to manual review.
Chen Zamir
Chen Zamir
12:19
Whatever you do, just keep in mind that fraud prevention should be more pragmatic than elegant. And sometimes it's not about being right, it's about being smart.