SardineCon SF/2026

Learn More
The Saturday Fraud Strategist

Clase magistral sobre falsos positivos, parte 2: Cómo identificar de dónde se originan los FPs

9 min

Uno de los errores más comunes que veo cometer a los equipos de fraude es atacar de frente los falsos positivos.

Un cliente se queja. El CEO dice que el modelo está bloqueando demasiado. Alguien abre un panel de control, ajusta el umbral de un modelo de fraude, quizá retoca algunas reglas de fraude y, de repente, todos sienten que se está avanzando.

Honestamente, no se ve bien.

No porque reducir los falsos positivos sea un objetivo equivocado. Es, sin duda, el objetivo correcto. El problema es que la mayoría de los equipos pasan enseguida a lo táctico. Y si tu enfoque para reducir los falsos positivos es táctico, probablemente solo deberías esperar mejoras tácticas.

En este episodio, abordamos la segunda parte de la masterclass sobre falsos positivos: cómo desglosar los modelos de detección de fraude con falsos positivos en grupos que realmente puedas priorizar y corregir. Analizamos de dónde provienen los falsos positivos, cuáles están impulsados por modelos de detección de fraude, reglas antifraude, revisión manual, socios upstream, analistas de fraude, problemas de calidad de datos, señales de fraude corruptas y flujos de trabajo de detección de fraude en pagos.

Cuantificar los falsos positivos es útil.

Pero no es un plan.

Lo que escucharás en este episodio:

  • Por qué reducir los falsos positivos requiere un análisis de causa raíz y no solo ajustar el modelo
  • Cómo identificar quién realmente rechazó el evento: una regla, el umbral de un modelo de fraude, un agente de IA, un analista de fraude, un equipo de revisión manual, el emisor, el adquirente o un proveedor de soluciones antifraude
  • Por qué los socios de pago upstream pueden generar falsos positivos que tus propios sistemas de prevención de fraude no pueden corregir directamente
  • Cómo se desglosa la toma de decisiones sobre fraude entre la detección de fraude en pagos, la puntuación de riesgo de fraude y los flujos de trabajo operativos
  • Por qué la optimización de los sistemas antifraude comienza identificando a los peores infractores
  • Cómo los problemas de calidad de los datos, las señales de fraude corruptas y el desvío de los modelos generan falsos positivos que parecen riesgo de fraude
  • Cómo los equipos de operaciones contra el fraude pueden priorizar los grupos que sean lo suficientemente grandes, lo suficientemente corregibles y lo suficientemente valiosos para abordar primero

Quién debería escuchar:

  • Líderes de operaciones de fraude que buscan mejorar la precisión en la detección de fraudes
  • Analistas de fraude que trabajan con colas de revisión manual
  • Equipos de riesgo que gestionan las reglas de fraude, los umbrales de los modelos de fraude y la puntuación de riesgo de fraude
  • Equipos de ciencia de datos responsables de los modelos de detección de fraude y del desvío de los modelos
  • Equipos de detección de fraude en pagos que se ocupan de rechazos de emisores y decisiones de socios upstream
  • Equipos de prevención de fraude que intentan reducir los falsos positivos sin aumentar las pérdidas
  • Cualquiera que haya mirado fijamente un panel de falsos positivos y haya pensado: «Vale, ¿y ahora qué?»
Notas del episodio

Un patrón frustrante pero familiar

Los equipos de fraude suelen reaccionar ante los falsos positivos solo después de una queja de un cliente, una escalada a la dirección o una preocupación repentina de que “el modelo está bloqueando demasiado”.

Vale, es justo.

Pero ahora tienes que preguntarte: ¿estás solucionando la causa raíz o solo ajustando la parte del sistema que es más fácil de ver?

Esa distinción es importante.

Los sistemas modernos de prevención de fraude ofrecen a los equipos una gran variedad de herramientas

Los modelos de detección de fraude, las reglas antifraude, la evaluación del riesgo de fraude, la revisión manual, los agentes de IA, los sistemas de detección de fraude en pagos y los analistas de fraude desempeñan un papel en la detención de actividades sospechosas.

El mismo sistema que te protege también puede bloquear a buenos usuarios si no entiendes de dónde proviene realmente cada decisión.

La gran brecha es la visibilidad

Un falso positivo puede deberse a tus propias reglas de fraude. Puede originarse en un umbral conservador del modelo de fraude. Puede provenir de una revisión manual. También puede venir de un emisor, adquirente, procesador, proveedor de soluciones antifraude u otro socio aguas arriba.

Si no sabes quién dijo “no”, no puedes saber qué debes corregir.

Los falsos positivos no son solo errores del modelo

Se bloquea a buenos clientes. Los analistas de fraude pierden tiempo revisando casos que se podrían evitar. Las colas de revisión manual crecen. Los equipos de detección de fraude en pagos persiguen el problema equivocado. Los líderes de operaciones de fraude tienen dificultades para explicar por qué está aumentando la fricción para el cliente.

Y en algún punto en medio de todo eso, un usuario perfectamente legítimo se pregunta por qué tu sistema decidió que parecía sospechoso.

No es ideal.

Problemas de calidad de los datos

A veces la lógica antifraude está bien. El problema es que el sistema está funcionando con señales de fraude corruptas. Tal vez el campo de IP sea incorrecto. Tal vez falten los ID de dispositivo. Tal vez los metadatos de pago nunca se hayan transmitido. Tal vez un error en el SDK móvil esté haciendo que el tráfico de iOS se vea extraño.

De repente, tus modelos de detección de fraude se desvían. Tus reglas contra el fraude fallan. Tu puntuación de riesgo de fraude parece peor de lo que realmente es.

Y ahora el equipo está “optimizando el modelo” cuando el verdadero problema es una entrada defectuosa.

Clásico.

El camino a seguir

La mejor opción es agrupar los falsos positivos por actor, socio, recorrido del usuario, flujo del producto, plataforma, método de pago, región geográfica y tipo de problema de datos.

Prioriza los problemas que sean:

  • Lo suficientemente grandes como para importar
  • Realmente bajo tu control
  • No causado por problemas de datos que no puedas solucionar
  • Vinculado a reglas de fraude de alto impacto, umbrales de modelos o políticas de revisión manual

Así es como la optimización del sistema de fraude se vuelve estratégica en lugar de reactiva.

Conclusiones clave

Reducir los falsos positivos no se trata solo de los umbrales de los modelos de fraude, las reglas de fraude o la revisión manual. Se trata de comprender todo el sistema de toma de decisiones sobre fraude, incluidos los socios, las plataformas, los flujos de trabajo, los analistas, las señales de fraude alteradas y los problemas de calidad de los datos.

Una vez que sabes de dónde viene el problema, por fin puedes decidir qué arreglar primero.

Y, sinceramente, ese es un lugar mucho mejor en el que estar.

¿No estás listo para dejar la conversación sobre mi, y ojalá también tu, tema favorito? Suscríbete a The Saturday Fraud Strategist newsletter.

Episode transcript
Chen Zamir
Chen Zamir
00:08
One of the common mistakes I see fraud teams make is attacking false positives head on. They'll get a customer complaint, and they'll now trace back why this false positive happened and how to make sure it doesn't happen again, or the CEO complains that the model is blocking too much, and now they try to optimize it. The problem, of course, isn't the fact that they're trying to minimize their system's false positives. It's the fact that they're being tactical about it. And if you're tactical about how you do things, you can only expect tactical gains at best. In part one of the false positives masterclass, and if you missed that one, the link is down below, we accepted the reality that your system isn't built to report its own mistakes. We discussed several methods that can help you work around that and generate a usable picture of your false positives. Then comes the next problem. Quantifying false positives is merely an observation. It is not a plan. To reduce false positives in a meaningful way, you have to go one level deeper and ask a different set of questions. Where do they come from? Which are driven by my own system, and which my partners? Which are actually data quality problems masquerading as fraud risk? And most importantly, where should I start? That's what the second part of the series is about, breaking down your false positives into buckets that you can prioritize and act on. How do you do that? Here's my seven-step process. The first and most important step is to attach every false positive to the actor that made a decision. When a transaction is declined or an onboarding attempt is blocked, someone or something said no. That someone might be a rule in your system, machine learning model threshold, an AI agent, human analyst, manual reviewer, even a third-party partner, an issuer, acquirer, or fraud vendor. Without this categorization, you will default to optimizing the things you can see, usually your rules and your models. That's how teams spend months fine-tuning rules, only to later learn that most of their declines were coming from an issuer they never spoke to. So, your first task is to map every decline to the actor that made that decision. Sometimes it will be a single rule. Sometimes it will be a decision based on a certain score threshold. Sometimes it will be a human decision. Sometimes it will be a response coming back from a payment partner. The point is that you don't have to do this perfectly on day one. Even a rough breakdown makes an enormous difference, because it will help you get a sense for where the most value lies. Once you have that basic map, you can ask the following question: How much of this is even under my control? This is particularly important in card payments, where a single card transaction can pass through half a dozen hands before the issuer finally says yes or no. Between the cardholder and the issuing bank, you often have the merchant or platform itself, a payment service provider, an acquirer, an acquiring processor, a card network, an issuer processor, and finally the issuing bank. And that's even without counting the third-party fraud vendors that some of these actors plug into their own stack. Each of these actors can decline a transaction. Each has its own risk logic, and each contributes its own false positives. Now, if 60% of your false positives are driven by upstream partners, then your maximum sphere of influence is capped at 40%. That doesn't mean you should give up and walk away, but it does mean you should rethink your goals, how you set expectations internally, and where you spend your energy. And I think that this is an important point that I see a lot of fraud leaders miss. You cannot tune rules you don't own. You can, however, quantify the impact, make it visible, make sure everyone in the organization understands where the limits of your influence actually are. And trust me on that. This saves a lot of frustration later. And also, strategically, if you're unhappy with your partner's performance and believe it's substandard, you can always work to replace them. Once you have your decision sorted into high-level buckets, it's time to go one level deeper.
Chen Zamir
Chen Zamir
03:50
For example, you take a rule engine bucket that is responsible, say, for 25% of all false positives, and you break those 25% down to individual rules. And you do that as best you can for each one of your buckets. Now, when I say as best you can, what I mean is that realistically this can easily become time-consuming and low-value work. But let's try to break it down into the top five offenders in each category, and remember that even if a rule is relatively accurate, high volume usually means it is a primary source of false positives. High-volume decision is a good rule of thumb in terms of where to start. The point is that you want to identify the specific actors which are responsible for the most amount of false positives in absolute terms. These are your quick wins. At the same time, your gut might tell you that some offenders are flying below radar, legacy rules, outdated policies, or solutions that were never properly validated. Don't ignore that gut instinct, even if these actors have low volumes, or at least don't ignore it before you collect data on it. So, we now know what is blocking your users, but we still need to know where those users are coming from, even within the part of the stack you control. False positives are rarely evenly distributed. They tend to cluster around specific flows, such as mobile versus web, iOS versus Android, maybe different products or payment methods, or maybe even like new versus established users. If you only look at the actor that declined the event, you will see a rule or a model misbehaving. But if you look at which flow the user went through, you might find a deeper issue. Specific flow that produces corrupted or missing data, causing your entire fraud stack to misfire. Here's an example. Imagine you have an integration bug in your mobile SDK that sends incorrect IP data for iOS signups. You don't see that bug at first. What you see is that several geo-based rules suddenly have higher false positive rates on that platform, or a model that uses IP-based features also seems to drift, or maybe analysts complain that events coming from mobile look weird. If you only look at the rule level, you waste weeks recalibrating rule sets. But breaking it down by flow, you quickly realize the logic is fine on web, and the issue is isolated to iOS. Now, keep in mind that this isn't always a data bug. It might also be that a specific flow concentrates many good users who behave differently than your general population, which on its own can drive false positives up. But the point is, once you've bucketed your false positives by actor, do the same by user journey. Which product, which platform, which payment method, which geography, which specific funnel? I guarantee you'll see patterns emerge very quickly. The moment you start seeing patterns by flow, you will almost always run into the same culprit: data quality. Sometimes the underlying fraud logic is actually fine, and yet on a particular slice of traffic, your performance tanks. That's often because the system is operating on corrupted or missing inputs, default IP addresses, where the real IP failed to capture placeholder emails or garbage values, device IDs that reset to null on some OS versions, payment metadata that is never passed through for a certain method, timeout, or integration errors with third-party intelligence sources. You want to locate the exact data fields that are affected in those population segments you've identified. A reliable method I use all the time is to simply group data points by their value and look for values that have a suspiciously high count. And if you think about it, how many times have you done exactly that to uncover fraud patterns just to stumble upon a bug? Once you identify data quality issues, you have some detective work to do. First, we need to remember that often corrupted data points have cascading effects. Corrupted email field would, of course, corrupt email velocity checks that are based on it. So, your first task is to identify all the impacted data points and link those to your misbehaving rules or models. For most organizations, this exercise can prove incredibly hard, but you have a shortcut: your list of top offender solutions from step three.
Chen Zamir
Chen Zamir
07:41
All you need to do is to cross-reference your corrupted data points with the inputs used by your worst-performing rules. This will save you days of analyzing data skills by associating data issues with solutions. You will be able to complete the last link in the chain, attributing a value to each of these issues, just as you've done with offenders. You have a dollar amount, which you can put on that email bug, at least in terms of false positives. And now it's time to bring it all together. Now you have a complete map, not only of how many false positives you have, but also where and why they are generated. With that, prioritization becomes much more straightforward. You start with buckets that are large enough to matter, under your control, and not caused by data quality issues that you cannot fix. In my experience, in many organizations, these are likely to be a small number of high-volume, high-impact rules, one or two model thresholds that were set conservatively, maybe a specific case management policy that encourages overblocking, a couple flows where your system produces data issues. These are exactly the areas we'll focus on in part three, where we'll get into the mechanics of fixing your decision logic to be less trigger-happy. The work would only be effective because you've done the root cause analysis first and know how to invest smartly. For now, if you've gone from, we have false positives, to, we know where most of them come from, and which parts we can actually fix, you are already well ahead of most teams, and that's a good place to be at.
Chen Zamir
Chen Zamir
09:07
You.