SardineCon SF/2026

Learn More
The Saturday Fraud Strategist

Cuando un ejecutivo de datos se une a un equipo de fraude, con Shachar Meir

Retrato en blanco y negro de un hombre sonriente con gafas y barba.

Así que este episodio comienza con un ejecutivo de datos que entra en un equipo de fraude.

Lo cual suena como el comienzo de un chiste muy de nicho. Y quizá lo sea. Pero, sinceramente, también describe un problema real que la mayoría de los equipos de fraude conocen muy bien: dependemos de los datos para casi todo, pero la relación entre los equipos de fraude y los equipos de datos suele ser más complicada de lo que nadie quiere admitir.

En este episodio me acompaña Shachar Meir, asesor de datos y exdirector de datos en Meta, para hablar sobre analítica con LLM, analítica de autoservicio, infraestructura de datos y por qué los equipos de fraude no pueden simplemente poner IA encima de datos desordenados y esperar que eso se convierta en una estrategia.

Porque sí, la analítica de autoservicio impulsada por LLM suena increíble. Haz una pregunta en inglés sencillo, obtén una respuesta, avanza más rápido y evita esperar tres semanas a un equipo de datos con 900 prioridades. Genial. Yo también quiero ese mundo.

Pero entonces tienes que preguntarte: ¿de dónde salió la respuesta? ¿Con qué tablas se unió? ¿Qué definición utilizó? ¿Entendió el contexto del negocio? ¿Alucinó? ¿Hiciste la pregunta correcta desde el principio?

No es un detalle menor.

Esta conversación trata realmente de la brecha entre la promesa de las herramientas de analítica de IA y la realidad operativa de la analítica de fraude. Los equipos de fraude y riesgo trabajan en entornos adversariales. El problema cambia constantemente. La relación señal-ruido es mala. El costo de equivocarse puede ser enorme, ya sea por dejar pasar el fraude o por bloquear a usuarios legítimos. Así que, si quieres que la analítica con LLM realmente ayude, necesitas más que una interfaz llamativa. Necesitas bases de datos sólidas, gobernanza, claridad semántica y la humildad para empezar en pequeño.

Lo que escucharás en este episodio:

  • Por qué los equipos de fraude y riesgo son clientes de datos inusualmente complejos
  • Cómo los equipos de fraude pueden trabajar mejor con los equipos de datos en lugar de solicitar infinitas soluciones puntuales
  • Por qué la analítica de autoservicio fracasó tan a menudo antes de la llegada de los LLM
  • Qué cambia y qué no cambia cuando la analítica con LLM se convierte en la interfaz
  • Por qué hacer la pregunta de datos correcta importa más que obtener una respuesta rápida
  • Cómo el gobierno de datos, las capas semánticas y la calidad de los datos dan forma a los resultados de la analítica de IA
  • Por qué los equipos de fraude deberían comenzar la adopción de IA con una sola tabla curada, un solo caso de uso y un solo resultado medible
  • Cómo pensar en las versiones de una hora, un día, una semana y un mes de un proyecto de datos

Deberías escuchar este episodio si:

  • Trabajas en operaciones de fraude y dependes de los equipos de datos para implementar proyectos de análisis de riesgo o fraude
  • Están considerando análisis con LLM o agentes de IA para el análisis de datos dentro de su stack de fraude
  • Te han prometido consultas de datos con IA que suenan demasiado fáciles, probablemente porque lo son
  • Necesitan mejores formas de alinearse con los equipos de ingeniería de datos, analítica o infraestructura de datos de riesgo
  • Quieren una forma práctica de pensar sobre la gobernanza de datos de IA antes de implementar herramientas que afectan a usuarios reales
Notas del episodio y conclusiones clave

Los equipos de fraude no son clientes de datos habituales

Uno de los puntos principales que Shachar plantea en este episodio es que los equipos de fraude son diferentes de la mayoría de los stakeholders de negocio. Y lo digo con cariño. Casi siempre.

Los equipos de marketing, ventas y finanzas necesitan absolutamente buenos datos. Pero en fraude y riesgo, la complejidad es diferente. Puedes ejecutar enormes procesos por lotes sobre años de datos históricos, pero el resultado aún tiene que respaldar una decisión en milisegundos. ¿Esta cuenta es legítima? ¿Esta transacción es arriesgada? ¿Este comportamiento es normal o es el inicio de un ataque?

Ese es un entorno operativo muy extraño.

La relación señal-ruido suele ser mala. Los patrones de fraude cambian. Los usuarios legítimos se comportan de maneras que parecen extrañas. Los estafadores se comportan de maneras que parecen normales. Y a veces la diferencia entre una cuenta robada y un usuario real que viaja entre países no es nada obvia.

Por eso la colaboración entre los equipos de fraude y los equipos de datos tiene que ser más estrecha que una relación estándar de solicitud y respuesta. Los equipos de fraude no pueden simplemente pasarles tareas a los de datos y esperar que las ejecuten en producción. Los equipos de datos necesitan contexto. Necesitan entender la misión. Necesitan saber por qué una decisión es importante y qué ocurre cuando el sistema se equivoca.

De lo contrario, todos están cumpliendo técnicamente con su trabajo y, aun así, el sistema falla. No es lo ideal.

Las soluciones puntuales generan problemas de datos a largo plazo

Shachar habla de uno de los errores más comunes que ha visto en los equipos de fraude y riesgo: pedir soluciones puntuales en lugar de infraestructura.

Un equipo de fraude detecta un problema. Vuelven los vendedores maliciosos. Vuelven los compradores abusivos. Aparece un patrón de estafa específico. Un nuevo flujo de riesgo. Entonces le piden al equipo de datos una solución para ese único problema. En el momento, eso parece razonable. Pero si cada solicitud se convierte en un proyecto aislado, todo el sistema se vuelve más difícil de mantener, más difícil de escalar y más difícil de reutilizar.

La mejor pregunta es: ¿qué infraestructura haría más fáciles muchas de estas soluciones?

Muchos problemas de análisis de fraude necesitan bases similares. Hay que calcular señales de backend a partir del comportamiento histórico, almacenarlas en algún lugar accesible y ponerlas a disposición en tiempo real o casi en tiempo real. Si construyes bien ese pipeline una vez, no estarás empezando desde cero cada vez que aparezca un nuevo patrón de fraude.

Eso no significa que todos los problemas tengan la misma solución. Obviamente no. El fraude sería mucho más fácil si eso fuera cierto. Pero sí significa que los equipos de fraude deberían alinearse con los equipos de datos en torno a una infraestructura reutilizable, no solo en solicitudes individuales.

Así es como haces que el próximo proyecto sea más barato, más rápido y menos doloroso.

El tramo final de los modelos de fraude importa más de lo que los equipos admiten

Hay un momento muy familiar en la conversación en el que Shachar describe a equipos que construyen modelos sofisticados y luego usan un umbral estático casi al azar.

De acuerdo, si la puntuación del modelo es superior a 800, rechaza la transacción.

¿Por qué 800?

A menudo, la respuesta no es muy buena.

Esta es una de esas cosas que pueden hacer que una persona de fraude se quede mirando en silencio a la distancia. Un equipo pasa meses construyendo un modelo sólido, midiendo el rendimiento, ajustando las características y evaluando los resultados estadísticos. Luego, la capa real de decisiones de negocio recibe mucha menos atención.

Pero la última milla es donde el modelo entra en contacto con los usuarios. Es donde decides cuánta actividad fraudulenta detienes, cuántos falsos positivos generas y qué compensaciones está haciendo realmente el negocio.

El AUC, o área bajo la curva, puede ser útil. Pero no te dice si el modelo se está utilizando bien. No te dice si el umbral tiene sentido para distintas regiones, segmentos, productos o tipos de usuarios. No te dice si el resultado empresarial está mejorando.

Por eso la capa de estrategia es importante. Si no validas periódicamente cómo se está utilizando el modelo en producción, gran parte del trabajo que se invirtió en él puede desperdiciarse.

La analítica de autoservicio fracasó antes de la IA por una razón

La analítica de autoservicio no es algo nuevo. Existía mucho antes de la analítica con LLM. Paneles con filtros. Herramientas de BI. Herramientas internas diseñadas a medida. Menús desplegables por todas partes. El sueño era sencillo: permitir que los usuarios de negocio respondieran sus propias preguntas sobre datos sin necesitar SQL, Python ni un analista.

Y, sin embargo, a menudo fracasaba.

¿Por qué? Porque la analítica de autoservicio depende de suposiciones que rara vez son ciertas.

Das por hecho que los usuarios saben lo que quieren. Das por hecho que pueden explicarlo con claridad. Das por hecho que las palabras tienen un significado determinista dentro de la empresa. Das por hecho que “ingresos”, “caída” o “usuario activo” significan lo mismo para todos los equipos. Una idea bonita. A menudo falsa.

También supones que los usuarios entienden la complejidad de los datos. Por lo general, no es así. Y eso no es un insulto. Los datos encierran por todas partes un conocimiento tribal. A veces un analista sabe que debe excluir el tráfico de un país específico para cierto informe porque, de lo contrario, los resultados son ruidosos. El usuario de negocio quizá nunca sepa que se tomó esa decisión, pero el informe solo tiene sentido porque alguien la tomó.

El autoservicio también supone que las personas quieren hacerse responsables de obtener sus propios números. A veces no es así. Si una cifra es incorrecta delante del CEO, muchas personas preferirían que fuera el analista quien quedara bajo el autobús. Comportamiento humano. No es elegante, pero es real.

Así que, cuando decimos que la analítica con LLM resolverá la analítica de autoservicio, deberíamos hacer una pausa. La interfaz cambió. Los problemas subyacentes no desaparecieron.

La analítica con LLM puede hacer que las malas preguntas sean más peligrosas

Lo inquietante de la analítica con LLM no es que pueda responder preguntas en inglés sencillo. Esa parte es impresionante.

La parte aterradora es que puede responder a la pregunta equivocada con mucha seguridad.

Con los paneles tradicionales, las definiciones suelen estar codificadas de forma rígida. Puede que odies el panel, pero al menos está extrayendo datos de una fuente conocida y aplicando una lógica conocida. Con las consultas de datos mediante IA, el sistema puede decidir qué tablas usar, cómo unirlas, cómo interpretar la solicitud y cómo presentar la respuesta.

Si el usuario no entiende SQL, la estructura de los datos, las definiciones de las métricas ni el contexto del negocio, ¿cómo puede validar la respuesta?

Exactamente.

Aquí es donde las alucinaciones de la IA en analítica se convierten en un riesgo operativo real. La respuesta puede parecer pulida. La consulta puede parecer plausible. El resultado incluso puede ser lo suficientemente cercano como para parecer correcto. Pero si la pregunta estaba mal planteada o faltaba el contexto de los datos, el resultado aún puede ser incorrecto.

Y en las operaciones de fraude, las respuestas incorrectas no son solo un asunto teórico. Pueden llevar a malas reglas, malas revisiones, malas decisiones de modelos, cambios de políticas equivocados y un impacto real en los usuarios o en las pérdidas.

Por eso el punto de Shachar es tan importante: el trabajo de un analista no es solo responder preguntas, sino ayudar a formular las preguntas correctas.

Sinceramente, esa es la parte que deberíamos automatizar al final, no al principio.

El gobierno de datos no es opcional para las herramientas de analítica con IA

Una de las partes más útiles de esta conversación es la distinción entre dos enfoques de las herramientas de analítica con IA.

Un enfoque es: darle al agente de IA acceso al almacén y permitir que la gente pregunte lo que quiera.

Esa es la versión que pone a todo el mundo nervioso, o al menos así debería ser.

La otra opción es: primero construir bases de datos sólidas. Definir las métricas. Crear una capa semántica. Añadir metadatos. Aclarar qué significa cada campo. Seleccionar y depurar las tablas. Supervisar la calidad. Y luego usar analítica con LLM sobre un entorno de datos controlado.

Ese segundo enfoque es menos emocionante en una demostración. También es mucho más probable que funcione.

Las herramientas de analítica con IA necesitan los fundamentos incluso más que los humanos. Un analista humano a veces puede notar cuando los datos se ven extraños, recordar el contexto interno o hacer una pregunta de seguimiento. Un LLM puede simplemente producir la respuesta más probable. Y si el contexto subyacente es confuso, esa respuesta también puede ser confusa.

Así que, si un equipo de fraude quiere adoptar analítica de autoservicio impulsada por LLM, el primer proyecto no debería ser «conectar todo el almacén de datos». Debería ser un caso de uso confiable, un único conjunto de datos o conjunto de tablas altamente curado, definiciones claras, buenos metadatos y una forma de validar si la herramienta realmente está ayudando.

Muy aburrido. Muy necesario.

Empieza con un caso de uso, no con todo el universo de datos

El consejo de Shachar para la adopción de IA por parte del equipo de fraude es sorprendentemente práctico: no intentes abarcarlo todo.

Empieza con un solo caso de uso. Una sola pregunta repetible. Un solo conjunto de datos confiable. Una sola tabla o un pequeño conjunto de tablas seleccionadas con definiciones que el equipo realmente entienda.

Crea la capa de validación. Crea la capa de contexto. Añade metadatos. Define qué significa “transacción rechazada”. ¿La rechaza tu sistema de fraude? ¿La rechaza el emisor de la tarjeta? ¿Se rechaza por falta de fondos? Estas distinciones importan.

Luego prueba si la herramienta de analítica con LLM realmente ayuda. ¿Responde a la pregunta correcta? ¿Utiliza los datos adecuados? ¿Reduce el tiempo necesario para obtener información útil? ¿Apoya una decisión más segura? ¿Hace que el equipo de fraude sea mejor, o solo más rápido a la hora de equivocarse?

Esa última pregunta es importante.

Si el primer caso de uso funciona, amplía gradualmente. Añade dimensiones. Añade métricas. Añade más preguntas. Pero no empieces con la fantasía de que la IA puede pegar con cinta adhesiva todas las bases de datos, entornos, tablas y definiciones desordenadas de la empresa.

Eso no es una estrategia. Es un encogimiento de hombros muy caro.

La IA puede ayudar con el contexto, la lluvia de ideas y un pensamiento más claro

Este episodio no está en contra de la IA. Eso sería perezoso. Y además, incorrecto.

Shachar ofrece una distinción útil: la IA puede ser muy buena para aportar contexto externo, ayudar en la lluvia de ideas, la generación de ideas y para desafiar tu forma de pensar. Si necesitas una idea del rango típico de un indicador de negocio, la IA puede ayudarte a obtener un contexto orientativo. Si quieres reflexionar sobre lo que podrías estar pasando por alto, la IA puede ser útil. Si quieres poner a prueba un plan bajo presión, también puede ayudarte.

La clave es usarla como una herramienta de apoyo, no como un sustituto.

Eso significa que la IA puede ayudarte a afinar tu pensamiento, pero no debería reemplazarlo. Puede sugerir preguntas, pero tú sigues necesitando saber cuáles importan. Puede ayudar a plantear opciones, pero aún necesitas entender la decisión. Puede ayudar a redactar una consulta, pero todavía debes validar la lógica si el resultado afecta a usuarios reales.

Esto es especialmente cierto en fraude y riesgo, donde las decisiones pueden desencadenar rechazos, revisiones, fricción, daños a los usuarios, pérdidas financieras o exposición a incumplimientos normativos.

Así que sí, usa IA. Pero quizá no le entregues las llaves de producción y te desentiendas. Suena razonable.

La versión de una hora suele ser mejor que la versión de un mes

Hacia el final de la conversación, Shachar comparte un marco sencillo pero útil: todo proyecto tiene una versión de una hora, una versión de un día, una versión de una semana y una versión de un mes.

Esto merece ser copiado. Con todo respeto.

Muchos equipos técnicos, incluidos los equipos de fraude y riesgo, ven un problema e inmediatamente imaginan la versión de un mes. Cada caso límite. Cada integración. Cada fuente de datos perfecta. Cada complejidad. Y entonces el proyecto se vuelve tan grande que se estanca, pierde prioridad o muere silenciosamente en un documento de hoja de ruta.

En su lugar, pregúntate cómo sería la versión de una hora. ¿Puedes crear valor direccional rápidamente? ¿Puedes estimar la métrica? ¿Puedes avanzar sin fingir que es la versión final?

Luego mejóralo. La versión de un día es más pulida. La versión de una semana está automatizada. La versión de un mes es más completa.

Esto no es nuevo. Es mentalidad de MVP. Es ágil. Pero, al parecer, todos necesitamos volver a aprender cosas sencillas cada pocos años bajo un nombre nuevo. Bien.

Para los equipos de fraude que trabajan con equipos de datos, esta mentalidad es especialmente útil porque los recursos de datos son limitados. Si puedes demostrar valor rápidamente, validar la dirección y luego construir a partir de ahí, es mucho más probable que obtengas apoyo que si propones un proyecto de dos años que nadie entiende del todo.

Conclusión final

La analítica con LLM puede ayudar enormemente a los equipos de fraude, pero solo si las bases no son un desastre.

Si las definiciones de tus datos no están claras, te falta una capa semántica, tus métricas significan cosas distintas para cada equipo y tus usuarios no saben cómo hacer las preguntas adecuadas, la IA no va a arreglar eso mágicamente. Solo puede hacer que la confusión vaya más rápido.

Esa es la conclusión ligeramente molesta pero útil de este episodio: los equipos de fraude no deberían esperar a tener datos perfectos, pero tampoco deberían fingir que la IA elimina la necesidad de gobernanza de datos, calidad de datos y una colaboración inteligente con los equipos de datos.

Empieza en pequeño. Crea un caso de uso confiable. Alinea la infraestructura. Haz mejores preguntas. Usa la IA como apoyo para el contexto y el pensamiento, no como una vía de escape de la responsabilidad.

De todos modos, si tu estrategia de analítica con LLM es “conectar todo y ver qué pasa”, supongo que técnicamente eso cuenta como una estrategia.

Solo que no es una que me gustaría tener que explicar cuando falle.

Recursos y enlaces

¿Aún no quieres dejar de hablar sobre mi tema favorito y, con suerte, también el tuyo? Suscríbete a The Saturday Fraud Strategist newsletter.

Conecta con Shachar Meir | LinkedIn
Asesor de datos
Exdirector de datos en Meta

Conecta con Chen Zamir | LinkedIn
Anfitrión de The Saturday Fraud Strategist
Ayudando a las fintech a construir defensas antifraude más inteligentes
Coautor de “The Fraud Fighter’s AI Playbook

Episode transcript
Chen Zamir
Chen Zamir
00:09
Welcome everybody to another episode of The Saturday Fraud Strategist. And today I have a special guest. Actually, more than a guest, I have a very dear friend joining us, Shachar Meir. Shachar, I think we've known each other for about 15 years or so. And all of this time you've been working on the other side of the fence, specifically in the data space, data management space. So it's a pleasure having you here.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
00:24
Yeah. It's my pleasure. Thank you so much for having me.
Chen Zamir
Chen Zamir
00:40
The pleasure is all mine. So Shachar, I'm guessing that most of my audience isn't really familiar with you, as you're coming from outside of the space. So before we talk about why I invited you and the conversation that we'll have today, maybe you can tell our listeners more about yourself and your career so far, and what have you been up to lately.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
01:02
Yeah. So hi everyone. My name is Shachar. I'm a data executive with over 20 years of experience. I've been leading and managing data teams in smaller companies as well as bigger companies such as PayPal and Meta. At Meta, I was London data engineering site lead. I've grown the site from four data engineers when I started to more than 200 at the peak. I managed the data for a portfolio of products in ads with a run rate of $7 billion a year. And in my very last role I was a director of data engineering for all of Facebook, Instagram, and Messenger's trust and safety problems. And I left Meta about three years ago because I realized that over the past 20 years, I've seen a lot of companies that had a lot of data. They had good data teams, they had decent data platforms, and nothing was working. No one was happy. It was a complete disaster. So I took a step back. I asked myself, why is it that in 2023 we're still failing as an industry, and why do so many companies struggle to get value from the data? And I investigated that, and I realized two things essentially. One is that I have the playbooks and I know how to think to fix it, because that's what I've been doing my entire career. And secondly, that I can achieve a lot more impact if instead of being chief data officer one company at a time, I would work with multiple companies helping them implement my playbook. So that's what I do today. And the interesting thing is I come from a data background, but throughout my entire career I've actually been working with risk teams and teams that deal with adversarial environments in different places, both at PayPal and WorldRemit, which is another fintech. But interestingly also at Meta with trust and safety, which is a different type of adversarial environment, but it's essentially the same kind of playbooks and methodologies and ideas. And then in my consulting career as well. So I've seen it a lot from the other side.
Chen Zamir
Chen Zamir
03:00
Yeah, I'm sure. You mentioned adversarial environments and it made me laugh because I was thinking, who's the adversary? Is that the risk team or the fraudsters? But we'll get to that in a second. I can only say that I definitely like how you described your career switch. It definitely resonated with me. And I think that one of the things that I truly appreciate in our relationship is that we've met about 15 years ago, and I think we've worked there for probably like two or three years. Then we lost touch for probably nearly a decade, and we had very different careers. You just described your career at Meta. I went and started Frockster as a startup. But we both ended up really in the same place at the same time, basically starting our own consultancy firms. And this is exactly when we got back in touch. I think since then we are probably in touch every week, maybe sometimes, most of the time, daily. And Shachar, by the way, now that I'm thinking about it, he's one of the main sources of memes that I post on LinkedIn. So if you like those memes, know that Shachar is the one to thank for.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
04:21
Guilty as charged.
Chen Zamir
Chen Zamir
04:27
Yeah, it's been a journey, man.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
04:30
Definitely a journey. And I appreciate that.
Chen Zamir
Chen Zamir
04:33
Yeah. Same here. So Shachar, you mentioned that already, that you yourself, you're not a fraud fighter, but you've been working with fraud fighters for long periods of time in all sorts of different places in your career, both in-house as a manager, but also as an advisor. And actually, that's the main reason why I wanted to have you on the podcast, because I think that data is such a core pillar to the work of every risk manager, and especially fraud managers. And that's also an area where many, many times we have, let's say it nicely, challenges or hurdles to overcome. And I wanted to invite you to the show to basically discuss exactly that, both the relationship, but mainly how can we leverage our peers and how can we leverage our own tools or infrastructure so our lives as fraud fighters are easier. That's the aim. Let's see if we can grab some takeaways out of that.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
05:58
Yeah, sure. Actually, maybe I can just start by telling about my journey starting from PayPal, because I think that was the inflection point when it comes to fraud fighting. It was the first time I was actually part of a fraud team. And PayPal, back in the day, acquired a startup company that became the PayPal office in Tel Aviv, and I was essentially part of it. It was a group of brilliant people, absolutely brilliant, who couldn't get anything done. And the reason was that they were really struggling with PayPal's data platforms and infrastructure. They were trying to build things, and they managed to build some solutions, but they couldn't get anything into production. It was a complete mess. And I was hired. I don't think that I was hired for a very specific role in that sense, but I sort of identified that as an opportunity. And I thought to myself, okay, I'm going to become the bridge between the data technologies team and the risk team. And I tried to investigate and understand why it is not working, and why these people are not able to deliver anything, and why are the people in data technology, why do they keep complaining about risk? And they had plenty of complaints because risk was so data-intense and data-heavy, and they were killing the clusters and running processes that the data team could not even understand what these people were trying to do. So it was fascinating to see this relationship. And I thought, okay, if I can bring the data team a little bit closer to risk and the risk team a little bit closer to data, and I can understand the gaps on both sides, I can make the situation better. And I think that specifically, unlike other professions like, let's say marketing, sales, finance, where things are, I don't want to use the word straightforward, but you can get away with not having perfect data, and you can get away with not being super data-intense. In risk, it just doesn't work like that. First of all, a lot of the decisions you have to make are in real time and in the moment, sometimes with very limited history, sometimes with a lot of history, but then you're looking for anomalies that are not always very obvious. I think the signal-to-noise ratio is very, very bad in most cases. And sometimes understanding the difference between a stolen account and somebody who's just traveling the world and happens to be in a different country trying to make a transaction that doesn't look like their normal profile or history, there's a very thin line between these two and it's hard to distinguish. So I think that these people were actually some of the most data-savvy and tech-savvy people that I've seen, but at the same time, I think there was a constant misalignment. And I noticed two very interesting things. One of the mistakes the risk team was doing is they kept asking for very pointed solutions. So they were working on a process to detect returning malicious sellers and then returning offensive buyers, and they kept working on specific pointed solutions for specific risk problems that they've identified. Whereas what they should have really done, or what would have been good, and this is what we did when I came in, is instead of asking for those pointed solutions, we started thinking about what is the infrastructure that we need to enable all of that. And if you think about it, the infrastructure for different risk solutions is often the same. You want to calculate some flags in the back end, using the entire customer's history, keep them in a front-end system, and then be able to access that in real time. It's the same infrastructure. So if you build it once and you model it in the right way and you create yourself a pipeline or a channel that you can send data to your front-end system, you're 80% done from a data perspective. So I think that understanding when is the point that you need to switch from thinking about the individual solutions to the infrastructure that enables all of that is something that really accelerated the fraud team there. And the second thing is that I think people often boil the ocean, in the sense that they were trying to build the most prolific solution that they could to a specific fraud problem. Whereas if they started simple, proved that it works, and then iterated on that, it would also give them something to work with and it would also give them the infrastructure to iterate. And I think people spent months and months building something that they could have shipped the simple version of in a couple of weeks, see if it works, and take it from there.
Chen Zamir
Chen Zamir
11:01
And it probably solved 90% of the problem. I say it all the time. It's hard.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
11:04
Most likely. And actually, if I can add one last point, I think it was very interesting to see the risk teams, the way they worked when they were building solutions. There were some things that they really thought through all the way and they came up with very interesting heuristics and ideas, and they managed to create pretty good models. But at the same time, sometimes they would end up writing a rule that says, okay, if the model score is greater than 800, then decline the transaction. And every time I would see a static number somewhere, I would always challenge that in a nice way. I would just ask, hey, by the way, how did you come up with this number? And almost always, it was random. So they spent all this time building this model and coming up with something that was, as far as I could say, state of the art. And then the last mile always seemed a little bit like guesswork. Not always, but sometimes. And then I would say, but 800, is it the same 800 for a customer in Europe and a customer in Vietnam? And I would start challenging that and people would be like, yeah, that's a good point. So it would be like, why did you spend so much time building the model, but then so little thinking about how you're going to use it?
Chen Zamir
Chen Zamir
12:30
Yeah. I must say you touched a nerve for me here, because what you just described is one of the main things that I've been speaking about for ages now, I feel. And it seems that many times we invest so much, as you said, effort and sophistication in the model, and we measure the model in, for example, AUC, area under the curve, right? Like how curved the model performance is. While in the end, the curve and the AUC doesn't matter and doesn't dictate performance and business performance at all. In the end, you need to ask yourself how much losses did I stop versus, or for the cost of, how many false positives did I create? And around this question, what we call the strategy layer, how you basically implement the model and bring it to production so it touches users and influences business performance, in that area, most teams spend very little sophistication and very little effort. And not only that, but also they very rarely look at that decision periodically and validate that it is still true. And yeah, all of that effort that was invested in the model many times just goes to waste because you didn't really look at the last mile.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
14:05
We should actually change the name from AUC to PAUC, right? Potential area under the curve. Which is, if you use the model correctly, then this is how much you can get. But if you just do whatever, then you'll get whatever.
Chen Zamir
Chen Zamir
14:19
Yes, yes. Potential, probable, maybe area under the curve. Yeah. I want to first of all say that both of these things that you've seen very much resonate and feel familiar. I want to take a step back and maybe talk a bit more, because you mentioned that risk teams are different than your usual or your other clients, right? The marketing team, sales team, customer support, and so on. I'm just wondering, as a fraud fighter, how should I think about it? I'm not thinking, why should I care, but how is the fact that I'm different, or what about this relationship is different? How should I transform that into a different way in which I align with a data team or communicate with the data team? So how does this difference matter in reality?
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
15:50
I would say there are two ways I would like to think about it. One is that from a technical perspective, I think risk teams have extreme technical complexity, because you can run these batch processes offline in the background for hours and hours to crunch years of historical data, but then in the end you do that because you need the data to be available to make a split-second decision, under 200 milliseconds or 100 milliseconds, and make a decision on a specific transaction or specific account. So I think this is a different level of complexity, I would say, from most other use cases and most other teams. And I think the other thing is the cost of mistake. In marketing, if your marketing is inefficient, then your CPMs might be a little bit too high or the cost of acquisition might be a little bit too high, or you might be spending money acquiring the wrong customers and then churning them, not achieving the lifetime value that you could be achieving. So I'm not saying that's not a bad thing. It is, but it's not the end of the world. But if you think about the consequences of a risk team getting it wrong, then you're essentially funding the next, you know, if you're allowing bad transactions to come through, you might be funding the next 9/11, or you might be helping build terrorist groups in Africa. Or on the other hand, the impact of people traveling the world, trying to buy something not in their country, and then getting a transaction declined when they're in the middle of nowhere and they have no other means. It's a terrible experience for the end users. So either way, the implications and repercussions can be pretty bad. And it has implications on everything. If I use PayPal, for example, and my transaction is declined once, maybe I'll try again. But if it keeps getting declined all the time, guess what? I'm going to stop using PayPal. So you have implications on marketing, and you have implications on your company's finances, and you have implications on everything. And while you're supposed to be an enabler for the company and you're supposed to save the company, doing a poor job is actually hurting the company more than people realize. And in most cases it's very difficult to measure that. It's very difficult to measure the false positives. There are some other examples, right? I know you talk mostly about fraud fighters in fintech, but if you think about some of the problems that we dealt with at Meta, like suicide and self-injury, the cost of getting it wrong is somebody committing suicide. Or bullying and harassment or these kinds of problems, the cost is horrendous. So I think that it's a higher level of complexity than any other problem, and the cost of getting it wrong either way is going to be more difficult. And I think that what does that actually mean for the risk teams working with data? I would say a couple of things. First of all, I think that it means that the collaboration must be a lot tighter. And I also think that there's an opportunity here to connect them to the mission and to what you're trying to do. Help them understand. Don't just throw jobs and solutions over the fence and say, hey, can you just run it in production for me or whatever. Try to bring them on the journey. Try to make them a real partner of yours. They really want to understand what are all these things that you're building, and they really want to become enablers for the business. So bring them on the journey. I think that's the most effective way most risk teams can get the support that they need from the data teams they work with.
Chen Zamir
Chen Zamir
19:58
That's super interesting. First, it comes to my mind that, you know, you spoke about the complexity. But as you were describing how data teams work with sales and marketing organizations, for example, if you're measuring CAC or CPM, these are very known definitions and there are no half measures here. Either you know your CAC and you're calculating it properly, or you don't. And with risk, first of all, you're dealing with undefined situations. You know, I want to run a batch that catches all the accounts that made these and these transactions in this and this period, that look like this and that. And it might be that I can do something a bit more simple, as you mentioned before, kind of like a simpler solution, a more basic solution that would be 10x simpler to the data team to achieve. And I think, you talked about bringing the data team into the journey, and I think one big aspect here is that the data team, as any type of engineering resource that you usually deal with, they have a lot of demand, very little supply. I haven't seen a data team that is bored yet, or an engineering team, and AI didn't change it. I don't think it will change it soon. So there are a few things here. Can you actually make sure that the cost or the resources that you need are minimized? And that is something that you need to not negotiate, but align with the data team. How can you create the most, or how can you ship the most efficient solution? And the other thing here is also to have them bought into the result, into the mission. So they are on your side and at least they are interested in this project as well. So I completely see it. I wonder, speaking, you know, I mentioned AI. You know what, before AI, you talk a lot about self-serving, and not in the context of fraud and risk teams in particular, actually very generally speaking about data teams. And I think that is exactly where, in this conversation, this topic can come in.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
22:41
Yeah, so I talk a lot about self-service analytics, right? The ability of users or stakeholders to use their own data, like non-technical users, non-savvy users, getting access to data without the need of learning SQL or Python or other languages like that. And I think that self-service analytics is not a new concept. It's been around for over two decades, where BI teams in the past were building dashboards with lots of selectors where you could slice and dice the data in whichever way you wanted. We've always tried to give business users or stakeholders the ability to get their own data. And it always failed, right? So it's not a new concept, but it's making a massive comeback now with AI and LLMs, because now you got to a point that you can ask an LLM a question in plain English and get an answer. It will get translated to SQL or whatever, and then...
Chen Zamir
Chen Zamir
23:35
Yeah. But before we go into LLMs, because that's another whole thing that I want to discuss, why originally was it the case that it always failed?
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
23:48
Yeah, it's a good question. I've been pondering on that a lot because in my career I've tried several times to introduce self-service. It didn't work, so I kind of gave up on that idea in a way. Or it worked in very specific cases, but I sort of stepped away from that. And now all of a sudden it's coming back and everybody's asking for it. It made me think, hang on a minute, that's not a new idea. It's always been around. So why did it fail back then, and how do we know it's not going to fail again? I thought about it, and what I realized is if you think about self-service as a concept, for self-service to work, you're making a few assumptions that you assume are true. Number one, that users know what they want. Number two, that they know how to put it into words, and they can explain what they want in words. You assume that words have deterministic meaning, so that revenue always means the same thing and not six different things depending on who in the organization you talk to. You assume that users understand the complexity and nuances and the intricacies of the data, which is very rarely the case. I had a friend who was an analytics manager, I think in Wix or one of the website-building platforms. And he said, Shachar, our analysts, whenever users would ask for reports, in some cases they would just automatically exclude all traffic from China because that created a lot of noise. And as soon as you take it out, it kind of sorts out the report. But they would do it in only some cases, and they knew when they needed to do it. So it was kind of like tribal knowledge. And the users were asking for reports, they didn't explicitly mention anything about China, and they didn't know that this is what they're getting, but this is what the analysts knew. So understanding the nuances and the complexities of the data is not something that most users understand. They might have an idea of what it is that they want, but they cannot put it into words. And the last thing is that you actually assume that users want to be self-served and be held accountable for pulling their own data. And I've seen so many cases where I've built teams dashboards, and I said, hey, here's all your data. You can answer 80% of your questions in this dashboard alone, thinking that it's going to remove a lot of the ad hoc questioning. And actually, it didn't. I was very surprised. People still insisted that they want the analysts to pull the data. But then I realized that if they go and present a number to the CEO and the CEO says, this number is wrong, where did you get that from? If they can throw the analyst under the bus, it makes them look better. But if they are responsible for pulling this number, it makes them look like a bunch of clowns. So obviously they don't want that. And you also assume that people need a high degree of flexibility in the way they consume data, which is not always the case as well. Sometimes you do if you're a fraud analyst and you're now trying to understand a certain modus operandi and trying to investigate something. Fine. But if you're the chief risk officer, you probably want your scorecard or list of metrics that are always there, and it always looks the same so you can manage the business. So I think that if you look at the assumptions behind self-service, many of them just don't work. It is not obvious to me that it's an interface problem and that LLMs and chatbots and free-text, plain-English questioning will make it much different.
Chen Zamir
Chen Zamir
27:35
Okay, so maybe just to step back. So when you're describing a self-service system in the old days, right, three years ago, we're talking about systems that are essentially dashboards that allow, in this case, fraud teams, with a lot of features, sorry, not features, filters, with a lot of filters and a lot of different dropdowns and so on, to ask data questions without needing to either have SQL skills or have SQL permissions even to go into the data and query themselves. And you're saying that you actually never saw that being a successful solution, not in general and not specifically when it comes to risk teams.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
28:30
No, so I won't say never. Self-service can work, but I think where people get it wrong is I think unbounded self-service doesn't work. In the sense that people can just ask whatever about any topic, and they just get it. I don't think that's something that's going to work. At the same time, bounded self-service, if you bound it to a very specific domain, to a set of questions, the underlying data sets are highly curated and very specific and purposely created for that mission, I think this is where it can work. And we've seen examples of that. By the way, self-service, when I say self-service in the old days, it doesn't need to be dashboards. It can be purpose-built tools. So at PayPal, for example, you remember we had a review tool that the fraud analysts were using, where they could put in an account ID and they would get all information on the account, and they had a bunch of ways they could explore the data. But it was a very purpose-built tool that gave them specific views, and they had specific workflows around it. But again, you had to invest a lot of time learning how to use it and to use it correctly. They kept enhancing it all the time, where every time you want to support more views and more ways of people looking at fraud, I think that's something that you have to keep building all the time.
Chen Zamir
Chen Zamir
29:52
Yeah. And that's something that because of the adversarial nature and the fact that the problem keeps shifting, a data point that was interesting yesterday would not necessarily be interesting tomorrow. It's very hard to bound these kinds of use cases for fraud teams.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
30:13
Exactly. And because of that, the fraud analysts would always know SQL on top of that. But that was their go-to tool and that was always their starting point, if it was good enough and if their data was there. But it was meant to answer very specific questions in specific ways, with specific views and specific information. So I think that somebody did spend a lot of time thinking about what it is that they need and structure the data in a way that would be useful to them. I think that just open-ended self-service for people who are not very technically savvy and not very data-savvy and don't understand these nuances, my view is, it's a little bit delusional. But maybe some people were more successful than I have.
Chen Zamir
Chen Zamir
30:59
I wouldn't bet my money on that. But I want now to speak about the self-service concept as we are hearing about it today. And today self-service means having an LLM be your basically bionic extension and replace your querying knowledge, capabilities, skill, and querying the data directly and not through a dashboard. I want to share my view and see how you see it. So I'm sharing my view as a data and/or a fraud analyst, right? I think it's even worse. Honestly, it scares me. It scares me a lot because in the old days of self-servicing data tools, at least I know there's a single definition to where the tool gets the data from and how it processes it. There are very clear definitions for each data point that I'm seeing on screen. And this doesn't change. Obviously, it's all hard-coded, so there are no hallucinations. And supposedly now in this new age, the major advantage is that as a user, I can ask any question that I want in plain English, and I will always get an answer. And that's seemingly very powerful. But as a data analyst, I'm like, okay, I'm not sure exactly where you took the data, I'm not sure exactly how you joined the data, and I anyway need to go through the SQL myself and the data scale myself and see and validate that you've done it correctly. And what scares me about it is not that, okay, you can write a query faster and better than me, maybe. That's great. But most users, and that's exactly the point, most users don't have SQL skills. So this is what scares me about it. What's your view?
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
33:13
Yeah. First of all, I'll say that I think that we're observing the biggest revolution of our lifetime right now, technological revolution of our lifetime. What we're seeing is probably unprecedented, and it happens at a scale that is probably too fast for comfort for most people. I think that the ability to have a chat with a bot or whatever in your own language, just explain what it is that you want, and then get a result, is mind-blowing for most people. And if you look at sci-fi movies from five, 10, 20, 30 years ago, that's always how humans wanted to interact with computers, right? And it's finally happening. So in a way, that's amazing. We just describe some sort of an outcome and we get it. But I think when people think about using AI for self-service, there are two camps. There's one that says, I'll just give my AI agent access to my BigQuery or my Databricks or whatever, and then I would be able to ask whatever questions I want. I think that's setting themselves to fail, and big time. And I'll explain why. And the other camp, and if you've seen Anthropic's recent paper on how they do data analytics, and they say they do 95% of their analytics with conversational interfaces and all of that, the basis of that paper is you need to have very, very strong data foundations, good semantic layer, clear definitions, like you need a lot of data governance. So basically what they say is you really need to get the basics right to be able to do that. And that's a company that's presumably already doing it. I think that a lot of the way other people are interpreting the situation is, hey, I have this amazing capability, let's just give it access and see what happens. Now, these models have one very big flaw. They have many flaws, but one of them is they never say, I don't know what the answer is. They just give you the most probabilistic answer, even if this answer is delusional or has no grounds in reality. So I think that we have to be super careful. Now, when can that work? It can work where you define a semantic layer, you define clear metrics, you have very good definitions, you give it a lot of metadata and a lot of context about these are my tables, this is the information that I have, this is what everything means, and all of that. So it can basically do the mapping between your questions to the underlying data. And if you don't have that layer in between, and if you don't have the governance and controls and monitoring and all of that, then you're just leaving yourself to chance, which is not a good way of doing things. Now, the one thing that people also don't realize is that this is all great assuming that you know how to ask the right questions. But actually, if you ask the wrong questions, even if you get the right answers to the wrong questions, that's not helpful. And I think never in the history of mankind was it so easy to ask the wrong questions and get the wrong answers at such a scale. So I think we're getting into a point that every person, and I've seen it a lot managing analysts and data scientists and analytics teams, is often the best service you do for the business or to your stakeholders is not actually running the queries for them, but doing the framing and the ideation and challenging their question. Saying, you're asking for this, why do you need this data? Okay, so this is what you're trying to do. This data is not going to help you. You need something completely different. So that's the kind of detailing conversation. And people can say, yeah, but I can build a skill in Claude and have Claude ask me these questions and whatnot. Possibly. We need to prove that. But people assuming that they will just ask whatever and it will be the right question, and they will get the right answer, and they will know what to do with this information, there are a lot of assumptions going on here.
Chen Zamir
Chen Zamir
37:22
This is great because I always say that, in my mind, the job of an analyst is not to answer questions. It is actually to ask the right questions. When you ask the right questions, you will get the answers that you need. I want to double-click on this for a second because I think it is truly, truly important. When we, data professionals, I would call myself a junior data professional for a second, talk about this, it is very, very clear to us. It's very clear to us that you need to ask the right questions, and it's very clear to us what is the right question. Well, how do you know whether a question is correct or not? Most of our listeners, especially I would assume the ones that would benefit the most from this self-service, LLM-powered oracle, are not necessarily people who were trained and qualified and educated. And this is not something that you are born with. This is something that you need to learn through many, many years, to know how to ask the right questions and how to identify bad questions. And you managed hundreds of data professionals. You mentored hundreds of data professionals. What would be your advice for someone who is not a skilled data person? How to ask the right question?
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
39:10
It's a good question. I don't know if I have an answer to this, because I think that some people just naturally have that and some people don't. I think that it's something you can learn. I think you can learn that from being in an environment where you get challenged a lot and you get a lot of external feedback on everything that you do, and you're a person that can learn from that feedback and reshape their framing. There are some things that people who ask good questions naturally do, and they keep challenging themselves, right? How do I know this is the right thing? How do I know it's going to work? What am I trying to do here? Is the question I'm asking actually going to help me achieve the goal? I think that a lot of people go into this process being very unclear on what their goals are in the first place. And then they don't actually sense check and say, okay, this is the goal. Is this approach right now going to help me achieve the goal? And we talked about asking the right questions. I think that there are three things here, which is: is the question right? Am I asking the right question? Is the answer correct? Is this answer that I'm getting correct? And then, is the decision correct? Is my interpretation of the answer and what I'm trying to do with it correct? And I think one of the canonical examples for that is the survivorship bias story. I'm sure you know that from the Second World War, where the German Air Force, I think, were losing a lot of fighter airplanes. At some point, they tried to approach it from different angles and engineers tried to fix it, and a lot of people tried to fix it until somebody thought, hang on, maybe it's a statistical problem, right? Some airplanes come back, some don't. There's some statistical angle to it. So they brought a statistician and asked him, hey, can you help us fix it? And he said, yeah, of course. So what used to happen is they would shoot the airplanes and then there would be bullets hitting the airframe and leaving holes in the airframe. So he drew a picture of an airplane and said, just put a dot every place you see a hole, like a bullet. And they did it, and then there were some clusters of areas where there were no bullets at all, and then the rest of the airframe was full of bullets. And their conclusion was, that's great. So now we need to reinforce all these areas where the bullets are and we'll be fine. And he said, no, you need to do the opposite. These are the airplanes that came back despite the fact they had bullets in these places on the airframe. So if you have it in the tip of the wing, that doesn't matter. But if you have it in the root of the wing, then that's a problem for the airplane and it cannot generate lift. So you need to reinforce the areas where you don't see the bullets. I think it's a marvelous example of people asking the right question, getting the right answer, and then reaching the wrong conclusion.
Chen Zamir
Chen Zamir
42:18
Yeah. I think one of the things that, at least how I operate, and this is how I'm trying to mentor analysts around me, is to actually, before you get to the question, work back your different options. What are the possible decisions that you can reach? Which answer would you need to get that would make you choose each of these decisions, each of these options? And then based on that, ask the question that would actually be able to tell you which of these decisions you need to make. And many times I see that people not only ask the wrong questions, they ask the wrong questions in the sense that all of the answers lead them into the same decision, picking the same decision without even thinking about it in advance. Yeah, it's very difficult.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
43:18
I love that. And I think an interesting example is that sometimes just reframing your questions and your goals... Let's say on Facebook ads, there are obviously people who use that as a tool, fraudsters who use ads to recruit victims and all of that. So there was an entire business integrity team whose purpose was to find these fraudulent ads and take them down, and fraudulent advertisers and that sort of stuff. It's an amazing channel for fraudsters, by the way. I'm sure you know that. So I think that there are two ways you can think about this problem. One is, okay, they're going to create all these fraudulent ads. How do I find them? How do I identify them at a high accuracy? How do I block them? And that sort of stuff. Which is assuming that they're going to do it. But the other approach could be, how do I prevent them? How do I make it so that they don't want to do it in the first place? And I believe there is actually a very simple solution here, which is not data and has nothing to do with technology. Just changing the policy and saying every advertiser needs to put in $50 in their account before they can start advertising, or $100 or whatever. And then maybe the payment terms are net plus 30 or something. So you basically ruin the economics of being a fraudster on this platform. Make it not worthwhile for most fraudsters. The problem is, yeah, you're also hurting the business, right? Because you're also hurting the economics for a lot of small- and medium-sized businesses in foreign countries, where for them, $100 is, I don't know, 10x their marketing budget or anything they can afford right now. So there's a real tradeoff here. But I think this is a really good way. And if your approach in the first place is, how do I detect that, you're already going down a certain path and you're completely eliminating this path, this branch of how do I make it not worthwhile for them in the first place?
Chen Zamir
Chen Zamir
45:35
Yeah. Analysts tend to try and solve all of their problems through the data set. And sometimes the solution is outside of the data set. Look, I just want to say that this is not easy. We're both sitting here and trying to come up with maybe some guidance on how to ask the correct, the right questions. And honestly, there is no trick or tip that we can give that would just make you now understand the framework with which you need to approach these AI self-service tools. And I think, in my mind, that's also part of my conclusion, that there isn't a simple way to upskill or train employees, especially in fraud, to work with self-service tools. And for me, that is more of a cautionary red flag than, this is the right way to do that. But we bashed on AI for a bit. That was fun. Still, there are many, many ways, and you said, we're living in exciting times. There are many ways in which fraud and risk teams can leverage LLMs and AI, GenAI in general. And I can share what probably most of our listeners have experienced: that the promise is very enticing. Just run this agent on your data and get everything you've ever imagined or fantasized on. The reality, when you already paid and booked and signed the contract, is that it is not that easy to make it work well, and many times exactly because of data integration issues. And I wonder if you have some words of wisdom and perhaps even tips about how fraud teams should think about preparing for implementing AI and even realizing or testing whether they are ready to adopt AI.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
48:17
Yeah. I think I'll go back to what I said earlier. Don't boil the ocean and focus on the basics. I think that humans need good data sets, good data foundations, good consistent metrics, high data quality to operate. AI needs that even more. And the idea that we don't have very good data foundations, so we're going to let AI do that, that's going to make things worse. So if you're a fraud team, and by the way, I appreciate that most fraud teams sit in a very tough spot because they're very data-intense, they're very data-savvy, and they're usually not the owners of their data platforms. So they do need to work with the data teams, the data organizations, to get things done. So I would say, first of all, get the basics right. Get one table or one data set that you can trust 100%, that you feel like is highly curated, high quality, very clear definitions. Make sure you have good metadata around it. Don't try to open the floodgates and ask any questions at any point in time, but choose one use case, questions that are highly repeatable, and you want to try and solve that and do a proof of concept and show that this actually helps you, that this new tool is actually helping you operate better. That's one thing that I would say. The second thing is I think that AI is really great for external context. Just to give an example, I was working with one of my clients. They produce electronic appliances, and they showed me their sales dashboard. Basically, from the top of funnel to the bottom of the funnel, for people who did the checkout, the number was about 1% to 1.5%. And I asked the person that was showing me, is that a good number or a bad number? I don't know what it means. I don't know, from the number of people who visit the website, who click on a specific product, who go to the add to cart and then checkout and complete payment, I don't know if 1.5% is good or bad. And AI is actually very good at giving you a sense of that external context. It's not scientific, but it's actually very, very good. Like, what should I consider? I asked Claude, I gave the context, I explained this is the situation, these are the steps, this is what's happening, these are the type of electrical appliances they sell. And it said the typical range is between this number and that number. Now, do I trust it 100%? I'm not sure. Is it fully accurate? No. But it kind of gives me an idea of the lay of the land. So I think that's pretty good. I also think that AI is very good for brainstorming, for ideation. If you just need somebody to challenge what you're doing, to tell you what are the things you're not thinking of, I think these are really good use cases. To basically detail your plans or help you sharpen your thinking process. That's good. But don't let it do the thinking instead of you. Use it as an augmentation tool.
Chen Zamir
Chen Zamir
51:47
That's super interesting, and I've written all of that down. There's one thing that you said that surprised me. I mean, it didn't surprise me, but I think it might surprise a lot of our listeners. And I think that you glossed over it quite quickly. You said get one view, one table correct. And I think many organizations that I see, in general on an organizational level, think about AI as the duct tape, as the pipelines. We don't need to build a view because we can just give it access to all of these different tables. And usually these are not different tables, these are different databases, different environments, different technologies, and let AI figure it out. But you're saying, if I understand you correctly, no, you want to prepare everything on an infrastructural level. So all the data for one use case is collected in this one table, one data view, and you plug AI on top of that. Did I get it correctly?
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
53:08
Yeah, that's absolutely how I would do that. Now, it doesn't have to be one table. It can be a set of multiple tables. I don't know if you know about medallion architecture, where the idea is that it's an architecture that's become very prevalent in the last few years, where you build your data in three layers. So the first layer is usually called bronze, silver, gold. The bronze layer is basically all your data coming from all your data sources and goes through very minimal adjustments or permutation. So it's basically what we used to call a data lake. And then the silver layer is your model data. So you clean the data, you model it, you validate it, it goes through quality checks and all of that, and you structure the data in a way that is useful for the business. But this data, which is what most companies have as their data warehouse, still has a lot of variance. There are a lot of different definitions, there are a lot of different ways of interpreting this data. So then on top of that, you build the gold layer, which is what you use to power your dashboards or your AI use cases and machine learning use cases. By the way, when you train machine learning models, you didn't just train it on your dim customer table or fact transactions or whatever. You'd always build a smaller data set that is highly curated, and you spend a lot of time defining the metrics and validating them. And sometimes you apply all sorts of different things on them to remove anomalies and reduce noise and all of that. So I think that using your data for LLMs is no different. You want to give it a data set that you trust, that you know what's in it, to reduce the variance and the errors, because there's going to be a lot of errors. So you want to make sure that at least on the surface, you give it a data set where it's very clear who is the customer, what is their history, how many transactions, how much total payment, how many declined transactions, etcetera. So you give it all the information, you give it context and metadata that says this is a customer, this is how many transactions, this is the definition of declined transactions. It's declined because we decided to decline, not because the card issuer or whatever decided to decline. So you give it all this extra information to reduce the variance, and then when you ask questions, the answer is going to be in a smaller range of error, which makes it more trustworthy than in a massive range where you just don't know what you're getting.
Chen Zamir
Chen Zamir
55:39
Yeah. That is massive. This is so, like when you say it, it sounds so trivial, but I don't think that even I had it in my mind. Of course. If you're not sure that all of your team can ask the right questions, at least make sure that the AI can give you the best answers. And for that you need simplification.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
56:08
Yeah. And when you're a fraud team or a risk team that doesn't own their data foundations, then agreeing with the data team, building together a data set or set of tables that are highly curated with the right set of metrics, and working on that, is actually a good project. And you don't need to start with what's the ultimate version. Start with what can I produce in one week. Build a table with them. Build all the validations on top of that. Build the context layer. Build everything. So you have the entire infrastructure end-to-end on just one dimension or just one metric. Show that it can work and then gradually start incrementing from there. I think that's a much more sensible approach than coming up with a two-year project that nobody understands why you're even doing it.
Chen Zamir
Chen Zamir
56:57
That's awesome. So true. I really love it, Shachar. Wow, we went through a lot. So I want to take a moment and gather my thoughts and replay all the takeaways that I've written down. We started a conversation speaking about how fraud teams and risk teams in general are a different client, or a different stakeholder, when you look at it from the perspective of the data teams. And you mentioned two common mistakes that you see in this relationship. The first one is that fraud teams tend to request point solutions, and you say that a better approach would be to actually align on the infrastructure and not necessarily a specific use case. Because when you align on the infrastructure, you have a way better chance to first of all get it right, but also to make other requests much cheaper. The second mistake that you mentioned was boiling the ocean, basically going for a full-fledged, highly sophisticated, highly advanced solution, when you can probably deliver an MVP, basically a pilot, in a fraction of the cost and time, and probably still achieve a significant amount of the value, 80%, 90%, and so on. So that was the first thing that I've captured. We also talked about, in general, that data teams are a coveted resource, like any engineering resource in an organization. And so you want to align not necessarily on the project goals, but at least on the architecture and the proposed solution, because that can really cut down the resources needed and increase the ROI, so increase the chances that you will get your project prioritized. And secondly, also to sell the result, to sell the mission, to get the data team bought into the project so they are on your side as well. Then we talked a lot about self-service. We talked about the fact that in general, this approach fails often, and especially with fraud teams, because of the adversarial, or let's say dynamic, nature of the problem. So there are very few use cases and questions that you can ask that are really repeating themselves in exactly the same way, especially in fraud. And we talked about the fact that this is now getting way worse because of LLMs and the way that a lot of tools nowadays offer self-service data querying in the form of LLMs. And we talked extensively about all of the red flags and the issues that you may encounter here. But we mainly touched upon the point that the core skill that you need here is to ask the right question. And that is not something that you can just understand or learn very, very quickly, despite all of our best efforts. And that this should serve as a warning that a lot of your efforts here can lead to, let's say, not the results that you are expecting. And I would say one thing, you mentioned it later, but I think it's very important here to mention it as well. You need to think very clearly about the impact of the decision that you make. If you're shooting to a range, or if you're brainstorming, as you put it, or if you're doing maybe base research analysis of your population, okay, maybe you can use it and see if it leads you to a ballpark that is within the expected results. But if you're going to deploy a solution that eventually would take action on real users, and you are using LLMs for that without any sort of sanity checks, that can become very, very dangerous. You mentioned all of that as part of speaking about how to get ready to adopt AI, not necessarily in a self-service context, but in general. And other than this point that I just mentioned, you also mentioned two other points. You said start with one simple, clear use case. And for that use case, prepare one table that you can implement your agent or your LLM or whatever you want to run there. And this is so important. I see so many businesses skipping this and basically going, as you mentioned, trying to boil the ocean with their first go, and they get stuck and disappointed with AI. That's basically the cycle that I'm seeing in the last 12 months repeating itself over and over. I hope, dear listeners, that maybe this conversation has helped you gain some sort of glimpse into what goes on on the other side of the fence in the data team, but also how to approach these projects in a way that is more realistic and that has a better chance of success. Shachar, did I miss anything?
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
63:25
No, I think that's good. And maybe I can just add one small thing, which is I have this framework or idea of four versions of every project that I shared with my friend. He's managing the analytics in a company that's building software for psychotherapists in the U.S. We looked at his dashboards and I said, show me your top-line company dashboard, whatever. And he showed me a dashboard where the top-line metric was number of users. And I said, look, this metric is pretty useless. It's interesting to know, but what I care about is the market penetration rate. You're showing a number with no context. I don't know if this number is 1% of all potential users, or 20%, or 70%. I think that's what matters. And he said, you're actually right, but it's actually very difficult to get this number. And I said, well, it's not very difficult to get this number because there are four versions of every project: the one-hour version, one-day version, one-week version, and one-month version. So in one hour, you can go on Claude and ask how many psychotherapists are there in the U.S. You get a number, and you go to your dashboard, you divide your metric by this static number, and you call this estimated market penetration rate, where the word estimated is the key word here. And you've already made progress in the right direction. It's not going to be highly accurate, but it's probably going to be within the right ballpark, and it gives you an idea of whether you're closer to the 1% or to the 70% penetration rate. And then in one day, you can do a little bit more research and maybe there is a database, because I'm sure all these practitioners need to register somewhere. So there must be some database where you have a list of all of them, and you can say, okay, let's download this list once, tidy it up a little bit, get a more accurate number. And what this one day of work allows you to do is remove the word estimate. So now it's not an estimated market penetration rate, you can just call it market penetration rate. So the one-week version, you can probably build an automated process that will download this list periodically, update the number and all of that, so your dashboard is always correct and it doesn't drift. And then the one-month version, and also some of these practitioners are probably registered but inactive and that sort of stuff, so spend a lot more time cleaning it up and getting to a more finely accurate number. Now the one-month version is a little bit more complicated because maybe some of these practitioners work in multiple clinics and they probably have to pay a license for every clinic that they work. So one person can have multiple licenses, and that changes a little bit your market penetration formula and all of that. So in one month, you can probably scrape and bring external data. You can get other data sources like their websites, their LinkedIn, look at big clinics and get the list of all practitioners and look for duplications, and come up with a factor of how different your number is from the real potential one. That's one month worth of work. Now, the interesting thing is that a lot of technical teams, and I would say a lot of risk teams in that sense also, operate in the mindset that when you put a problem in front of them, what they tend to do is to think of the one-month version. They think of all the edge cases and all the complexities of the problem and everything. Whereas what you can do is really ask yourself, how can I create enough value? How can I create progress or create some value in one hour of work? And I think that people usually start with this version, start working on something, realize that it's even more complicated than they'd thought, and they stop in the middle and the project fails. It's not a new concept. You call it MVP. I think it's the same idea, or just agile. But the same project has multiple versions, and start from the small and then build it from there, rather than starting from the biggest one and fail.
Chen Zamir
Chen Zamir
67:31
Yeah. Or agile. Just agile. Be modest. Just be modest. Go back to basics and be modest. Awesome, Shachar. Thanks a lot for all of these valuable insights. For our listeners that wonder where they can hear more from you, or if they want to work with you, where should they go?
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
68:07
Yeah. So first of all, I'm on LinkedIn. I post on LinkedIn every day. Check me out. If what I say is merely interesting to you, feel free to follow and connect. I have my YouTube channel where I talk mostly about data career growth, but also some of these concepts that can also be applicable for a lot of risk professionals. And most importantly, if you're struggling with your risk team or with your data team, or something about this connection of risk and data doesn't work for you, I'm always happy to explore and see if there are ways I can help. This is what I do today as an advisor. I help companies get value from the data, and that includes risk teams. So of course, any problem you have, feel free to reach out and let's have a quick chat.
Chen Zamir
Chen Zamir
68:54
Shachar, thank you very much. It was a pleasure.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
68:58
Pleasure is all mine.
Chen Zamir
Chen Zamir
69:00
And for the rest of you, I'll see you next Saturday.
Black and white close-up of a smiling man with glasses and messy hair, looking to the right.
Shachar Meir
69:05
Thank you.