SardineCon SF/2026

Learn More

¿Qué es Threshold tuning?

SUBSCRIBE

Threshold tuning is the disciplined adjustment of scenario thresholds to balance detection coverage against false-positive volume, backed by testing and a documented reason for each change. Thresholds decay over time, so examiners expect tuning to be an ongoing, evidenced activity, not a one-time event.

What is threshold tuning, in plain English?

Threshold tuning is the practice of deliberately adjusting the values inside monitoring scenarios so they keep catching real risk without drowning analysts in noise. It is how a program keeps its thresholds calibrated as the world changes, using data rather than guesswork to decide where each line should sit.

The reason it is necessary is that thresholds decay. A setting that worked when it was chosen slowly drifts out of alignment as customer behavior shifts, products change, and new typologies appear. Left alone, a threshold either starts missing genuine risk or starts generating floods of false positives.

Good tuning is evidence based. The core technique is above and below the line testing: checking that the alerts a threshold produces are productive, and sampling activity just below the threshold to confirm you are not missing anything you should catch. Every change is tied to data and kept with a documented rationale.

How a tuning cycle works

Tuning is a repeatable loop, not a one off tweak. A typical cycle runs like this:

  1. Measure — Look at the numbers. Review alert volumes, false-positive rates, and productive outcomes for each scenario and threshold.
  2. Test — Above and below the line. Confirm current alerts are useful, and sample activity just below the threshold to check for missed risk.
  3. Adjust — Change with a reason. Move the threshold based on the evidence, tying every change to data, not to analyst workload alone.
  4. Document — Keep the before and after. Record the rationale, testing, and impact so the change is defensible to model validation and examiners.

What it looks like in practice

In practice

A rapid movement scenario is generating thousands of alerts a month, and analysts clear most as false positives. The team is tempted to simply raise the threshold to cut the volume, but they run the analysis properly first.

Above the line testing confirms the current alerts are almost all unproductive at the low end. Below the line testing, sampling activity just under a proposed higher threshold, shows no genuine cases would be lost. They raise the threshold, document the testing, alert volume, and the reasoning, and keep the before and after. The change cuts noise and, crucially, it is defensible because it was driven by evidence, not by the size of the queue.

Why threshold tuning matters to operators

Examiners expect to see tuning as a continuous, evidenced discipline. A program that set its thresholds once and never revisited them is a finding, because those thresholds have almost certainly decayed into either blind spots or noise. Showing a regular, documented tuning process is part of proving your monitoring is actually calibrated to risk.

The classic criticism is tuning for the wrong reason. Loosening thresholds purely to shrink an alert backlog, driven by analyst workload rather than risk, is a recurring exam finding, especially when there is no documentation behind the change. The whole point is that every adjustment must be justified by data and preserved with its before and after analysis.

What to watch when tuning

  • Workload driven changes. Loosening thresholds to clear a backlog rather than because risk changed is the classic criticism.
  • Missing documentation. A change with no recorded rationale or testing cannot be defended, even if it was sensible.
  • One and done. Tuning treated as a single project rather than an ongoing cycle lets thresholds silently decay again.
  • No below the line testing. Raising a threshold without checking what sits beneath it can quietly create a blind spot.
  • Ignoring segments. A single threshold across very different customer types often misfits both; segment before you tune.

Quick questions

How often should tuning happen?

There is no fixed rule, but it should be a recurring, scheduled activity rather than a one-time project, with additional tuning whenever products, risk, or customer behavior change materially. Examiners look for evidence of an ongoing process.

What is above and below the line testing?

Above the line testing examines the alerts a threshold currently produces to see if they are productive. Below the line testing samples activity just under the threshold to confirm you are not missing genuine risk. Together they justify where the line sits.

Is it wrong to tune to reduce alerts?

Reducing noise is a legitimate goal, but the change must be justified by evidence that the removed alerts were unproductive and no real risk is lost. Loosening thresholds purely because the queue is long, with no testing, is the classic finding.

Who should tune thresholds?

Usually the financial crime or monitoring team, often with data analytics support and independent review by model validation. Segregating who tunes from who validates strengthens the defensibility of the changes.

What documentation is expected?

The rationale for each change, the testing that supported it, the data on alert volumes and outcomes, and the before and after impact. This record is what makes tuning defensible to examiners and validators.

Can machine learning replace tuning?

Models can help score and prioritize alerts, but they do not remove the need for governance. Any model that adjusts detection still needs validation, documentation, and testing, so the discipline of evidencing changes remains.

Go deeper

Qué saber junto con Threshold tuning