AML Decision Readiness

Your AML team passed the training. But how will they make the call?

A 14-session decision readiness baseline that shows how your analysts investigate, reason and escalate when the evidence is incomplete. Team-level patterns. No customer data. No verdict on any individual.

RCM ThinkLabs (rcmlabs.io) is a game-theory engine for measuring AML decision readiness. One decision session per business day, each up to fifteen minutes, records how analysts investigate ambiguous activity, test assumptions and escalate. The result is behavioral evidence that completion metrics cannot provide.

Built for second-line teams at banks, payment firms and fintechs. Output written to sit inside your compliance function’s annual plan.

The gap

Completion is documented. Judgment is not.

Every remediation plan says the same thing: enhanced training, sufficient staff with appropriate skills. Your institution can prove completion for every analyst. It cannot show a regulator how any of them makes the call when the signals are incomplete and the queue is full. That is the gap the baseline reads.

What conventional programs measure

Training completion. Policy recall. Quiz performance. Attendance. Certification.

What operational risk depends on

Signal interpretation under incomplete evidence. Evidence seeking before a decision. Assumption control and hypothesis updating. Escalation timing and thresholds. Consistency under volume pressure.

The founding cohort

Three founding seats. Then the benchmark closes.

The engine has run one full deployment, with a mission-critical engineering team at a U.S. defense contractor. It has not yet run inside an AML function, and we are not going to ask you to pay full price to be first. The first three institutions run on founding terms.

What a founding institution gets

The 14-session baseline for 10 to 25 analysts at founding pricing, confirmed on the call and honored on the continuing program through the end of 2026. The scoring rubric before day one, the full method note with the map, and a say in what the AML benchmark measures: the founding cohorts define the reference read every later institution is compared against.

What a founding institution gives

Permission to publish anonymized, aggregate results, with no institution named unless it chooses to be. And one 30 minute reference conversation with a peer institution, if the read was useful. If it was not, tell us, and we publish that too.

Founding terms end with 2026. Claim a founding seat

How it works

Fourteen sessions. Fifteen minutes each. No customer data.

Step 1

Select the team

10 to 25 analysts from one AML workflow: monitoring, triage, CDD and EDD, screening, or the SAR/STR desk. No integration, no customer data, no personal data beyond a work email.

Step 2

Run the sessions

One scenario per business day, up to fifteen minutes, for fourteen business days. Incomplete signals, competing explanations, real trade-offs; every decision is scored as it is made.

Step 3

Receive the Executive Decision Risk Map

Team-level patterns across six behaviors, delivered with the method note and scoring rubric, written to be attached to an internal report. Patterns, not people.

Sample map · illustrative

What the output looks like.

Every baseline reads the same six dials. Findings from AML cohorts will replace this sample as founding institutions complete their baselines.

Dial 01

Escalation timing

Whether the team shares one escalation threshold or several, and how far apart they sit.

Dial 02

Evidence seeking

How many analysts check a second source before deciding, and when they stop looking.

Dial 03

Assumption testing

How often a working theory is challenged before it becomes the conclusion.

Dial 04

Contradiction response

What the team does when new evidence contradicts the first read: update, hold, or split.

Dial 05

Red flag sensitivity

Whether recognized typologies register consistently across the team.

Dial 06

Decision consistency

Whether the same analyst makes the same call under volume pressure as under none.

No fabricated findings appear on this page. When a founding institution’s anonymized map is published here, it will be labeled with cohort size, workflow and month.

Where it lives

Evidence that you looked at competence, not only at completion.

The baseline is not a training record and not a monitoring control. It is behavioral evidence about the second line, produced from scenarios rather than customer files, and it lives in three places.

Training needs analysis

The team pattern summary

Six behaviors, one page, attached to the annual training needs analysis. Read by the Head of Financial Crime and the MLRO.

Governance and resourcing

The map plus the method note

Evidence of adequacy and competence when governance is reviewed. Read by the audit committee, internal audit, and examiners on request.

Change baseline

Before and after, same team

A read on either side of a new monitoring system, an AI agent in the queue, or a hire wave. Read by the CRO and the program owner.

Use cases

Built for the decisions that carry the risk.

Use case 01

SAR/STR escalation

When the evidence is suggestive but not conclusive, who files, who waits, and what separates them.

Use case 02

PEP and sanctions pressure

Holding a screening decision when the relationship manager is on the phone.

Use case 03

Monitoring over-reliance

Whether analysts investigate the alert or restate it.

Use case 04

CDD and EDD narratives

Whether the narrative tests the customer’s story or repeats it.

Use case 05

Alert triage under volume

What changes in the call when the queue doubles.

Use case 06

Emerging typologies

How the team reasons about a pattern nobody has trained them on yet.

Engine proof · defense sector deployment · by day 39

The engine, proven where the stakes were highest.

A 39-day deployment with a mission-critical engineering team at a U.S. defense contractor, roughly 40 people. Not AML figures: they show the engine reads decision behavior, and that people come back unassigned. AML cohort one publishes in the fifth tile.

84%of regulars improved by day 39.
15,000+decisions scored across 1,088 sessions.
70%practicing daily, nobody assigned.
+57%in the habit of surfacing contradictions.
AML cohort oneresults published Q4 2026.
Method and limits

What it reads, and what it does not.

Scenarios are authored against published AML typologies and reviewed before a cohort starts. Six behaviors, scored from multiple decisions per session, produce roughly 190 scored decisions per analyst. The rubric is shared before day one, the full method note ships with the map, and no individual scores leave the platform.

The baseline does not test typology knowledge, does not score decisions against your case files, and does not read live alerts. A team can be consistent and still be wrong: the baseline shows the consistency, your QA shows the correctness. It does not determine whether a SAR or STR must legally be filed, does not judge individual competence or fitness for employment, and does not replace mandatory training, monitoring systems or QA.

The engine draws on advanced game theory, developed through research at MIT with Prof. Muhamet Yildiz, and on behavioral science, including the work of Karl Kapp. Read the deployment evidence

Request the method note
Common questions

What compliance leaders need to know.

01What is AML decision readiness?

AML decision readiness is how reliably a financial crime team investigates, reasons and escalates when the correct answer is not obvious: when evidence is incomplete, signals conflict, throughput pressure is real and a pattern does not match a familiar typology. It is different from knowledge. A team can score perfectly on mandatory training and still clear alerts passively, escalate late or carry unverified assumptions into a decision. Readiness is observed behavior under realistic pressure, measured against the team’s own baseline, and it is the layer conventional completion metrics do not reach.

02How do you measure how AML analysts actually make decisions?

By watching decisions develop rather than grading final answers. RCM ThinkLabs places analysts inside fictional cases where evidence is incomplete and incentives conflict, then records the path each person takes: what they check before acting, which assumptions they test, how they respond when a fact contradicts their working theory, and when they choose to escalate. Because every choice is a move in a modelled game, the record is behavioral evidence rather than self-report, and it is read as team-level patterns, with each person measured against their own starting point.

03Why is the engine domain agnostic instead of built only for AML?

Because it separates judgment from knowledge. Score judgment with typology quiz content and every miss is ambiguous: a knowledge gap and a judgment gap look identical, so you cannot tell which one to fix. The engine reads how a person decides, independent of what they know, and your training records already certify the knowledge. It also means the instrument arrives proven rather than new. The same measurement ran in production with a mission-critical engineering team at a U.S. defense contractor, and its validity carried over when the content changed to AML, because the behaviors it reads, evidence-seeking, surfacing contradictions, escalation timing, are the same in any domain. And it is what makes a benchmark possible: one taxonomy of 3 vectors, 16 capabilities and 64 microskills reads every cohort on a common scale. A test built only for AML could never tell you how your team compares, because nothing would sit on the other side of the comparison.

04Is this a replacement for mandatory AML training?

No. Mandatory training, institutional policy, transaction-monitoring systems, case management, quality assurance and regulatory reporting all remain in place and are not touched. RCM ThinkLabs adds a behavioral evidence layer on top of them: it shows how consistently people apply what the training delivered once the correct response is debatable. The two answer different questions. Completion records show that information was delivered and recalled. The decision record shows what analysts do with it under ambiguity, which is where operational exposure actually lives.

05Does RCM ThinkLabs decide whether a SAR or STR should be filed?

No. RCM ThinkLabs does not determine whether a suspicious activity report or suspicious transaction report must legally be filed, and it makes no legal or regulatory determinations of any kind. It provides behavioral evidence about how a team investigates and escalates inside fictional scenarios, as an input to institutional oversight and management judgment. Filing decisions, regulatory interpretation and legal review remain entirely with the institution’s own compliance professionals, counsel and governance processes.

06What is the AML Decision Readiness Baseline?

A focused diagnostic engagement. The institution selects 10 to 25 analysts, investigators or reviewers from one AML workflow. Each participant completes one scenario-based decision session per business day for 14 business days, with each session taking up to fifteen minutes. The scenarios are fictional, so no live customer or transaction data is required. RCM ThinkLabs records how decisions develop during the sessions, and compliance leadership receives an Executive Decision Risk Map: aggregated team patterns, inconsistencies, blind spots and areas worth targeted follow-up. It is a baseline, not an audit or certification.

07What does the Executive Decision Risk Map show?

An aggregated view of how the team decides across six behaviors: escalation timing and thresholds, evidence-seeking tendencies, assumption testing, responses to contradictory information, red flag sensitivity and consistency under volume pressure. It highlights patterns and distributions rather than ranking individuals, and it names the blind spots that deserve further investigation or reinforcement. It is delivered with a method note and the scoring rubric, written to be attached to an internal report, with access and visibility aligned to the institution’s governance requirements.

08Does the baseline require integration with our monitoring systems?

No. The baseline runs on fictional scenarios and requires no integration with transaction-monitoring platforms, case-management systems or customer records, and no live customer or transaction data. No personal data beyond a work email is needed. Participants need a browser and up to fifteen minutes on each business day. That separation is deliberate: fictional cases reduce reliance on memorized answers and keep every live decision unaffected, so the baseline can start quickly and sit safely outside production systems.

09How long does each session take?

Up to fifteen minutes per session, one session per business day, for fourteen business days. The sessions are short on purpose. Decision behavior is built and revealed through frequent repetition under realistic pressure, not through occasional long exercises, so the format is a brief daily decision rather than a workshop. Each of the six behaviors is scored from multiple decisions per session, so the baseline produces roughly 190 scored decisions per analyst without pulling anyone away from their queue for meaningful stretches of the day.

10Can scenarios reflect our institution’s risk environment?

Yes. Scenarios are authored against published AML typologies and reviewed for plausibility before a cohort starts, and the decision tensions can be weighted toward the risks that matter most in your environment: escalation under volume pressure, sanctions exposure, reliance on monitoring output, or due diligence narratives that do not add up. The instrument underneath does not change, which is what keeps results comparable over time. During scoping, RCM ThinkLabs confirms the relevant workflow, participant roles, priority decision risks, governance expectations and reporting requirements.

11Does RCM ThinkLabs judge whether an individual analyst is competent?

No. RCM ThinkLabs does not determine individual competence, fitness, promotion readiness or employment status, and the findings should never be used as a standalone verdict on a person. No individual scores leave the platform. The reporting emphasizes aggregated team patterns and distributions, each person is measured against their own baseline rather than ranked against colleagues, and the purpose is developmental: to show leadership where reasoning is inconsistent, where escalation drifts and where reinforcement would pay off.

12What happens after the baseline?

The baseline ends with an executive briefing on the Executive Decision Risk Map: what the team’s decision patterns show, where the priority blind spots are and which areas deserve targeted follow-up. From there, institutions that continue move into RCM ThinkLabs’ standard deployment model: an Initial Operating Capability cycle to establish a full measured baseline across the cohort, then a Full Operating Capability rhythm of two fixed days per week, so decision readiness compounds and leadership keeps a continuous read instead of a one-time snapshot.

Request

Request the baseline.

Twenty minutes: play one AML scenario yourself and tell us where the scoring is wrong. Then we agree which team runs the baseline and when. The first three institutions run on founding terms, confirmed on the call.

Or send the details below and we will come back within 48 hours. Ask for the method note in the last field and it comes with the reply.

No participant, customer or transaction data is required to submit this form.

Prefer LinkedIn? Follow RCM ThinkLabs →

Fourteen sessions from now, you could know how your team really decides.