← Back to all projects

PROJECT 13 / Financial Crime Analytics · Graph ML

Financial Crime Risk Intelligence

A scalable blockchain anti-money-laundering decision-support project that turns Elliptic2 transaction subgraphs into a validated investigator priority queue, explicit workload trade-offs, calibration controls, case-specific statistical evidence, and a matched graph-native benchmark.

INGEST · VALIDATE · RANK · EXPLAIN
DATASETElliptic2 blockchain transaction graph
SIZE49.3M background nodes · 196.2M edges · 121,810 labeled components
SOURCEElliptic2 public dataset ↗

MEASURED EVIDENCE

What the validated analysis shows.

These metrics come from the locally executed Elliptic2 pipeline and validation runs. Published paper benchmarks are kept separate from project results.

0.5279Mean PR-AUC

The preferred node-enriched random forest remained stable across five stratified 80/20 splits with PR-AUC SD of 0.0081.

41.53×Top-0.5% lift

The preferred model captured about 115 suspicious components on average in roughly 122 reviews at the tightest tested review budget.

196.2MBackground edges scanned

All 367,137 labeled edges matched exactly once to the large background-edge table using the full source, target, and transaction key.

0.2498GraphSAGE PR-AUC

A directed GraphSAGE model on the exact seed-42 RF test components materially underperformed the node-enriched RF at 0.5306 PR-AUC.

VALIDATION

Repeated-split stability, shuffled-label sanity, schema-leakage checks, feature-dominance review, edge-match integrity, held-out calibration, and a matched graph-native comparison were evaluated before the preferred model was finalized. Shuffled-label RF PR-AUC collapsed to 0.0210 against 0.0227 prevalence.

PLAIN-LANGUAGE INTERPRETATION

The model is designed to order connected transaction patterns for human review when investigation capacity is limited. The raw score is treated as a ranking signal rather than proof or a literal probability of criminal activity.

01

Why this project matters

Financial-crime monitoring is not only a classification problem. Real review teams operate under severe workload constraints, suspicious cases are rare, and network behavior can be distributed across connected transactions. The useful question is therefore not simply whether a model can separate classes, but whether it can put the right connected patterns near the top of a finite investigator queue.

02

What was developed

The project builds an out-of-core DuckDB and Parquet feature pipeline across tens of millions of nodes and hundreds of millions of edges, creates structural and anonymized node-feature aggregates, compares logistic regression and random forest baselines, validates repeated review-budget performance, tests a 380-feature edge enrichment, compares sigmoid and isotonic calibration, produces a capacity-ranked queue with case-specific statistical evidence cues, and benchmarks a directed GraphSAGE classifier on the exact same seed-42 held-out components as the RF.

03

What it means

The strongest operational result came from the node-enriched random forest rather than the largest feature set or the graph-native model. Node features lifted mean PR-AUC from near-random structural performance to 0.5279; edge enrichment reduced mean RF PR-AUC to 0.5022; and directed GraphSAGE reached 0.2498 versus 0.5306 for RF on the matched seed-42 test set. The project therefore treats both the edge and GraphSAGE experiments as evidence for parsimony: more data, features, or model complexity do not automatically create more investigator value.

04 / INVESTIGATOR WORKFLOW

From score to review evidence.

The final queue uses capacity-based tiers such as top 0.5%, 1%, 2%, 5%, and 10% instead of arbitrary probability-like thresholds. Each queued component can be paired with three statistical evidence cues describing unusual values among globally important anonymized features, including percentile, standardized deviation, and model importance. The evidence is designed to support review, not to invent causal explanations.

05 / CALIBRATION & GOVERNANCE

Ranking and probability are separated.

A held-out 60/20/20 experiment showed that sigmoid calibration reduced expected calibration error from 0.00881 to 0.00272 while preserving ranking and constrained-review results. Isotonic calibration achieved better probability metrics but reduced PR-AUC and top-budget capture, so the raw random-forest score remains the operational ranking signal and sigmoid is optional for research probability estimates.

06 / LIMITATION

What the project does not claim.

Elliptic2 features are anonymized, so the project does not assign unsupported business meanings to feature numbers. Scores and evidence cues do not identify real people, establish criminal conduct, make legal determinations, or automate regulatory reporting. The completed GraphSAGE benchmark uses the compact labeled-subgraph universe and is not a reproduction of published GLASS; a true full-background-graph GLASS reproduction remains an optional research extension.

TECHNICAL TOOLKIT

PythonDuckDBParquetPandasscikit-learnPyTorchPyTorch GeometricSeabornPlotly
NEXT PROJECTWastewater Infrastructure Analytics→