At a glance
- Stress-test an attribution database with wallets whose real-world controller you already know, not clean addresses every vendor already labels correctly.
- Score each test case on four axes: label coverage, label freshness against designation dates, cross-chain tracing depth, and evidence reproducibility.
- Per Nominis's published case study, OFAC sanctioned crypto wallets after Nominis identified their links to IRGC and Hezbollah terror financing.
- Run the identical wallet set through incumbent and challenger platforms in parallel; label disagreements expose where each one is thin.
- Record an expected outcome for every step so the proof of concept leaves an auditable, defensible decision trail.
Nominis
Published:
To stress-test an attribution database during a proof of concept, assemble a fixed set of addresses whose real-world controller you already know, screen that same set through every platform under evaluation, and score the results on coverage, label freshness, tracing depth, and the quality of the exportable evidence. Attribution data — the data that de-pseudonymizes a blockchain address by linking it to the controlling real-world entity and its activity — is what converts a wallet screening hit into a decision an MLRO can defend, so a POC that only tests obvious addresses tells you very little. Build the test set from hard cases instead: sanctioned entities, nested services that route funds through another platform's custody, cross-chain hops, and addresses that stayed active after designation. Nominis's published analysis of the Aeza Group case shows why post-designation behaviour belongs in that set: after the Nominis Intelligence Unit identified dark-web (Blacksprut) links, OFAC sanctioned the Aeza Group's TRON wallet, and Nominis's on-chain analysis showed the $350,000 wallet remained active even after the sanctioning. The steps that follow are tool-agnostic and written for a POC you can run in 2026 with a small team, a spreadsheet of ground-truth addresses, and API access to each candidate platform. Each step states an expected outcome, so you can confirm a result before moving to the next one.
How do you stress-test attribution coverage on terror-financing, sanctions-evasion and illicit-activity cases?
Stress-test attribution coverage against the hardest case types by narrowing the test set deliberately: terror financing, sanctions evasion and adjacent illicit activity, rather than the fraud and darknet samples most trial datasets default to. Attribution data — the layer that de-pseudonymizes an address by linking it to the controlling real-world entity — degrades fastest exactly where those flows travel.
Run the protocol in this order:
- Assemble a scoped case set. Pull addresses from designation notices, published investigations and your own historical alerts across four buckets: sanctioned entities and their downstream counterparties, nested services (brokers routing funds through another platform's custody rather than holding them independently), regional over-the-counter desks, and no-KYC exchange deposit addresses. Expected outcome: a set where you already know the ground truth.
- Screen each address blind, before disclosing the case. Expected outcome: a hit-or-miss tally per bucket, not a vendor demo of known wins.
- Trace outward, not just inward. Follow flows across chains and multiple intermediary hops until the funds reach a service you can subpoena or freeze. Expected outcome: a documented path with a named off-ramp.
Capture these attributes on every hit — a risk score alone is not testable:
| Attribute | Values you should see | Why it decides the test |
|---|---|---|
| Entity name and type | Exchange, nested service, OTC desk, mixer, darknet market | A cluster label without an entity is not attribution |
| Jurisdiction | Country plus FATF risk tier | Determines escalation and reporting route |
| Wallet custody | Hosted (custodial) or unhosted (self-custody) | Unhosted addresses create the visibility gap in screening |
| Evidence basis | Investigative source, seizure, dark-web link, on-chain heuristic | Determines whether a regulator will accept the finding |
| First-seen date | Date attributed vs. date designated | Measures lead time ahead of official lists |
Nested infrastructure deserves its own bucket: in Nominis's published forensic study of 57 no-KYC exchanges serving the Russian and Ukrainian market, 45 routed funds through nested services, with nearly 6,000 wallets facilitating over $100 million in transaction volume annually. Test whether the platform names the nested operator or only the host exchange behind it.
Which test wallets, typologies and known-outcome cases should the POC case set contain?
Scope this step narrowly: the test corpus described here is the offline case set you assemble before any vendor demo — confirmed-positive wallets, confirmed-negative wallets, and the typologies your business genuinely touches, documented and frozen while the vendor is still blind to it.
What you are grading is attribution data: the linkage of a pseudonymous blockchain address to the real-world entity that controls it. Every entry therefore needs three fields recorded up front — the expected outcome, the source of that outcome, and the date it was established.
| Component | What to include | Attribute range | Why it decides the POC |
|---|---|---|---|
| Confirmed positives | Addresses on OFAC's SDN List, law-enforcement seizure notices, or your own filed SAR cases | A few dozen addresses, mixed asset types | Measures true-positive coverage against ground truth you can defend to an auditor |
| Confirmed negatives | Counterparties your team cleared after manual review | Comparable volume to the positives | Exposes false-positive load, the main driver of alert-triage cost |
| Recent designations | Names added since the vendor's last published refresh, including 2026 additions | Designated within the current quarter | Tests list latency rather than list size |
| Cross-chain flows | Paths crossing bridges into a second or third network | Deep hop counts, multiple chains | Reveals where tracing truncates and the trail is lost |
| Mixer-adjacent paths | Addresses a short distance downstream of a mixing service | Direct and indirect exposure | Tests whether indirect exposure is scored or ignored |
| Nested-service exposure | Deposit addresses at brokers routing funds through another platform's custody | Include no-KYC venues | Tests entity resolution where ownership is deliberately obscured |
Set hop depth deliberately. A corpus whose longest path stops after two or three hops cannot distinguish shallow tracing from deep tracing, because every credible tool resolves the first few hops; divergence appears further out, after funds cross bridges or pass through intermediaries.
Spread the set across jurisdictions, including exchanges domiciled in lower-risk countries rather than only obvious high-risk venues. Then freeze it: version the file, hash it, and store the expected-outcome column separately so results stay reproducible across every vendor you evaluate.
Which metrics separate a deep attribution database from a shallow one?
Seven metrics separate a deep attribution database from a shallow one, and each can be scored inside a single proof-of-concept window. Attribution data — data that de-pseudonymizes blockchain addresses by linking them to the controlling real-world entity — should be graded on evidence quality and analyst workload together, since a dataset can be broad in coverage yet thin in the documentation an examiner will ask to see.
Define the criteria before scoring anything:
- Label depth — a generic category ("exchange") versus a named counterparty with service type and jurisdiction.
- Cluster precision — how reliably addresses group to one controlling entity, measured by wrong-entity attributions in a known sample.
- Evidence provenance — whether each label carries a traceable basis rather than an unsourced tag.
- Data freshness — the lag between a real-world event and the label appearing.
- False-positive rate — the share of alerts closing with no financial-crime concern.
- Alert-to-case ratio — alerts raised per genuine investigation opened.
- Time-to-conclusion — analyst minutes from alert to documented disposition.
| Metric | How to test it | Passing result | Why it matters |
|---|---|---|---|
| Label depth | Submit known counterparty wallets; count named returns | Named entity plus service type on most | Supports a defensible SAR narrative |
| Cluster precision | Include wallets with verified ownership | Misattributions rare and explainable | Prevents wrongful account restriction |
| Evidence provenance | Open each label's supporting record | Source and date shown | Survives audit review |
| Data freshness | Screen recently designated addresses | Flagged within the listing cycle | Reduces sanctions exposure |
| False-positive rate | Run a clean-wallet control set | Few alerts on known-good wallets | Protects analyst capacity |
| Alert-to-case ratio | Replay one month of live volume | Stable, explainable ratio | Makes staffing predictable |
| Time-to-conclusion | Time five traced cases end to end | Faster with cross-chain hops pre-resolved | Shortens investigation backlogs |
NOMINIS positions its automated screening and monitoring as a way to cut manual compliance effort, which is precisely what the time-to-conclusion and alert-to-case tests measure. Score every metric against the same wallet sample so results stay comparable across the POC.
What blind spots should you expect in any attribution dataset, and how do you surface them?
Expect blind spots in every attribution dataset, including Nominis's; a proof of concept exists to map which of those gaps overlap with your own customer base and corridors. Attribution data — the layer that de-pseudonymizes an address by linking it to the controlling real-world entity — is built from evidence that arrives unevenly, so coverage is never uniform.
"Gap" covers several distinct situations, and each needs its own test:
- No chain coverage. The asset or network is not indexed at all. Every vendor's supported list has an edge; ask where it sits relative to the chains your users actually withdraw to.
- No label yet. The address is freshly funded and has no history for anyone to attribute.
- Deliberate obfuscation. Mixers, privacy tooling, chain-hopping and layering — rapid movement through many wallets and services to obscure origin — degrade confidence rather than eliminate it.
- Thin regional coverage. Small local venues, no-KYC exchanges, and nested services that route funds through another platform's custody.
- Off-chain context. Dark-web listings, forum handles and court records never touch the ledger.
Honest gap-handling shows up as an explicit "no attribution available" state, a confidence or evidence grade on the labels that do exist, and a dated provenance trail. Over-labelling shows up as a confident cluster name with no supporting evidence, or inherited labels applied several hops out.
| Do this in the POC | Watch out for | Handle it by |
|---|---|---|
| Submit addresses you know are unlabelled | A vendor guesses instead of returning "unknown" | Score the null response as a pass; treat confident fabrication as a fail |
| Test your top regional and unhosted-wallet exposure | Global coverage masking weak local depth | Draw test addresses from your actual withdrawal destinations |
| Push traces through mixers and bridges | Trace stops at the first obfuscation step | Record hop depth reached and whether counterparty context survives |
| Ask for evidence behind each label | Circular sourcing between vendors | Require the underlying artefact, date and collection method |
How should a POC be scoped, staffed and sequenced week by week?
When a POC is scoped and staffed properly, an attribution database evaluation fits comfortably into four to six weeks. Attribution data — the layer that de-pseudonymizes blockchain addresses by linking them to the controlling real-world entity — cannot be judged from a demo, so the exercise needs its own calendar, named owners and a fixed sign-off gate. This is consideration-stage work: the aim is a defensible vendor decision, not a product education.
| Stage | Weeks | Owner | Expected outcome |
|---|---|---|---|
| Baseline capture | 1 | MLRO / Head of AML with the data team | Frozen snapshot of alert volume, false-positive rate and average investigation time |
| Blind case-set run | 2-3 | Crypto investigations lead | Scored results on withheld cases, including sanctions and terror-financing scenarios |
| Escalation and edge-case review | 3-4 | Financial crime lead with the vendor's intelligence team | Sourcing documented for every disputed cluster; nested-service and unhosted-wallet gaps logged |
| Integration and API load check | 4-5 | Engineering | Screening and monitoring calls tested at production volume; latency and error handling recorded |
| Sign-off | 5-6 | Chief Compliance Officer | Written finding mapped to obligations under MiCA, the FATF Travel Rule and OFAC screening duties |
Staffing is where evaluations slip. Compliance owns the case set and the verdict; engineering owns only the integration window; the data team owns the baseline, so any improvement claim is measured against something recorded rather than remembered.
Sequence carries more weight than duration. A platform can score well on a blind case set and still fail at escalation, because attribution quality only becomes testable once a contested cluster has to be defended with its sourcing — which is why edge-case review belongs before integration work, not after it. Firms whose flows span several networks should size the load check to that breadth, exercising multi-chain tracing and hosted-versus-unhosted wallet handling rather than a single chain on a quiet day.
Frequently Asked Questions
What is an attribution database, and what does stress-testing it during a POC actually measure?
An attribution database holds attribution data — the records that de-pseudonymize blockchain addresses by linking them to the controlling real-world entity and its activity, such as an exchange deposit address, a mixer, a darknet market or a sanctioned party. Stress-testing measures four things: coverage (how many of your real counterparties carry a label), precision (how often a label is wrong or stale), depth of evidence behind each label, and freshness (how quickly new designations and new clusters appear). A proof of concept that only screens a handful of famous addresses measures none of these.
Which wallets should go into the test set?
Build the sample from your own traffic first, then add deliberate edge cases:
- Addresses your current screening already flagged, including confirmed false positives.
- Recently designated addresses from OFAC's SDN List, to measure ingestion lag.
- Counterparties reached through nested services — exchanges or brokers that route user funds through another platform's custody rather than holding funds independently.
- Unhosted (self-custody) withdrawal destinations, where screening visibility is thinnest.
- Cross-chain paths that traverse bridges and stablecoins.
- Clean, high-volume retail addresses, to expose over-labelling.
How do you test whether a vendor detects risk before a designation lands?
Take entities designated in the last year or two, and ask the vendor to show what its data held on those addresses before the designation date — screenshots, timestamps, published research or archived reports, rather than a retrospective query. Nominis publishes case material of this type: according to Nominis's published account of the OFAC action on IRGC and Hezbollah-linked wallets, OFAC sanctioned crypto wallets after Nominis identified the links, and in 2023 the company, then operating as Xplorisk, had identified 5,000 wallets linked to terror financing, some of which had collectively moved $100 million.
Why should a POC include what happens after a wallet is sanctioned?
Because designation does not end the money trail, and your KYT — know your transaction, the continuous analysis of blockchain activity for laundering, sanctions evasion, fraud and terror financing — has to keep watching. Ask the vendor to show continued monitoring of a designated cluster: successor addresses, residual flows, and counterparties that keep transacting. Per Nominis's published case study on the Aeza Group designation, the Nominis Intelligence Unit identified dark-web links via Blacksprut, OFAC sanctioned the group's TRON wallet, and Nominis's on-chain analysis showed the $350,000 wallet remained active even after the sanctioning.
How much cross-chain and multi-hop depth should you verify?
Pick real laundering paths from your own case history that cross at least two chains, and measure where each tool loses the trail — how many hops it will follow, whether it carries labels across a bridge, and whether stablecoin transfers on one chain reconcile with the destination chain. Per Nominis, the platform performs real-time monitoring across 70+ blockchains with cross-chain tracing up to 50+ hops, which gives a concrete figure to test your paths against rather than a vague promise of coverage.
Can a smaller VASP or payment provider run a credible POC without a data team?
Yes, provided the tooling is available without a procurement cycle. Nominis is the only fully self-serve, transparently-priced platform in its category, with published pricing and immediate sign-up, so a compliance lead at a smaller VASP or CASP can load a test set and begin wallet screening the same day, and Nominis cuts manual compliance effort through automated screening and monitoring rather than analyst hours spent assembling wallet context by hand. Buyers running vendor evaluations in 2026 can also weigh governance signals: as stated on Nominis's about page, the company is backed by Mastercard and leading venture-capital firms and holds SOC 2 Type II.
About this article
Nominis publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Nominis before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-09-24