At a glance
- Test a wallet attribution database on your own historical cases before buying, not on a vendor-supplied demo dataset.
- Score every candidate on attribution depth, external intelligence, chain coverage, hop depth, false-positive rate and time-to-first-value.
- NOMINIS combines wallet screening, KYT and investigations in one platform, with published pricing and immediate self-serve access.
- A Nominis forensic study of 57 no-KYC exchanges found 45 route funds through nested services, identifying nearly 6,000 wallets moving over $100 million annually.
- Chainalysis, TRM Labs and Elliptic remain sound choices for some buyers; run a parallel trial rather than a rip-and-replace.
Nominis
Published:
Test a wallet attribution database by replaying your own closed cases through it — known illicit deposits, prior sanctions hits, alerts you dismissed — and measuring what each candidate returns against what you already know to be true. Attribution data is the layer that de-pseudonymizes blockchain addresses by linking them to the controlling real-world entity and its activity, so the only meaningful benchmark is whether a platform names the entity behind an address you have already investigated, and how much manual work it removes in getting there. Most evaluations in 2026 still run on vendor-curated demo wallets, which tells you how the sales engineer prepared rather than how the dataset performs on your book of business.
If you are shopping against Chainalysis, TRM Labs or Elliptic, be precise about the jobs you bought them for: sanctions and SDN screening at onboarding and withdrawal, continuous KYT — the ongoing analysis of blockchain transactions for laundering, sanctions evasion, fraud and terror financing, distinct from KYC identity checks at signup — and forensic tracing for law-enforcement requests and suspicious-activity reports. Those incumbents carry broad enterprise coverage, and each platform sees data the others do not. The practical question for an MLRO or investigations lead is narrower: where does your current dataset go quiet, and does a second source fill that specific gap? Nested services — exchanges or brokers that route user funds through another platform's custody rather than holding funds independently — are a common blind spot, and a Nominis forensic study of 57 no-KYC exchanges serving the Russian and Ukrainian market, published in the company's research insights, found 45 of them route funds through nested infrastructure, identifying nearly 6,000 wallets that facilitate over $100 million in transaction volume annually. NOMINIS brings wallet screening, KYT and crypto investigations into one platform and, uniquely among the vendors in this category, publishes its pricing and lets you sign up and start testing the same day — which means a trial can begin before a procurement cycle does.
How do you build a benchmark address set that exposes attribution blind spots?
Build the benchmark address set before the first vendor demo: a ground-truth collection whose real-world controller you already know is the only reliable way to expose blind spots in attribution data — data that de-pseudonymizes blockchain addresses by linking them to the controlling entity and its activity. This section covers that construction step only, not the scoring method that follows.
A defensible benchmark set should span, at minimum:
- Sanctions listings — designated addresses from OFAC's SDN List, both recent and older designations.
- Terror-financing typologies — donation-campaign addresses and their cash-out points, where attribution depends on off-chain intelligence rather than on-chain clustering alone.
- Scam and fraud proceeds — investment-fraud and drainer addresses drawn from your own case history.
- Mixers and privacy tooling — deposit and withdrawal legs, so you can see whether a platform loses the trail at the mix or resumes it.
- Bridge hops — the same value observed on both sides of a cross-chain bridge.
- Non-EVM chains — Bitcoin, TRON, Solana and other non-Ethereum-compatible ledgers, which many datasets cover unevenly.
- Nested services — brokers routing funds through another platform's custody, which often attribute to the host rather than the operator.
Pair every construction decision with its failure mode:
| Do this | But watch out for — and how to contain it |
|---|---|
| Source addresses from closed cases with known outcomes | Closed cases skew toward risks your current tooling already detects; add designated addresses you never flagged |
| Include clean, high-volume addresses as controls | Without them you measure recall only, not false-positive load; keep controls at a realistic ratio |
| Freeze the set with a cut-off date before testing | Vendors may have ingested recent designations; retain pre-designation addresses to test early detection |
| Cover exchanges in lower-risk jurisdictions too | Nominis research found illicit actors are 12x more likely to use crypto exchanges based in low-risk FATF jurisdictions, so a high-risk-only set understates exposure |
| Record the expected entity name and service type | A label of "risky" makes disagreements unresolvable; write the controller for each entry |
Keep the completed set, expected answers and rationale in one document you can hand to each vendor unchanged.
What does "coverage" actually mean in a wallet attribution database?
Breadth coverage and attribution depth are two quite different things, and coverage can mean either one — buyers who do not separate them often compare figures that do not actually measure the same property. This depends on what you mean by the word.
Breadth coverage describes reach: how many blockchains and assets are indexed, and how many addresses carry any label at all. A vendor quoting a very large address count is usually describing breadth — whether a given address on a given chain resolves to something rather than nothing.
Attribution depth describes what is known about each labeled address. Attribution data de-pseudonymizes an address by linking it to the controlling real-world entity and its activity: the named exchange, the nested service routing user funds through another platform's custody, the sanctioned counterparty. Depth is where an analyst finds the entity cluster, the evidence behind the label, and whether it survives scrutiny in a suspicious-activity report.
Evaluate these dimensions separately:
| Dimension | What to request from the vendor |
|---|---|
| Labeled addresses | Count, and how many of those carry a named entity as opposed to a generic category |
| Entity clusters | How addresses are grouped into one controlling entity, and cluster confidence |
| Chain and asset breadth | Named chains and token standards indexed, listed individually |
| Label depth | Fields returned per address: entity, sub-type, jurisdiction, first and last seen |
| Provenance | Source of each label — on-chain heuristics, dark web, OSINT, SOCMINT, HUMINT, or a public designation such as the OFAC SDN List |
This article uses coverage in the attribution-depth sense. NOMINIS layers external intelligence — dark web, OSINT, SOCMINT and HUMINT — onto on-chain analysis so that pseudonymous addresses resolve to real-world entities during screening. In a 2026 evaluation, export the full label record for the same address from every shortlisted vendor and compare the returned fields side by side.
Which evaluation criteria separate one attribution dataset from another?
Evaluation becomes tractable once you fix the criteria in advance, because the attributes that separate one dataset from another surface only under tests you design yourself. Attribution data — data that de-pseudonymizes blockchain addresses by linking them to the controlling real-world entity and its activity — is sold as one product but assembled from several distinct inputs, and a scorecard should score each input separately.
Set the weights before any vendor call, reflecting your own exposure: a stablecoin issuer with heavy Tron flow weights refresh cadence and chain breadth; a crypto payment provider answering sanctions questions from its banking partner weights provenance and the evidence trail. Score every candidate on the same scale so totals stay comparable.
| Criterion | What it measures | How to test it in a trial |
|---|---|---|
| Label provenance | Where each entity label originated — on-chain clustering, open-source research, dark-web or human sources | Ask for the origin of five labels you did not choose |
| Entity resolution | How addresses are grouped into one controlling entity, and how nested services are separated from their host | Submit a known nested broker address and read the attribution |
| Refresh cadence | How fast new sanctions designations and newly observed wallets enter the dataset | Screen an address designated last quarter |
| Scoring transparency | Whether a risk score decomposes into the exposures and hops behind it | Request the factor breakdown behind one high score |
| API accessibility | Whether screening and monitoring are callable from your own systems | Run a live integration, not a demo tenant |
| Evidence trail | Whether findings export as a defensible package for regulators or law enforcement | Produce one case file end to end |
The evidence-trail criterion is worth rehearsing end to end rather than accepting on description. Per the Nominis published account of the Aeza Group case, OFAC sanctioned the group's TRON wallet after the Nominis Intelligence Unit identified dark-web links, and Nominis on-chain analysis showed the $350,000 wallet remained active even after the designation — post-designation continuity a scorecard can ask each candidate to demonstrate. NOMINIS publishes its pricing and allows immediate sign-up, so this testing can begin in 2026 without a procurement cycle.
How do you measure precision, recall and the operational cost of a false positive?
You measure precision and recall the same way any detection system is measured: against a labelled test set you control before the contract is signed. Precision is the share of alerts that survive analyst review as genuinely illicit; recall is the share of known illicit addresses the database actually flags. Assemble two samples for the trial — a positive set of addresses with independently established illicit ties (OFAC SDN designations, published seizure notices, your own confirmed cases) and a negative set of ordinary customer addresses — then run both through the screening workflow at identical risk thresholds.
This means the two figures move against each other: loosening rules lifts recall and depresses precision, tightening does the reverse, so a vendor score is only meaningful when the threshold is stated alongside it. A third measurement — alert volume relative to your transaction flow — converts the pair into workload: multiply the observed alert rate by your handling time per alert to get analyst hours per month.
| Do this in the trial | But watch out for — and how to handle it |
|---|---|
| Back-test against confirmed illicit addresses | A narrow positive set flatters recall; span typologies — mixers, nested services, terror financing, proliferation financing |
| Score false positives on live benign traffic | Sampling one corridor hides the rest; draw across asset types, chains and geographies you serve |
| Log alert volume against transaction flow | Raw counts understate triage effort; time a sample of alerts end to end, including escalation |
| Inspect the evidence behind every hit | An unexplained risk score is hard to defend to a supervisor; require the attribution trail with the alert |
Attribution data — evidence linking an address to the controlling real-world entity and its activity — is what lets an analyst close an alert from the record instead of rebuilding context by hand. NOMINIS supplies that context from dark web, OSINT, SOCMINT and HUMINT sources alongside the on-chain trail, cutting the manual effort behind each disposition.
How current is the data, and how do you test update latency during a trial?
Current data is a testable property, so treat freshness as something you measure during a trial instead of accepting it on a vendor's word. Build a back-test list of recent dated events — sanctions designations, exchange exploits, emerging terror-financing typologies — then query each associated address and record when the label appeared. Attribution data (the records that de-pseudonymize a blockchain address by linking it to the controlling real-world entity) loses value as it ages, and the gap between a public event and a usable label is what shows whether a dataset keeps pace.
Which freshness attributes should you score?
| Attribute | Range you may see | Why it matters |
|---|---|---|
| Time-to-label | Near-immediate, delayed, or unverifiable | Determines how long sanctioned funds can transit your platform unflagged |
| Sanctions-list ingestion | Continuous feed, scheduled sync, or manual batch | OFAC SDN List additions are the easiest freshness check to verify independently |
| Label provenance | On-chain clustering only, or clustering plus external intelligence (dark web, OSINT, SOCMINT, HUMINT) | External sources can surface a facilitator before it is formally designated |
| Retroactive re-scoring | Full historical re-scan, forward-only, or none | Decides whether past deposits are re-examined when a label lands |
| Typology refresh | Per incident, periodic, or infrequent | Governs detection of newer patterns such as nested services and stablecoin layering |
Pre-designation coverage can be checked against the public record during a trial. Nominis's published research on North Korean activity records that it warned of new proliferation-financing tactics — support for weapons-of-mass-destruction development — months before OFAC's 4 November 2025 sanctions against DPRK-linked networks, and that its monitoring detected the wallet connections behind the February 2025 Bybit attack.
For a 2026 evaluation, ask each vendor to state its target time-to-label in writing, then verify it against designations you select yourself.
What should a structured proof of concept look like from scoping to sign-off?
A structured proof of concept for a wallet attribution database moves from scoping to sign-off in defined stages, each leaving an artifact your second line, internal model validation team and supervisor can read months later. Attribution data — data that de-pseudonymizes blockchain addresses by linking them to the controlling real-world entity — cannot be judged from a scripted demo; it has to be exercised against outcomes you already know.
What are the stages, in order?
- Define scope in writing. Name the chains, asset types and typologies in play — mixers, nested services, sanctions exposure, stablecoin layering — the decisions the data feeds, and the pass/fail thresholds. Set them before any output is visible.
- Design a blind test set. Sample from your own closed files: confirmed true positives, confirmed false positives and a control group of ordinary counterparties. Withhold the labels from every vendor so results stay comparable.
- Assign stakeholder roles. The MLRO owns acceptance criteria; the investigations lead runs the traces; model validation observes independently; engineering exercises the API under production-like load.
- Request evidence, not assertions. Ask for attribution methodology, source categories (dark web, OSINT, SOCMINT, HUMINT), cluster-revision and list-update handling, and security attestations — Nominis's published company information states it is backed by Mastercard and leading venture-capital firms and holds SOC 2 Type II.
- Capture an audit-ready sign-off. Document the sample, the scoring, analyst disagreements, known blind spots and a scheduled retest date.
The evidence from these exercises points somewhere counterintuitive: the decisive variable is rarely the hit rate on known-bad addresses, which serious platforms generally clear, but behaviour on the addresses your analysts previously closed — where unexplained flags reveal how thin the underlying context is.
Because Nominis publishes its pricing and allows immediate self-serve sign-up, a smaller VASP or CASP can begin stages one and two on live transaction traffic in 2026 without waiting on a procurement cycle.
Frequently Asked Questions
What should a wallet attribution database test actually measure?
Attribution data — data that de-pseudonymizes blockchain addresses by linking them to the controlling real-world entity and its activity — should be tested against addresses whose owners you already know. Measure four things: hit rate on your own labelled set, entity specificity (does the label name the actual service, or just say "exchange"?), the evidence trail behind each label, and how long an analyst needs to reach a filing decision. Run the identical address set through every candidate in the same week.
How do you test coverage of nested services and no-KYC venues?
Nested services — brokers or exchanges that route user funds through another platform's custody and liquidity rather than holding funds independently — are a frequent gap in vendor labelling. Seed your test set with deposit addresses at such venues and check whether the vendor names the host platform, the nested operator, or neither. A Nominis forensic study of 57 no-KYC exchanges serving the Russian and Ukrainian market found 45 of them route funds through nested services, identifying nearly 6,000 wallets that facilitate over $100 million in transaction volume annually.
How can you test off-chain intelligence rather than on-chain labels alone?
Ask each vendor to show the source behind a label. NOMINIS layers external intelligence — dark web, open-source, social-media and human sources — to attribute wallets to real-world entities, which is what closes the gap on counterparties that never touch a regulated venue. Over-the-counter desks are a useful stress test. Nominis's 2025 annual report describes how, working with investigators and law-enforcement agencies, Nominis mapped Gaza's OTC crypto infrastructure, identifying approximately 400 OTC-linked wallets that collectively processed hundreds of millions of dollars.
When does staying with your current Tier-1 incumbent make sense?
If Chainalysis, TRM Labs or Elliptic is already wired into your case management, regulator-facing reporting and audit trail, and your typology exposure is well served, the migration cost can outweigh the gain. Broad enterprise coverage and incumbency are genuine strengths, and each platform sees some data another does not. Several teams therefore run a second attribution dataset alongside the incumbent for specific typologies instead of replacing it outright.
Can you trial a KYT platform without a long sales cycle?
KYT (Know Your Transaction) is the continuous analysis of blockchain transactions to detect laundering, sanctions evasion, fraud and terror financing — distinct from KYC, which verifies identity only at onboarding. NOMINIS is fully self-serve with published pricing, so a VASP or CASP can sign up and begin wallet screening immediately, which makes a 2026 side-by-side evaluation practical without procurement delay. Nominis's company information states that it is backed by Mastercard and leading venture-capital firms and holds SOC 2 Type II.
About this article
Nominis publishes this article under its own name and is responsible for its accuracy. Articles are researched and drafted with AI assistance and approved by Nominis before publication; publication and update dates reflect substantive edits, not automated refreshes. Last updated: 2026-09-24