Deterministic vs Probabilistic Identity Resolution: What Changed When Cookies Didn’t Die
Most guides to this question were written for a future that never arrived. The standard framing — choose deterministic matching because third-party cookies are about to vanish — rested on a deprecation Google called off in April 2025, then followed in October 2025 by retiring the replacement technologies it had spent years building. Cookies did not die. They fragmented, which is a harder thing to plan around than a clean cutover would have been. This guide covers what each matching method actually does, where each one genuinely fails, and how the current browser reality changes which one you should want.
Quick answer: deterministic vs probabilistic identity resolution
Deterministic matching links records through stable person-level identifiers — an email address, a phone number, a postal address — so a match is either confirmed against a shared identifier or it is not made at all. Probabilistic matching infers that two records belong to the same person from circumstantial signals such as IP address, device type and user agent, trading precision for reach. Deterministic matching is more accurate and easier to defend to a regulator; probabilistic matching reaches the traffic deterministic matching cannot see. Most real systems use both, and the question worth asking a vendor is which one produced any given match.
What deterministic matching actually is
The IAB Tech Lab’s Identity Solutions Guidance defines deterministic identifiers as those relying on attributes that are “relatively permanent and associated with one person or household,” naming email address, phone number and home address as the working examples. That permanence is the whole point. A deterministic match is a claim you can audit after the fact: two records either shared a verified identifier or they did not, and you can go back and show which one.
This matters more than it sounds. When a campaign underperforms, deterministic matching lets you separate a targeting problem from a data problem, because the match itself is not in question. With probabilistic matching you are debugging two uncertainties at once, and you generally cannot tell them apart from the outside.
What probabilistic matching actually is
The same IAB Tech Lab guidance describes probabilistic methods as using “attributes from devices a consumer uses, the attributes of the way a consumer connects to the internet or specific applications” — it lists IP address, user agent, timestamps, and device details or settings. The document is direct about the trade: probabilistic methods “allow marketers to achieve greater scale, with potentially less precision.”
The structural weakness is in the inputs rather than the maths. IAB Tech Lab notes these methods rely on “relatively temporary attributes or attributes that change often.” An IP address changes when someone moves between home Wi-Fi, an office network and a phone connection. A user agent changes with a browser update. A household shares all of it. A probabilistic graph built on signals that shift weekly needs continuous rebuilding, and its accuracy on any given day is not something you can verify from the output.
The guidance also raises a consent problem that is easy to miss: these attributes are “often collected without user knowledge,” which means they require “proper permissions” rather than being free for the taking because they happen to be technically visible.
The ceiling on deterministic matching that vendors skip
Anyone selling deterministic matching, this firm included, should be straight about its limit. Deterministic matching needs someone to have identified themselves, and most of the web never asks. IAB Tech Lab puts a number on it: “only about 10%-20% of ‘open-web’ content sits behind an authentication wall.” Outside that slice there is no email address to match on, because nobody logged in.
That is the real reason probabilistic methods exist, and it is why a vendor claiming near-total deterministic coverage of open-web traffic is describing something other than deterministic matching. If the coverage number is very high and the method is said to be purely deterministic, one of those two claims is doing work it has not earned. Ask which.
What actually happened to third-party cookies
This is where most explainers on this topic are now out of date, and it changes the analysis rather than decorating it.
In April 2025 Google announced it would “maintain our current approach to offering users third-party cookie choice in Chrome, and will not be rolling out a new standalone prompt for third-party cookies.” Then in October 2025 it retired the Privacy Sandbox technologies built as replacements — among them Topics, Protected Audience, the Attribution Reporting API, Private Aggregation and IP Protection. The industry spent years preparing for a transition that was called off, and the standardised alternatives were withdrawn along with it.
Meanwhile the other browsers never moved. Apple’s WebKit documentation states that Intelligent Tracking Prevention “by default blocks all third-party cookies. There are no exceptions to this blocking.” Firefox takes a different route to a similar end: Mozilla’s documentation describes Total Cookie Protection, on by default, as giving “third-party cookies a separate cookie jar per site, preventing cross-site tracking” — the cookies exist but cannot follow anyone between domains. Chrome, per the same Mozilla reference, “doesn’t block third-party cookies by default.”
Why fragmentation is harder to plan around than deprecation
A deadline is something you can build toward. What the market got instead is a permanent split in which cookie-based matching works fully in one browser, is neutered in another, and is blocked outright in a third. Your effective coverage is now a function of your audience’s browser mix, which you do not control and which differs by device, demographic and channel. Two campaigns with identical targeting can return materially different match performance for no reason visible in the campaign itself.
This is what quietly strengthened the case for deterministic matching, though not for the reason the old “cookieless future” pitch gave. Deterministic matching was never improved by any of this. It simply does not depend on the thing that fragmented, and the standardised replacements that might have offered a neutral middle path no longer exist. A verified email address behaves the same way in Safari, Firefox and Chrome because it is not a browser artefact.
Where probabilistic matching is the right call
There are cases where insisting on deterministic matching costs more than it returns. Measurement and modelling work — understanding roughly how often the same person sees a campaign across devices, or sizing an addressable audience before committing budget — tolerates probabilistic input well, because an error on an individual record washes out in aggregate and no single person is contacted on the strength of it.
The line worth holding is between analysis and action. A probabilistic match that informs a budget decision is reasonable. A probabilistic match that triggers an email to a named individual is a guess about who someone is, and the consequences of being wrong land on your sending reputation and potentially on your compliance position rather than on a dashboard.
Compliance: consent does not travel between channels
A matching method is not only a technical choice. Linking records creates a more complete picture of a person than any single source held, and that picture carries obligations.
The most common failure is assuming permission transfers. M3AAWG’s position on email appending is explicit that consent given in one context does not extend to another, and that sending to someone who never gave informed consent for their address to be used that way “is never acceptable.” Resolving an identity does not create a right to contact it. The match tells you two records describe one person; it says nothing about whether that person agreed to hear from you.
Under the California Consumer Privacy Act, as summarised by the state Attorney General, consumers can request disclosure of what personal information a business has collected, used, shared or sold about them, request deletion, correct inaccurate information, and opt out of the sale or sharing of their data, with businesses required to respond within 45 days. A resolved identity graph has to be able to answer those requests at the level of a person, which means knowing which records were linked and on what basis. A probabilistic graph that cannot reconstruct why two records were joined is harder to defend when someone asks what you hold about them.
How to pressure-test a vendor’s matching claims
Four questions separate a real answer from a confident one. First, for a given match, which method produced it — if a vendor describes the system as deterministic but cannot label individual matches by method, the system is a blend and the deterministic claim covers the whole of it. Second, what is the match key, specifically: hashed email, phone, postal address, or a device signal wearing one of those names. Third, what is the refresh cadence, since a graph built on attributes IAB Tech Lab calls “relatively temporary” is only as good as its last rebuild. Fourth, what is the provenance and consent basis of the underlying records, and can it be evidenced per source rather than asserted across the file.
Be equally wary of the reverse failure. A vendor who will not name any limitation, or who quotes a single headline match rate that holds across every vertical and file quality, is quoting a sales number. Match rates move with file age, match key, record completeness and the attributes requested, and anyone who has run enough of them will say so unprompted.
Where this fits
Digital Bulldogs builds identity resolution on deterministic matching, uses probabilistic signals only where they are appropriate and says which is which, for the reasons set out above rather than because cookies were supposed to disappear. It feeds the same records behind our audience data and custom audience work, and it is the reason list hygiene and identity resolution tend to get discussed together: a graph built on stale or unconsented records produces confident matches to people who should never have been contacted.
If you want to know which method would actually serve your file, and how much of it is realistically matchable, book a free consultation and we will tell you plainly, including when the answer is that your file does not need this.
Sources
- IAB Tech Lab, Identity Solutions Guidance v1.0 — definitions of deterministic and probabilistic identifiers, the scale-versus-precision trade, and the 10-20% authentication-wall figure.
- Google, Next steps for Privacy Sandbox and tracking protections in Chrome (22 April 2025) — Chrome retains third-party cookie choice; no standalone prompt.
- Google, Update on Plans for Privacy Sandbox Technologies (17 October 2025) — the list of retired technologies, including Topics and Protected Audience.
- Apple WebKit, Tracking Prevention — Intelligent Tracking Prevention blocks all third-party cookies by default, with no exceptions.
- MDN, Third-party cookies — Firefox Total Cookie Protection partitions third-party cookies per site; Chrome does not block them by default.
- M3AAWG Position on Email Appending — consent does not transfer between channels.
- California Attorney General, California Consumer Privacy Act (CCPA) — consumer rights to know, delete, correct and opt out, and the 45-day response requirement.
