The term “first-party data” shows up in nearly every measurement article now, but it only makes full sense next to its three neighbors: zero, second, and third-party data. All four mean “data about a customer” — the difference is where the data comes from and how reliable it is. Saying “let’s build a first-party data strategy” without pinning down that distinction is like writing a recipe without knowing which ingredient you’re working with.
Four Layers, One Table
| Layer | Source | How it’s collected | Reliability |
|---|---|---|---|
| Zero-party | The customer themselves | Surveys, preference forms, calculators, product finders — information the user knowingly hands over | Highest — the user states what they want directly |
| First-party | Your own user behavior | Site/app usage, purchase history, CRM records, email subscriptions | High — collected with your own consent, stays in your own system |
| Second-party | A trusted partner’s first-party data | Mutual, contractual data sharing between two businesses | Medium-high — the source is known, but the data arrives indirectly |
| Third-party | Unrelated data intermediaries | Ad networks, data brokers, cookie-based tracking chains | Shrinking — source is hard to verify, the layer most exposed to browser restrictions |
Zero-Party Data: What the User Chooses to Tell You
Zero-party data isn’t inferred — it’s stated directly. A skincare brand’s “what’s your skin type?” quiz, an insurance site’s premium calculator, or an e-commerce site’s “which categories interest you?” preference form all fall into this layer. According to Forrester’s definition, zero-party data is data “that a customer intentionally and proactively shares with a brand” — which is what separates it from first-party data inferred from behavior.
Its value comes from that directness: it requires no interpretation. A user browsing the running-shoes page three times is a guess — maybe they’re shopping for a gift. A user typing “I’m looking for running shoes” is a statement. The tradeoff is scale — the share of users who fill out a survey or form is usually small, so zero-party data alone doesn’t provide reach.
First-Party Data: What Stays in Your Own System
First-party data is what a business collects through its own channels — site, app, CRM, email list, point of sale. Our first-party data strategy article covers in detail how this data gets connected to the ad system through tools like Customer Match and custom audiences. The critical distinction here is ownership: this data depends on your own infrastructure, not on a browser setting or a platform’s policy.
Second-Party Data: A Trusted Partner’s First-Party Data
Second-party data is the least-known layer but a practical one: a business contractually opening up its own first-party data to a partner. An airline and a hotel chain, for example, might access each other’s customer segments under a shared loyalty program — each side shares data it collected itself, with no intermediary data broker involved. That’s the key difference from third-party data: the source is known, and there’s usually a direct commercial relationship between the two parties.
The risk follows from that same relationship: the quality of shared data depends on the partner’s own collection discipline, and access ends when the contract does.
Third-Party Data: Data From the Intermediary Chain
Third-party data is what an intermediary with no direct relationship to the user — an ad network, data broker, or measurement company — collects, combines, and sells across various sites. The classic example is the third-party cookie: a user visits site A, the same ad network’s cookie also runs on sites B and C, and those three visits get merged into a single profile. This is exactly the layer targeted by the restrictions covered in our Safari ITP and Google’s third-party cookie decision articles.
This layer’s advantage was scale — advertisers could reach millions of users they never collected themselves. Its weakness comes from the same place: the data’s source can’t be verified, user consent is indirect, and it’s vulnerable to decisions made by browser and OS makers.
Why This Distinction Matters Now
As third-party data’s reliability drops, the other three layers aren’t affected at the same rate. Restrictions like iOS App Tracking Transparency and Meta’s Limited Data Use specifically target the third-party tracking chain; zero and first-party data, because they come from a user’s direct relationship with the business itself, stay mostly exempt from these restrictions. That’s why measurement strategy has shifted from “what replaces the third-party cookie?” to “how do we grow zero and first-party data?”
The practical outcome: when a business adds one preference question to a signup form (zero-party), keeps its CRM regularly updated (first-party), or sets up a data-sharing relationship with a trusted partner (second-party), it builds a measurement foundation that doesn’t depend on the third-party cookie surviving.
Summary
The difference between the four data layers comes down to source and reliability: zero-party is the user’s own statement, first-party is behavioral data you collect in your own system, second-party is a trusted partner’s first-party data, and third-party is data from intermediary chains that keeps shrinking. As cookie-based measurement gets restricted, strategy’s center of gravity is shifting toward the first two layers — the ones you actually control.