Amazon DSP Supports Creation of Lookalike Audiences From: The Three Seed Sources
Amazon DSP builds lookalike (similar) audiences from an advertiser's own existing audiences — specifically hashed customer lists (CRM or email data, uploaded and hashed for privacy), pixel or Amazon Ad Tag event data (site visitors and converters), and audiences transferred in from a connected data management platform (DMP). All three are advertiser-owned seeds, not audiences Amazon builds on its own.
What this looks like across the book we manage
What's actually being asked here
This phrasing shows up most often as a fragment of an Amazon Ads certification exam question, cut off before the answer choices. The full question asks which sources an advertiser can use as a seed to build a lookalike audience inside Amazon DSP — not what Amazon's own shopper data looks like, and not a general definition of lookalike modeling. The honest, complete answer is three seed types, all of them the advertiser's own existing audience data, not Amazon-sourced.
The three seed sources, one at a time
Hashed audiences are customer lists an advertiser already has — email addresses or other identifiers from a CRM system — uploaded to Amazon Ads and hashed before matching, so Amazon never receives the raw customer data itself. This is the most common seed for brands with an established customer database outside of Amazon.
Pixel or Amazon Ad Tag event audiences are built from behavior Amazon can observe directly through a tracking tag placed on the advertiser's own site — page views, add-to-carts, purchases — turned into an audience of people who took a specific action, which then becomes the seed for a similar-audience model.
DMP-transferred audiences are segments an advertiser has already built inside a connected data management platform and passes into Amazon DSP through that integration, rather than uploading a raw list or relying on Amazon's own tag.
Amazon's own advertiser-audiences documentation describes this as giving advertisers "a way of incorporating their existing audiences into their Amazon Ad campaigns," and describes building a lookalike "from their hashed audience" specifically — the seed is always something the advertiser already owns or has already defined, even when Amazon's modeling is what expands it into a larger similar audience.
How the modeling actually turns a seed into a bigger audience
Once a seed is in place, Amazon's similar-audiences capability uses machine learning to score the broader pool of Amazon shoppers against that seed's behavioral pattern, surfacing new people who resemble it without Amazon ever having to reveal who's in the original seed. This is where Amazon's retail signal becomes relevant — not as the seed itself, but as the pool the model searches to find new matches. The seed has to come from the advertiser; the matching happens against Amazon's own shopper data.
It's worth being precise about what this means for privacy: the advertiser never learns the identity of any individual matched into the resulting similar audience, and Amazon never receives the advertiser's raw customer data in the hashed-audience case — only the hashed values, matched against Amazon's own hashed records. Neither party sees the other's underlying identity data at any point in the process, which is also why Amazon can activate this feature without the advertiser needing its own separate consent or data-sharing infrastructure beyond the standard upload flow.
A worked example: what a thin seed actually costs you
Say a brand uploads a hashed CRM list of 400 past customers as its lookalike seed. That's below the volume most platforms need to find a stable pattern — a seed that small produces a similar audience the model is essentially guessing at, because there isn't enough shared behavior across 400 people to separate signal from noise. The fix isn't a bigger percentage or a different targeting setting; it's a bigger or more homogeneous seed. A pixel-based seed of everyone who purchased in the last 90 days, even if it's a few thousand people who all did the exact same thing, will consistently outperform a larger but mixed CRM list where browsers, one-time buyers, and repeat customers are all lumped in together.
The mistake, and what to check when a lookalike isn't working
The common mistake is treating any of the three seed types as interchangeable and picking whichever is easiest to export, rather than the one that best matches the campaign's actual goal. A DMP segment built for a different channel's definition of "high value" won't necessarily translate into a useful Amazon seed. If a lookalike audience isn't performing, check seed size first, seed homogeneity second — mixed intent inside one seed dilutes the pattern the model is trying to find — and seed freshness third, since a seed built from data six months old is matching against who the customer was, not who they are now.
In our own accounts, DSP prospecting audiences built this way run around two-thirds new-to-brand, and the visible effect tends to show up first in rising branded search on Amazon rather than in the DSP report itself — a brand watching only its DSP dashboard will systematically undervalue what a well-built lookalike is actually doing.
Where Dr. DSP fits
Dr. DSP is Amazon DSP run as a managed product by Full Circle, which has managed more than $500M in revenue across 100+ brands. Picking the right seed — and checking its size and freshness before blaming the audience — is a small, specific decision that quietly decides whether a lookalike campaign works. A reader who never buys anything from us should still leave this page knowing the exam-style answer to this question and, more usefully, which of the three seed types actually fits their own customer data.
| Seed source | Where it comes from | Best fit |
|---|---|---|
| Hashed audiences | Uploaded CRM or email list, hashed before matching | Brands with an established customer database off Amazon |
| Pixel / Amazon Ad Tag events | Behavior tracked directly on the advertiser's own site | Brands wanting a seed built from recent, specific actions |
| DMP-transferred audiences | A segment already built in a connected data management platform | Brands with an existing cross-channel data stack |
Which one you should actually pick
A brand asking this question is almost always studying for a certification exam or evaluating whether it has enough of its own data to use the feature at all. Either way, the honest answer is the same: all three seed types belong to the advertiser, not Amazon, and a brand with none of them yet — no CRM list, no pixel history, no DMP — isn't ready for lookalike audiences regardless of budget.
Shortlist on the job, not the feature grid. Pull your search-term report for the last 90 days and total the spend against terms that produced no orders — 48.5% across the 47 brands above. Then ask each vendor on your list what they would do about it in week one, and see who answers with a process rather than a screenshot.
Common questions
Amazon DSP supports creation of lookalike audiences from what, exactly?
Three advertiser-owned seed sources: hashed customer lists (CRM or email data, uploaded and hashed), pixel or Amazon Ad Tag event data from the advertiser's own site, and audiences transferred in from a connected data management platform. All three require the advertiser to already have or define the seed audience — Amazon's modeling expands it, but doesn't create the seed itself.
Can Amazon DSP build a lookalike audience without any advertiser data at all?
No. Every lookalike or similar audience on Amazon DSP starts with a seed the advertiser provides — hashed, pixel-based, or DMP-transferred. A brand with no existing customer data, tag history, or DMP segment has nothing to seed a lookalike with, and is generally better served by in-market or contextual targeting until it accumulates one.
What's the minimum seed size for an Amazon DSP lookalike audience to work well?
Amazon doesn't publish a hard minimum, but the general principle holds across every platform that does this kind of modeling: a seed in the low hundreds is too thin for the model to separate a real pattern from noise, and a seed a few thousand strong that's all done the same thing outperforms a much larger but mixed one.
Is a pixel-based seed better than a CRM-based one?
Neither is universally better — they answer different questions. A pixel-based seed reflects recent, specific behavior on the advertiser's own site, which tends to be fresher. A CRM-based seed can be filtered by lifetime value or purchase history, which a pixel event alone can't tell you. The right choice depends on which signal the campaign actually needs.
Dr. DSP is Amazon DSP — the Demand-Side Platform, not the delivery franchise — run daily by Fable 5 with operators from a $500M+ Amazon team supervising. You pick the approval level, we reconcile in Amazon Marketing Cloud, and Orbit is included. First 30 days free, priced on the call.
Book a Dr. DSP demoRead next
- Salsify Pricing: No Public Price, and a Category NotePricing · salsify pricing
- ChannelAdvisor Alternative: Two Jobs, One Hard ExitAlternative · channeladvisor alternative
- Intentwise Pricing: Quote-Only, and What It BuysPricing · intentwise pricing
- Pacvue Pricing: What Quote-Only Really Means for BuyersPricing · pacvue pricing
- Quartile Pricing: What the Terms Commit You ToPricing · quartile pricing
- Perpetua Pricing: What the Page Shows, and What It Doesn'tPricing · perpetua pricing