Key takeaways
Duplicate content refers to the same text accessible at several URLs, on one site or across several sites. Google does not penalize it as such: it groups duplicates and shows only one. The risk lies elsewhere: the wrong URL displayed, scattered signals, wasted crawling. Copying for manipulation purposes, however, counts as spam.
- One page = one URL: a 301 redirect or a canonical tag for variants.
- Search Console shows which version Google has chosen.
- Copying third-party content to manipulate rankings is spam.
What is duplicate content?
The term “duplicate content” means content that exists in more than one copy. We speak of duplicate content when identical or very similar content is accessible at several addresses.
Definition
Duplicate content is identical or near-identical content available at more than one URL. It is internal within the same site, external across different sites.
There are two cases. The exact duplicate: the same HTML, for example a page accessible with and without a tracking parameter. The near-duplicate: the same main text with small variations. That is the case with two product pages that differ only by color.
Most of the time, duplication is technical and unintentional. The CMS, filters, URL parameters or regional versions multiply the addresses of a single page. Nobody decided it.
Does duplicate content lead to a Google penalty?
No, as long as there is no intent to manipulate. Google says so in its page on canonicalization (updated on August 20, 2026). Some duplicate content on a site is normal and does not violate its spam policies.
Its guide to how Search works describes what happens. During indexing, Google groups similar pages and chooses a canonical URL, the most representative of the group. That is the one that appears in the results. The others remain known but are rarely shown.
Duplicate content therefore has a cost, without any penalty:
- The wrong URL displayed: Google may choose a version with parameters or a sorting page.
- Scattered signals: the links received are split between several addresses instead of adding up.
- Wasted crawling: the bot spends time on copies instead of new pages.
- Skewed tracking: traffic and positions split across several URLs in your reports.
The line is clear. Google’s spam policies (updated on August 28, 2026) target “scraping”. It consists of taking content from other sites, often in an automated way, to manipulate rankings. There, the risk of a penalty is real.
Where does duplicate content on a site come from?
Google cites five families of causes: regional variants, device variants, protocol variants, site functions and accidental variants. In practice, they take these forms:
| Cause | Example | Usual treatment |
|---|---|---|
| Protocol and host | http:// and https://, with and without www | 301 redirect to a single version |
| Trailing slash, capital letters | /page and /page/, /Page/ | 301 redirect and consistent internal links |
| URL parameters | ?utm_source=, ?sort=price, ?session= | Canonical tag pointing to the URL without parameters |
| Filters and sorting | Category filtered by color or price | Canonical or noindex depending on search demand |
| CMS taxonomies | Tags, archives by date or by author | Noindex if no value of their own |
| Regional versions | fr-FR and fr-BE with identical text | hreflang between the versions |
| Reused texts | Supplier descriptions, republished articles | Rewriting or canonical to the original |
| Test environments | Staging site accessible and indexed | Password protection |
For regional versions, Google adds a clarification in its documentation on localized versions (updated on September 21, 2026). They only count as duplicates if the main content is not translated. Pages in the same language for two countries must therefore be linked with hreflang annotations.
How do you detect duplicate content?
- Read Search Console: in URL Inspection, compare “User-declared canonical” and “Google-selected canonical”. The indexing report also lists “Duplicate” pages. See our article on Google indexing.
- Crawl the site with a tool such as Screaming Frog. According to its tutorial on duplicates, it detects exact duplicates using an MD5 hash of the HTML. Near-duplicates go through the minhash algorithm, with a similarity threshold set to 90% by default.
- Compare the tags: identical titles, meta descriptions and H1s across several URLs often signal pages that are too similar.
- Search for an exact sentence in quotation marks on Google to find external copies of your texts.
Tip
The 90% threshold is a tool setting, not a Google rule. No “acceptable” similarity rate has been published. Judge case by case. Two pages that meet the same intent often benefit from being merged, even if their texts differ.
Which solution should you choose against duplicate content?
Google ranks canonicalization signals by strength in its page How to specify a canonical URL (updated on July 10, 2026). A permanent redirect and the rel="canonical" tag are strong signals. Inclusion in the sitemap is a weak signal.
| Solution | When to use it | Effect |
|---|---|---|
| 301 redirect | The variant does not need to remain accessible | Strong signal, the user lands on the right page |
| Canonical tag | The variant must remain visible (sorting, tracking, print) | Strong signal, a hint rather than a directive |
| Sitemap | List only canonical URLs | Weak signal, as a complement |
| Noindex | Page useful to visitors but with no search value | Removes the page from the index |
| Hreflang | Language or regional versions | Associates the variants instead of pitting them against each other |
| Merging or rewriting | Two pages on the same intent | Removes the cause |
The canonical tag is defined by IETF RFC 6596 (April 2012). It goes in the head of each variant:
<link rel="canonical" href="https://www.example.com/running-shoes/">
Each indexable page also carries a canonical pointing to itself. It neutralizes parameters added by external links. The implementation rules are detailed in our guide to the canonical URL.
Warning: robots.txt and URL removal
Google advises against using robots.txt or the URL removal tool to handle duplicates. A blocked URL cannot pass its signals to the canonical version.
What should you do if another site copies your content?
Start by checking that your page is indexed and that it appears before the copy for an exact sentence. If so, the impact is often small.
- Syndication partner: ask for a canonical tag pointing to your original article, or a noindex on the copy.
- Unauthorized copy: contact the publisher, then the host. As a last resort, file a copyright removal request with Google.
- Massive, automated copying: it counts as scraping under Google’s spam policies. You can report it.
Handling duplicate content is part of technical SEO, a workstream presented in our SEO guide. The best prevention remains editorial: one page per intent and texts specific to each page.
Frequently asked questions
What does “duplicate content” mean?
It refers to identical or very similar content accessible at several URLs. It covers internal duplication, within the same site, and external duplication, across different sites.
Does Google penalize duplicate content?
No. Google states that some duplicate content is normal and does not violate its policies. It chooses a canonical URL and shows that one. Only reusing content from other sites in order to manipulate rankings is treated as spam.
What similarity rate between two pages is acceptable?
Google publishes no threshold. The 90% often cited is Screaming Frog’s default setting for detecting near-duplicates. The right criterion is intent: two pages that answer the same question benefit from being merged.
Should you use a canonical tag or a 301 redirect?
A 301 redirect if the old URL no longer needs to be accessible. A canonical tag if the variant must remain viewable, for example a sorted page or a URL with a tracking parameter. Google treats both as strong signals.
Are manufacturer product descriptions a problem?
They create external duplication with every retailer that reuses them. Google does not apply a penalty, but it has no reason to prefer your page. Rewrite first the pages of the products that matter for your revenue.
Sources
- Google Search Central, What is canonicalization, updated on August 20, 2026. Accessed on September 26, 2026.
- Google Search Central, How to specify a canonical URL with rel=”canonical” and other methods, updated on July 10, 2026. Accessed on September 26, 2026.
- Google Search Central, In-depth guide to how Google Search works, updated on December 18, 2025. Accessed on September 26, 2026.
- Google Search Central, Spam policies for Google web search, updated on August 28, 2026. Accessed on September 26, 2026.
- Google Search Central, Tell Google about localized versions of your page, updated on September 21, 2026. Accessed on September 26, 2026.
- Screaming Frog, How To Check For Duplicate Content, published on July 7, 2020, updated on October 7, 2025. Accessed on September 24, 2026.
- IETF, RFC 6596: The Canonical Link Relation, April 2012. Accessed on September 24, 2026.
Cite this article
, . (2026, September 26). Duplicate content: how to find it and fix it. Elev8 Lab. https://elev8-lab.fr/en/seo/duplicate-content/