There is no penalty, and it still costs you
Duplicate content is the most misunderstood subject in search. There is no penalty for it in the way people assume, and duplication genuinely damages small sites, for a different reason.
Duplicate content is the most misunderstood subject in search. There is no penalty for it in the way people assume, and duplication genuinely damages small sites, for a different reason.
Short answer
Duplicate content is the same or near identical content at more than one address. Search engines filter duplicates rather than penalising them: one version is kept and the rest are set aside. The real cost is that your pages compete with each other and your effort is split. Deliberate scraping is a separate matter.
A search engine finds several pages saying the same thing, picks the one it considers the best representative, and shows that one. The others are held back rather than punished.
Nothing is subtracted from your site for having them. There is no scoring deduction, and the phrase duplicate content penalty describes something that does not exist in the way it is usually meant.
What you lose is control over which version is shown, and the benefit of everything you have done being concentrated in one place rather than divided.
The phrase duplicate content penalty describes something that does not exist in the way it is usually meant, which is worth saying plainly because it drives a lot of wasted effort.
Address variants, which are the largest single source. The same page reachable with and without a slash, with and without www, with a tracking parameter attached.
Boilerplate service descriptions repeated across twenty pages with only a place name changed. That is the version that most damages a local business site.
Manufacturer product descriptions used verbatim, which every competitor selling the same item also used. And a staging or development copy of the site that was never made private.
Exact copies are easy to spot and easy to fix. Near duplicates are the real difficulty, because they look like distinct pages to the person who wrote them.
A set of city pages where the only differences are the place name and a phone number is one page in the engine's view, however many addresses it occupies.
This is the pattern our own uniqueness check exists to catch, and it measures what remains once the shared words are stripped out. A page that fails that check would fail the engine's judgment too.
Exact copies are easy to spot. Near duplicates are the real difficulty, because they look like distinct pages to whoever wrote them.
Start with the plumbing, because it is mechanical and reliable. Serve one address form, redirect the rest, and give every page a self referencing canonical.
Then look at the content. Where two pages genuinely serve the same intent, merge them into one better page and redirect the loser. That almost always performs better than either did alone.
Where pages should be distinct, make them distinct. Different examples, different local detail, different questions answered. Swapping a noun is not differentiation.
Merging two pages that serve one intent almost always outperforms either alone, which is covered in keyword cannibalization.
Copying somebody else's content onto your site is a copyright matter quite apart from any search consequence, and it can be reported.
Publishing large volumes of content that adds nothing, simply to occupy more addresses, is described in search spam policies and does attract action.
The distinction is intent and value. Accidental address variants are a tidying problem. Mass produced near identical pages are a policy problem.
It happens, usually to sites that are doing well. Most of the time the engine correctly identifies the original and nothing needs doing.
Where a copy outranks the original, that is worth reporting, and the report route is published and free to use.
Publishing first, being indexed first and having a stronger site all help you win that contest without any intervention.
Every page written by Website Builder Studio is measured against its siblings for how much survives once shared wording is removed, and a page that is too close to another does not publish.
Each page also has one address, one canonical and one entry in the sitemap, which removes the mechanical sources before they appear.
That is the same judgment an indexer makes, applied before the page exists rather than discovered months later in a report.
The measure strips the distinguishing token first, so a page that is only a place-name swap fails before it publishes rather than after it is crawled.
Not in the way the phrase suggests. Duplicates are filtered rather than punished, and one version is kept. The cost is lost control and divided effort.
It is not punished, and every competitor has the same words, so nothing distinguishes your page. Writing your own description is what gives it a reason to rank.
If you could swap two pages by changing a single noun, they are too similar. The practical test is whether a reader would get anything different from each.
Short quoted passages within your own work are normal and fine. The concern is pages that are largely somebody else's content with little added.
15-day free trial. Card required. Cancel before day 15 and you pay nothing.