Find duplicate and thin pages across the whole catalog.
A filter or variant combination that turned 4,000 URLs into 40,000. Near-duplicate pages clustered by section, prioritised by crawl-budget and traffic impact — the pattern nobody sees reviewing pages one at a time.
You are here · the workflow
The technical surface · this page covers
Part of Technical Suite's six-stage loop. See the full system →
One filter combination. Thousands of near-duplicate pages.
- A filter, a size variant, a colour variant and a sort parameter turn 4,000 real product pages into 40,000 URLs, most of them near-identical to each other.
- Crawl budget goes to combinations nobody searches for, while the handful of pages that actually convert get buried under the ones that don't.
- A content team ships thin pages to fill out a category structure — a handful of products, no unique copy — because the template requires a page to exist at all.
- Nobody sees the pattern reviewing pages one at a time. It only shows up once something clusters the whole catalog at once.
What needs to happen: near-duplicates need to be found as clusters, prioritised by how much traffic and crawl budget they're actually costing, not reviewed page by page.
How SEORCE does it: the crawl covers every URL a filter or variant combination generates and clusters near-duplicate content by section. Which page should canonicalise where is a strategy decision, so it's diagnosed and ranked by impact — not applied automatically.
Duplicate content, by section.
Clustering is mechanical. Canonical strategy isn't.
Finding and clustering duplicates is mechanical, so it's automatic. Deciding which page should represent a cluster is a content and business decision, so it stays with your team, with the ranked cluster list ready.
Inside content issues.
For large sites, duplication is a crawl-efficiency problem first.
What does Content Issues check?
Content Issues clusters near-duplicate and thin pages across a site during the continuous crawl, grouped by section — every URL a filter or variant combination generates, not reviewed one page at a time. Clusters are ranked by estimated crawl-budget and traffic impact, so the highest-cost duplication surfaces first. Which page should represent a cluster, and how the canonical structure should resolve, is a content and business decision, so it's diagnosed and left for your team rather than applied automatically.
Key facts
- Clusters near-duplicate pages by similarity and by section
- Covers every URL a filter or variant combination generates, not a sample
- Ranks clusters by estimated crawl-budget and traffic impact
- Canonical and content strategy decisions stay with your team — clustering is automatic, the decision isn't
Frequently asked about content issues.
The crawl covers every URL a filter or variant combination generates, and clusters near-duplicate content by similarity and by section, rather than requiring someone to compare pages one at a time.
No. Which page should canonicalise where is a strategy decision — SEORCE clusters the duplicates and ranks them by impact, but a person decides the canonical structure.
Pages with very little unique text relative to the template around them — often a symptom of a page that exists mainly to fill out a URL pattern, such as a filter combination with a handful of products and no unique copy.
Every near-duplicate page the crawler visits is a page it isn't spending time on instead. At scale, thousands of near-duplicates can absorb a large share of the crawl budget that would otherwise go to pages that actually convert.
Find the duplicates costing you crawl budget.
Clustered by section, ranked by impact, across the whole catalog.