Technical Suite: 1.5M+ pages crawled every day, 3M+ issues fixed automatically. See the numbers →Add SEORCE as a preferred source on Google →
TECHNICAL SUITEOverviewSite AuditCrawlabilityRedirectsSchemaContentInternal LinksPerformanceAssetsReportingAutoFix
●Clustered by section, not reviewed page by page. Free to start, no card.

Find duplicate and thin pages across the whole catalog.

A filter or variant combination that turned 4,000 URLs into 40,000. Near-duplicate pages clustered by section, prioritised by crawl-budget and traffic impact — the pattern nobody sees reviewing pages one at a time.

You are here · the workflow

CRAWL→DIAGNOSE→PRIORITISE→FIX→VERIFY→MONITOR

The technical surface · this page covers

Crawl & discoveryIndexabilityRenderingRedirectsSchemaContentInternal linksPerformanceCrawlabilityAssets

Part of Technical Suite's six-stage loop. See the full system →

Common challenges

One filter combination. Thousands of near-duplicate pages.

  • A filter, a size variant, a colour variant and a sort parameter turn 4,000 real product pages into 40,000 URLs, most of them near-identical to each other.
  • Crawl budget goes to combinations nobody searches for, while the handful of pages that actually convert get buried under the ones that don't.
  • A content team ships thin pages to fill out a category structure — a handful of products, no unique copy — because the template requires a page to exist at all.
  • Nobody sees the pattern reviewing pages one at a time. It only shows up once something clusters the whole catalog at once.

What needs to happen: near-duplicates need to be found as clusters, prioritised by how much traffic and crawl budget they're actually costing, not reviewed page by page.

How SEORCE does it: the crawl covers every URL a filter or variant combination generates and clusters near-duplicate content by section. Which page should canonicalise where is a strategy decision, so it's diagnosed and ranked by impact — not applied automatically.

Where it shows up

Duplicate content, by section.

Near-duplicate clustersthis crawl
/products/shoes/*2,840
/products/apparel/*1,910
/products/accessories/*920
/blog/*310
/docs/*180
!One filter combination created 2,840 near-identical shoe pages. Canonicalising the top cluster recovers an estimated third of crawl budget spent on that section.
Found vs decided

Clustering is mechanical. Canonical strategy isn't.

100% diagnosed, ranked by impact

Finding and clustering duplicates is mechanical, so it's automatic. Deciding which page should represent a cluster is a content and business decision, so it stays with your team, with the ranked cluster list ready.

What you'll see

Inside content issues.

✓Near-duplicate clusters, grouped by section
✓Thin-content flags, by page and by section
✓Estimated crawl-budget impact per cluster
✓Clusters ranked by traffic and scale, not just count
✓Change tracking as clusters grow or shrink crawl over crawl

What does Content Issues check?

Content Issues clusters near-duplicate and thin pages across a site during the continuous crawl, grouped by section — every URL a filter or variant combination generates, not reviewed one page at a time. Clusters are ranked by estimated crawl-budget and traffic impact, so the highest-cost duplication surfaces first. Which page should represent a cluster, and how the canonical structure should resolve, is a content and business decision, so it's diagnosed and left for your team rather than applied automatically.

Key facts

  • Clusters near-duplicate pages by similarity and by section
  • Covers every URL a filter or variant combination generates, not a sample
  • Ranks clusters by estimated crawl-budget and traffic impact
  • Canonical and content strategy decisions stay with your team — clustering is automatic, the decision isn't
Last reviewed 18 September 2026 by the SEORCE team.
Common questions

Frequently asked about content issues.

Updated 18 September 2026
How does this find duplicate pages across a large catalog?+

The crawl covers every URL a filter or variant combination generates, and clusters near-duplicate content by similarity and by section, rather than requiring someone to compare pages one at a time.

Does this fix duplicate content automatically?+

No. Which page should canonicalise where is a strategy decision — SEORCE clusters the duplicates and ranks them by impact, but a person decides the canonical structure.

What counts as thin content?+

Pages with very little unique text relative to the template around them — often a symptom of a page that exists mainly to fill out a URL pattern, such as a filter combination with a handful of products and no unique copy.

Why does duplicate content matter for crawl budget?+

Every near-duplicate page the crawler visits is a page it isn't spending time on instead. At scale, thousands of near-duplicates can absorb a large share of the crawl budget that would otherwise go to pages that actually convert.

Free to start

Find the duplicates costing you crawl budget.

Clustered by section, ranked by impact, across the whole catalog.

Free forever planNo card neededFull-site crawl, first runRanked by impact