Indexing problems kept showing up in client audits with such regularity that we started keeping score. So CS Web Solutions turned it into a proper study: 500 websites, 15+ industries, examined with the same tools and the same checklist we’ve refined over 10+ years of working with startups, mid-sized companies, and enterprise organizations.
The headline finding surprised even us. Most websites publish far more than Google ever agrees to index, and most owners have no idea. Our toolkit was the usual suspects: Google Search Console, Screaming Frog, Ahrefs, Semrush, and PageSpeed Insights. Every automated finding got a human second look. Here’s what the research showed us.
Top Indexing Problems We Found in 500 Websites
The same ten issues surfaced everywhere. SaaS platforms, ecommerce stores, healthcare providers, law firms, manufacturers, real estate portals, education sites, restaurants, financial services firms, agencies, different businesses, identicalproblems. The only thing that changed with company size was how many zeros were on the traffic being lost.
1. Pages Discovered but Not Indexed
The most common status in Search Console, sitting on 61% of sites, and quietly the cruelest. Google knows your page exists. It wrote the URL down. Then it decided other things were more worth its time. Crawl budget, weak internal links, and low perceived value drive the snub, and because the page technically “exists,” nobody ever investigates. Our fix was consistent across the sample: link to these pages from your strongest, most-crawled pages, and give Google a reason to come back.
| Why Google Skipped the Page | Share of Affected Sites |
| Crawl budget limits on large sites | 38% |
| Weak or missing internal links | 34% |
| Low perceived page value | 28% |
2. Crawled but Currently Not Indexed
This one stings more, because Google actually showed up, read the page, and passed. That’s a quality verdict, delivered politely. Thin content, near-duplicates, low-value tag archives, and weak linking filled this bucket, and ecommerce sites were the repeat offenders, we found one store where 340 product variants differed by colour swatch alone.
3. Poor Internal Linking
Google discovers the web by following links, which makes this next number painful: 44% of sites buried key pages four or more clicks from the homepage. We saw money pages with fewer internal links than the privacy policy. Repeatedly. Menus linked to top categories and stopped, while blog posts floated in space, connected to nothing.
Flatten the architecture, add contextual links from pages with authority, and watch stranded pages get indexed within weeks. It worked on dozens of sites in this very sample.
4. Orphan Pages
Some 39% of sites carried pages with zero internal links pointing at them, alive in the sitemap, dead everywhere else. Retired campaign URLs, old landing pages, service pages from three redesigns ago. Someone paid good money to build every one of them. Screaming Frog’s crawl-versus-sitemap comparison finds them in minutes, and a single contextual link usually brings them back from the beyond.
5. Incorrect Canonical Tags
Canonical chaos touched 35% of sites, and the varieties were genuinely creative. Canonicals pointing at a domain the company had abandoned two migrations ago. Parameter URLs proudly self-referencing when they should defer to a parent page. Sitemaps and canonicals contradicting each other like witnesses in a courtroom. When your signals argue, Google picks a winner itself, and its taste is questionable.
6. Blocked by Robots.txt
A quarter of sites were blocking content they desperately wanted indexed. The greatest hits: a disallow rule copied from staging to production and forgotten, blocked CSS and JS folders that left Google rendering blank pages, and revenue pages sitting behind a disallow line while development files roamed free. One misplaced slash in that little text file can erase an entire site from search.
| Robots.txt Error | Share of Affected Sites |
| Staging disallow rules copied to production | 36% |
| Blocked CSS/JS folders breaking rendering | 31% |
| Revenue pages behind a disallow line | 22% |
| Development files left crawlable | 11% |
7. Noindex on Important Pages
Noindex leftovers haunted 43% of sites. Tags applied during a redesign and never removed. Plugins cheerfully no indexing entire post types. Template-level tags inherited by pages that deserved better. The good news is that recovery after removal tends to be quick, some pages we cleaned up were back in the index within days.
8. Duplicate Content Issues
URL parameters, filterable categories, pagination, and CMS quirks multiplied thousands of duplicate URLs on affected sites. Ecommerce platforms led the parade, minting a fresh URL for every sort order and filter combination a shopper could dream up. Canonical consolidation and disciplined parameter handling brought crawl budgets back from the brink.
| Duplication Source | Share of Affected Sites |
| URL parameters (filters, sort orders) | 44% |
| Category and tag page overlap | 27% |
| Pagination sequences | 17% |
| CMS auto-generated duplicates | 12% |
9. JavaScript Rendering Problems
React, Vue, and Angular builds produced most of the 18% with rendering failures. Content that only exists after client-side JavaScript runs reaches Google late, partially, or occasionally never. Server-side rendering or pre-rendering fixed it in every case we tested, and the indexation gains showed up within weeks. Every single time.
| Framework / Setup | Share of Rendering Failures |
| React (client-side only) | 46% |
| Vue | 24% |
| Angular | 19% |
| Other JS-heavy builds | 11% |
10. Slow Page Speed Affecting Crawl Efficiency
Slow servers make Google impatient. Sites failing Core Web Vitals, server responses past 600ms, bloated JavaScript, heavyweight images, consistently got fewer pages crawled per day. Think of speed work as indexing work wearing a different hat: every millisecond you save lets Googlebot see more of your site per visit.
| Speed Factor | Share of Slow Sites Affected |
| Server response above 600ms | 47% |
| Oversized JavaScript bundles | 39% |
| Unoptimized images | 52% |
| Failing Core Web Vitals overall | 58% |
Â
Final Thoughts
Five hundred audits later, our conclusion is almost embarrassingly simple: indexing problems hide in plain sight, and fixing them delivers some of the fastest wins in all of technical SEO. CS Web Solutions brings 10+ years of technical SEO experience as an award-winning digital agency serving startups, mid-sized businesses, and enterprise organizations, comprehensive technical ai powered SEO audits, enterprise-scale crawl and indexing optimization, advanced schema and structured data expertise, data-driven strategies, and ongoing monitoring that catches problems before they cost you traffic.
Call us at 905-890-2222 or email [email protected], and we’ll show you exactly what Google has been missing on your site.

Vin Sonpal is based in Mississauga, Ontario, and is the founder of CS Web Solutions, established in 2015. He works across web, mobile, and digital platforms, helping businesses build online systems that are practical, scalable, and designed to support long-term growth.

