Why is my Siteimprove crawl failing?
Your Siteimprove crawl might be failing for a few reasons. Here are the most common:
Your security tool might see the Siteimprove crawler as a threat and block it.
To fix this, you need to add Siteimprove’s IP addresses to your site’s safelist.
If your site is very large, your crawl might be taking too long to complete and a new crawl starts before the first one is finished.
This is more common on very large sites (over 10,000 pages).
To fix this, please submit a Siteimprove Support Request and ask for less frequent crawls.
You have thousands of pages of repetitive results in your Siteimprove report.
This can happen with some calendar plug-ins and increases crawling time.
To fix this, please submit a Siteimprove Support Request. We can help you identify URLs that can be excluded. (Exclusions must be approved by DAP.)
Your homepage is redirecting.
Siteimprove can't use a redirected homepage.
To fix this, either:
- Remove the redirect or,
- Use a sitemap as the indexing URL instead of the homepage. Please submit a Siteimprove Support Request and we can help you with next steps.
There's an issue with your robots.txt file.
Make sure your robots.txt file isn't blocking all crawlers. To allow Siteimprove to crawl your site, add:
User-agent: SiteimproveBot
User-agent: SiteimproveBot-Crawler
Allow: /
For more information, refer to Siteimprove's update to the robots.txt parser
Your sites are on old, shared servers.
-
Use a robots.txt file to reduce the crawl speed Review the Crawl-delay on dap.berkeley.edu site’s robot.txt file (TXT file) as an example.
-
If you can’t add a robots.txt file, submit a Siteimprove Support Request. We can review other options including:
-
Having Siteimprove reduce the crawl speed.
-
Reduce the frequency of site crawls.
-
Limiting the number of pages crawled. (Requires DAP approval.)
-
Adding URL exclusions. (Requires DAP approval. This can only be used for large numbers of highly repetitive pages.)
If none of these reasons apply, submit a Siteimprove Support Request for help.
Why can't Siteimprove find all my pages?
Siteimprove works by using a domain or subdomain’s homepage as the indexing URL. It starts on this page and then crawls through the site to find all of the website pages.
'Orphan' pages
If a page isn't linked anywhere, it is called an ‘orphan’ page. Siteimprove can’t find orphan pages. The Siteimprove crawler also can’t follow a link to a page formatted as a button.
To fix this, create a sitemap then submit a Siteimprove Support Request to add it as an indexing URL.
Websites with a homepage that is no longer active.
Some departments have a lot of websites, and these sites have evolved over many years. For some of these sites, the homepage is no longer used. It might have been unpublished or it might redirect to a different website. However, the website still has active subsites.
Siteimprove can’t use a defunct or redirected homepage as an indexing URL. We can’t add URLs for subdirectories or subsites in Siteimprove.
Domains and subdomains can be used in Siteimprove.
- Example of a domain: berkeley.edu
- Example of a subdomain: plantsaregreat.berkeley.edu
Individual pages, subsites, or subdirectories cannot be used in Siteimprove.
-
Example of an individual page: plantsaregreat.berkeley.edu/hydrangeas
Note: ‘plantsaregreat.berkeley.edu/sitemap’ is an individual page also, but an exception will be made for this URL only.
Example: The homepage of an old website (plantsaregreat.berkeley.edu) now forwards visitors to a new address (plantsareawesome.berkeley.edu). The old site still has a lot of helpful content that we want to keep active but it cannot be moved to the new address.
To fix this, create a sitemap with all the old pages. Then, submit a Siteimprove Support Request. We can use the sitemap to update the Siteimprove crawl automatically, instead of manually moving all the old pages to the new site.
Website inclusions and exclusions
Important: Only Administrators can remove links, or add inclusions, exclusions, or extra index URLs. If you are a User or Power User, and you need help with these features, please submit a Siteimprove Support Request
What are inclusions and exclusions?
- An inclusion is the URL of an external page that is not on your site, but Siteimprove is crawling. Siteimprove will treat this external page the same as your site’s internal pages.
- An exclusion is a URL or URL segment that Siteimprove’s crawlers ignore on your site. You can add exclusions when website features generate a large number of repetitive pages. When Siteimprove crawls these pages, the site’s servers may be burdened. This results in long crawl times (30+ days).
Additional notes:
- An extra index URL is an additional entry point for the Siteimprove crawler. These are used when sections of a site are not linked from a homepage.
- The 'remove links’ feature tells the crawler not to check or follow a specific link. This results in no content check on these pages and no HTTP status check).