Website health
Scheduled site crawls that find broken links, images and scripts, mixed content and insecure forms, plus Lighthouse scores and alerts on new problems.
Website health crawls a site on a schedule and produces a report: broken links, broken images, scripts and stylesheets, mixed content and insecure forms, plus Lighthouse scores for performance, accessibility, best practices and SEO. Each run is kept as its own report, so you can compare runs and share them with clients.
Website health is separate from uptime monitoring. It does not open incidents or change a monitor's status. It tells you about problems that a single-URL check would miss.
Website health is a plan feature. See Plans & limits for which plans include it. After a downgrade, existing reports stay readable and you can still pause or delete sites.
Add a site#
- Go to Website health in the sidebar (under Checks & insights).
- Click Add site.
- Enter a Name and the Start URL, for example
https://www.example.com. - Adjust the crawl, schedule and Lighthouse settings if needed.
- Click Add site.
On a Daily or Weekly schedule, the first scheduled run starts at a random time within the next 24 hours, so crawls are spread out. To get a report straight away, click Run now on the site's page. Sites set to Only when I run it run only when you click Run now.
Settings#
| Field | What it does | Default / limits |
|---|---|---|
| Name | Label for the site, for example Client: Acme marketing site | Required, up to 255 characters |
| Start URL | The page the crawl starts from | Required. Must be a public http:// or https:// URL |
| Pages to crawl (max 500) | Most pages the crawler visits per run | 100. 1–500 |
| Link depth from the start page | How many clicks away from the start page the crawler follows links | 3. 0–10. With 0 only the start page is crawled, but its links are still checked |
| Schedule | Daily, Weekly or Only when I run it | Weekly |
| Check links to other sites | Also checks links that point to other domains | On |
| Run Lighthouse | Runs a Lighthouse audit after the crawl | On |
| Lighthouse device | Mobile (throttled) or Desktop | Mobile (throttled) |
| Extra pages for Lighthouse (up to 3, one per line) | More pages to audit besides the start URL | Optional, up to 3 |
| Alert me about new problems | Sends a notice when a run finds new errors or score drops | On |
| Alert when a score drops by … points | How far a Lighthouse score must fall to count as a drop | 10. 1–100 |
A workspace can have up to 50 sites.
What the crawler checks#
The crawler starts at your Start URL and follows links to other pages on the same site. The same site means the same host name, ignoring a leading www. (so example.com and www.example.com are one site, but blog.example.com is a different site).
On every crawled page it collects links and subresources (images, scripts, stylesheets, frames and so on), then checks each one once:
| Issue | Severity | What it means |
|---|---|---|
| Broken link | Error | A link returns HTTP 4xx or 5xx (404 and 410 show as "page not found"), the domain does not resolve, the connection fails, or the TLS certificate is invalid |
| Broken resource | Error | The same, for an image, script, stylesheet or other subresource |
| Mixed content | Error or warning | An HTTPS page loads something over plain HTTP. Active content (scripts, frames, stylesheets, plugins) is an error because browsers block it; passive content such as images is a warning |
| Insecure form | Warning | An HTTPS page has a form that submits to an HTTP address |
| Restricted | Warning | The link answered 401, 403, 429 or 999 (often sign-in pages or bot protection, not a dead link), or points to a private or internal address and was not checked |
| Redirect problem | Warning | Too many redirects or a redirect loop |
| Page error | Error | The start page itself could not be loaded |
A link that times out is reported as a warning, not an error, because slow third-party servers are common.
Crawl limits#
| Limit | Value |
|---|---|
| Pages per run | Your Pages to crawl setting, at most 500 |
| Links and resources checked per run | 1,000 |
| Links checked per external host per run | 25 |
| Time per page | 15 seconds |
| Time per link check | 10 seconds |
| Total crawl time | About 10 minutes |
| Redirects followed | 5 |
When a run hits a limit, the report says Stopped at the page or link limit. To cover a large site, add several sites with different start URLs (for example /blog and /docs).
How the crawler behaves#
The crawler is polite and identifies itself:
User-Agent: Mozilla/5.0 (compatible; SutramX-SiteHealth/1.0; +https://sutramx.com/bot)- It obeys
robots.txt. It looks for rules forSutramX-SiteHealth(orsutramx), then for*. Pages yourrobots.txtdisallows are skipped and counted as "skipped by robots.txt" in the report. - It honours
Crawl-delay(up to 5 seconds) and otherwise waits a short moment between requests to the same host. - It does not follow
nofollowlinks. - It checks links with a
HEADrequest first and falls back toGET. - It never requests private or internal addresses.
If robots.txt disallows your start page, the run fails with a message like:
robots.txt disallows / for our crawler. Allow "SutramX-SiteHealth" in robots.txt or choose another start page.To allow it explicitly:
User-agent: SutramX-SiteHealth
Allow: /If your firewall or bot protection blocks the crawler, allow the user agent above. See Regions & confirmation for allowlisting options.
Lighthouse scores#
When Run Lighthouse is on, each run audits the start URL and up to 3 extra pages, on the device you chose. The report shows four scores (0–100) per page:
- Performance
- Accessibility
- Best practices
- SEO
Each score shows the change since the previous report. The Lighthouse part of a run can end as:
| Message | Meaning |
|---|---|
| Lighthouse is running | The audit is in progress |
| Lighthouse could not finish | The audit failed, for example because the page did not load in time |
| Lighthouse is turned off for this site | Run Lighthouse is off |
| Lighthouse was unavailable for this run | The audit service was not available. The crawl results are still complete |
Read a report#
Open a site to see its latest report. You can switch to any of the last 20 reports.
| Section | What it shows |
|---|---|
| Pages crawled | Pages visited, and how many were skipped by robots.txt |
| Links and resources checked | Unique links and subresources tested |
| Broken | Broken links and resources, with the previous report's count for comparison |
| Mixed content | Mixed content and insecure form findings, and the total number of warnings |
| Lighthouse | Scores per audited page |
| Issues | Every finding with its Type, URL, Found on (a page that links to it; Details says when several pages do) and Details |
Filter issues by severity (Errors, Warnings, Errors and warnings) and by type. Click Issues CSV to download all findings, for example to send to a client or a developer.
Run statuses: Queued, Crawling, Running Lighthouse, Completed and Failed.
Alerts#
With Alert me about new problems on, SutramX sends a notice when a run:
- finds errors that were not in the previous report (new broken links, broken resources or active mixed content),
- sees a Lighthouse score fall by at least your threshold, or
- could not finish.
Problems that were already in the previous report are not repeated. The notice lists what changed (up to 10 items of each kind) and links to the report.
Notices go to your workspace's alert recipients by email and to your connected chat channels and webhooks (event site_health.issues). See How alerting works.
Manage sites#
On a site's page:
- Run now starts a run immediately. Only one run per site can be in progress (
A check of this site is already running). You can start up to 20 runs an hour across all your sites. - Edit changes the settings. Changing the schedule queues the next scheduled run straight away.
- Pause stops scheduled runs. Resume turns them back on.
- Delete removes the site and all its reports.
Common questions#
Why is a working link reported as broken? Some servers refuse HEAD requests or automated clients. Responses of 401, 403, 429 and 999 are reported as Restricted warnings rather than errors for this reason. If a site you own blocks the crawler, allowlist SutramX-SiteHealth.
Why were pages skipped? They were disallowed by robots.txt, deeper than Link depth from the start page, beyond Pages to crawl, or on a different host.
Does website health affect my uptime? No. Uptime comes only from your monitors.
Related
Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.