Monitors

Website health

Scheduled site crawls that find broken links, images and scripts, mixed content and insecure forms, plus Lighthouse scores and alerts on new problems.

Website health crawls a site on a schedule and produces a report: broken links, broken images, scripts and stylesheets, mixed content and insecure forms, plus Lighthouse scores for performance, accessibility, best practices and SEO. Each run is kept as its own report, so you can compare runs and share them with clients.

Website health is separate from uptime monitoring. It does not open incidents or change a monitor's status. It tells you about problems that a single-URL check would miss.

Website health is a plan feature. See Plans & limits for which plans include it. After a downgrade, existing reports stay readable and you can still pause or delete sites.

Add a site#

  1. Go to Website health in the sidebar (under Checks & insights).
  2. Click Add site.
  3. Enter a Name and the Start URL, for example https://www.example.com.
  4. Adjust the crawl, schedule and Lighthouse settings if needed.
  5. Click Add site.

On a Daily or Weekly schedule, the first scheduled run starts at a random time within the next 24 hours, so crawls are spread out. To get a report straight away, click Run now on the site's page. Sites set to Only when I run it run only when you click Run now.

Settings#

FieldWhat it doesDefault / limits
NameLabel for the site, for example Client: Acme marketing siteRequired, up to 255 characters
Start URLThe page the crawl starts fromRequired. Must be a public http:// or https:// URL
Pages to crawl (max 500)Most pages the crawler visits per run100. 1–500
Link depth from the start pageHow many clicks away from the start page the crawler follows links3. 0–10. With 0 only the start page is crawled, but its links are still checked
ScheduleDaily, Weekly or Only when I run itWeekly
Check links to other sitesAlso checks links that point to other domainsOn
Run LighthouseRuns a Lighthouse audit after the crawlOn
Lighthouse deviceMobile (throttled) or DesktopMobile (throttled)
Extra pages for Lighthouse (up to 3, one per line)More pages to audit besides the start URLOptional, up to 3
Alert me about new problemsSends a notice when a run finds new errors or score dropsOn
Alert when a score drops by … pointsHow far a Lighthouse score must fall to count as a drop10. 1–100

A workspace can have up to 50 sites.

What the crawler checks#

The crawler starts at your Start URL and follows links to other pages on the same site. The same site means the same host name, ignoring a leading www. (so example.com and www.example.com are one site, but blog.example.com is a different site).

On every crawled page it collects links and subresources (images, scripts, stylesheets, frames and so on), then checks each one once:

IssueSeverityWhat it means
Broken linkErrorA link returns HTTP 4xx or 5xx (404 and 410 show as "page not found"), the domain does not resolve, the connection fails, or the TLS certificate is invalid
Broken resourceErrorThe same, for an image, script, stylesheet or other subresource
Mixed contentError or warningAn HTTPS page loads something over plain HTTP. Active content (scripts, frames, stylesheets, plugins) is an error because browsers block it; passive content such as images is a warning
Insecure formWarningAn HTTPS page has a form that submits to an HTTP address
RestrictedWarningThe link answered 401, 403, 429 or 999 (often sign-in pages or bot protection, not a dead link), or points to a private or internal address and was not checked
Redirect problemWarningToo many redirects or a redirect loop
Page errorErrorThe start page itself could not be loaded

A link that times out is reported as a warning, not an error, because slow third-party servers are common.

Crawl limits#

LimitValue
Pages per runYour Pages to crawl setting, at most 500
Links and resources checked per run1,000
Links checked per external host per run25
Time per page15 seconds
Time per link check10 seconds
Total crawl timeAbout 10 minutes
Redirects followed5

When a run hits a limit, the report says Stopped at the page or link limit. To cover a large site, add several sites with different start URLs (for example /blog and /docs).

How the crawler behaves#

The crawler is polite and identifies itself:

http
User-Agent: Mozilla/5.0 (compatible; SutramX-SiteHealth/1.0; +https://sutramx.com/bot)
  • It obeys robots.txt. It looks for rules for SutramX-SiteHealth (or sutramx), then for *. Pages your robots.txt disallows are skipped and counted as "skipped by robots.txt" in the report.
  • It honours Crawl-delay (up to 5 seconds) and otherwise waits a short moment between requests to the same host.
  • It does not follow nofollow links.
  • It checks links with a HEAD request first and falls back to GET.
  • It never requests private or internal addresses.

If robots.txt disallows your start page, the run fails with a message like:

text
robots.txt disallows / for our crawler. Allow "SutramX-SiteHealth" in robots.txt or choose another start page.

To allow it explicitly:

text
User-agent: SutramX-SiteHealth
Allow: /

If your firewall or bot protection blocks the crawler, allow the user agent above. See Regions & confirmation for allowlisting options.

Lighthouse scores#

When Run Lighthouse is on, each run audits the start URL and up to 3 extra pages, on the device you chose. The report shows four scores (0–100) per page:

  • Performance
  • Accessibility
  • Best practices
  • SEO

Each score shows the change since the previous report. The Lighthouse part of a run can end as:

MessageMeaning
Lighthouse is runningThe audit is in progress
Lighthouse could not finishThe audit failed, for example because the page did not load in time
Lighthouse is turned off for this siteRun Lighthouse is off
Lighthouse was unavailable for this runThe audit service was not available. The crawl results are still complete

Read a report#

Open a site to see its latest report. You can switch to any of the last 20 reports.

SectionWhat it shows
Pages crawledPages visited, and how many were skipped by robots.txt
Links and resources checkedUnique links and subresources tested
BrokenBroken links and resources, with the previous report's count for comparison
Mixed contentMixed content and insecure form findings, and the total number of warnings
LighthouseScores per audited page
IssuesEvery finding with its Type, URL, Found on (a page that links to it; Details says when several pages do) and Details

Filter issues by severity (Errors, Warnings, Errors and warnings) and by type. Click Issues CSV to download all findings, for example to send to a client or a developer.

Run statuses: Queued, Crawling, Running Lighthouse, Completed and Failed.

Alerts#

With Alert me about new problems on, SutramX sends a notice when a run:

  • finds errors that were not in the previous report (new broken links, broken resources or active mixed content),
  • sees a Lighthouse score fall by at least your threshold, or
  • could not finish.

Problems that were already in the previous report are not repeated. The notice lists what changed (up to 10 items of each kind) and links to the report.

Notices go to your workspace's alert recipients by email and to your connected chat channels and webhooks (event site_health.issues). See How alerting works.

Manage sites#

On a site's page:

  • Run now starts a run immediately. Only one run per site can be in progress (A check of this site is already running). You can start up to 20 runs an hour across all your sites.
  • Edit changes the settings. Changing the schedule queues the next scheduled run straight away.
  • Pause stops scheduled runs. Resume turns them back on.
  • Delete removes the site and all its reports.

Common questions#

Why is a working link reported as broken? Some servers refuse HEAD requests or automated clients. Responses of 401, 403, 429 and 999 are reported as Restricted warnings rather than errors for this reason. If a site you own blocks the crawler, allowlist SutramX-SiteHealth.

Why were pages skipped? They were disallowed by robots.txt, deeper than Link depth from the start page, beyond Pages to crawl, or on a different host.

Does website health affect my uptime? No. Uptime comes only from your monitors.

Last updated . Something unclear or missing on this page? Tell us at support@sutramx.com.