A 503 Service Unavailable means the server, or a proxy in front of it, is temporarily unable to handle the request: it is overloaded, in maintenance, has no healthy backends, or is rate limiting you. Unlike a 502, a 503 is usually a deliberate answer, and it may include a Retry-After header saying when to try again.
What does 503 Service Unavailable mean?
HTTP 503 means the server is currently unable to handle the request because of a temporary overload or scheduled maintenance, which is expected to clear after some delay. The key word is temporary. A 503 tells the client, the browser, a crawler or another service, that retrying later is the right response, and the server can say how much later with a Retry-After header, given either as a number of seconds or as an HTTP date.
In practice, 503s come from four places: the application itself (a maintenance switch, an overload guard), a reverse proxy or load balancer with no healthy backend to route to, a rate limiter, or a security layer in front of the site. Working out which of the four said it is most of the diagnosis.
What causes a 503 error?
No healthy backends
Most load balancers and proxies answer 503 when they have nowhere to send a request. The wording varies by product, and the wording is a useful clue:
- HAProxy"503 Service Unavailable: No server is available to handle this request." Every server in the backend is down or in maintenance according to its health checks.
- Envoy and service meshes"no healthy upstream" when the cluster has no healthy hosts, or "upstream connect error or disconnect/reset before headers" when connecting failed.
- Kubernetes ingress-nginxA 503 when the Service behind an Ingress has no ready endpoints: every pod is failing its readiness probe, the selector matches no pods, or the deployment is scaled to zero.
- AWS Application Load BalancerA 503 when the target group has no registered targets. Note that if every registered target is unhealthy, an ALB fails open and keeps sending traffic to all of them, so you see the targets' own errors instead.
- Apache mod_proxyA 503 when Apache cannot connect to the backend. Apache then puts that backend in an error state and keeps answering 503 for a retry period (60 seconds by default) even after the backend comes back.
- IISA 503 from HTTP.sys when the application pool is stopped, often because it crashed repeatedly and rapid-fail protection disabled it.
Overload
A server with every worker busy has two options for the next request: queue it or refuse it. When queues fill up, many stacks refuse with 503. PHP-FPM logs "server reached pm.max_children setting" when every child is busy; application servers and managed platforms reject requests when their queues overflow. Overload is usually a symptom, not a cause: a slow database query, a downstream API that got slower, or a traffic spike that arrived faster than autoscaling could add capacity.
Autoscaling and cold starts
Autoscaling reacts to metrics, and metrics lag. A traffic spike can saturate the existing instances in seconds while new ones take minutes to boot, pass health checks and join the pool. Scale-to-zero platforms add a cold start on the first request. Both produce bursts of 503 at exactly the moment traffic is highest, which is the worst time for them.
Health checks that are too strict
A health check that fails when an optional dependency is slow can take every instance out of rotation at once, turning a minor degradation into a full 503 outage. The same happens when health check timeouts are tighter than the endpoint's real response time under load: instances are marked unhealthy because they are busy, which makes the remaining ones busier.
Maintenance mode
Many applications return 503 during maintenance by design. WordPress does this while it installs updates: it writes a .maintenance file to the site root and serves "Briefly unavailable for scheduled maintenance". If an update is interrupted, the file is left behind and the site stays in maintenance mode until you delete it.
Rate limiting
The correct status for "too many requests" is 429, but plenty of rate limiters answer 503. nginx is the common example: limit_req and limit_conn reject excess requests with 503 unless you set limit_req_status 429 and limit_conn_status 429. If 503s cluster around one client IP or one busy endpoint while everything else is fine, look at rate limits first.
WAF, DDoS protection and bot challenges
Security layers sometimes answer with 503, along with 403 or 429, when they challenge or block a request. To the blocked client it looks exactly like an outage; to everyone else the site is fine. If an uptime monitor reports a fast, identical 503 on every check while the site loads in your browser, read why your monitor says 403 while the site loads fine. SutramX recognises the common vendors' block responses and labels them "Blocked by bot protection (WAF)" rather than reporting a plain 503.
How to diagnose a 503, step by step
1. Find out who sent it
Run curl -sv https://example.com/ -o /dev/null and read the headers and the body. A Server header, a vendor header (cf-ray, x-amz-cf-id, x-vercel-id), or the text of the error page usually names the layer. "No server is available" is HAProxy; "no healthy upstream" is Envoy; a branded challenge page is your CDN or WAF; your own maintenance page is your application.
2. Check for Retry-After
If the response carries Retry-After, something deliberately declared a temporary outage: maintenance mode, a rate limiter or an overload guard. curl -sI https://example.com/ | grep -i retry-after shows it. No Retry-After usually means a proxy with no backend.
3. Is it everyone, or just you?
Test from somewhere else: your phone on mobile data, or several regions at once with the multi-region website checker. A 503 that only you get points at a rate limit or WAF rule triggered by your IP. A 503 from every location points at the service itself.
4. Check backend health
Look at the load balancer's view of its targets: the HAProxy stats page, the ALB target group's health tab (aws elbv2 describe-target-health --target-group-arn ...), or kubectl get endpoints yourservice and kubectl describe pod for failing readiness probes. If every target is unhealthy, the question becomes why the health check fails, which is often a different question from why the site is down.
5. Check saturation
Look at CPU, memory, worker counts and queue lengths at the time of the errors. Search the PHP-FPM log for max_children, the application logs for queue or pool exhaustion, and the database for long-running queries and connection limits. If saturation lines up with the 503s, the fix is capacity or efficiency, not configuration.
6. Check rate limit and WAF logs
nginx logs "limiting requests" in its error log when limit_req rejects a request. CDN and WAF dashboards list blocked and challenged requests with the rule that matched. If a rule matches your monitoring, your own office, or a partner's integration, that is your 503.
How to fix a 503 Service Unavailable
- No healthy backendsRestore at least one healthy instance first, then fix the health check: make it test what the instance needs to serve traffic, give it a realistic timeout, and require several consecutive failures before removing an instance.
- OverloadFind the bottleneck before adding capacity. A slow query that holds every worker will saturate ten servers as easily as two. Cache expensive responses, add database indexes, and set timeouts on outbound calls so one slow dependency cannot hold every worker.
- Autoscaling lagScale on a leading signal such as request count or queue depth rather than CPU, keep a minimum number of warm instances, and pre-scale before known spikes such as launches and campaigns.
- Stuck maintenance modeDelete the maintenance flag (for WordPress, the .maintenance file in the site root) and finish or roll back the update that was interrupted.
- Rate limitsRaise limits that are tighter than real traffic, use burst allowances (limit_req ... burst=20 nodelay in nginx), allowlist known clients such as your monitoring and partners, and return 429 instead of 503 so the problem is obvious in logs.
- WAF and bot protectionAllowlist the clients that should get through, by IP, user agent or a secret header, rather than turning protection off.
Serving 503 correctly during maintenance
If you take a site down on purpose, a 503 is the right status, and getting it right protects your search rankings. Search engines treat a 503 as temporary and come back later. A maintenance page served with 200 tells them the page content is now "we will be back soon", and that can be indexed.
- Return status 503 for every page, not a redirect to /maintenance (a redirect returns 301 or 302).
- Add Retry-After with a realistic number of seconds, for example Retry-After: 3600 for an hour.
- In nginx: error_page 503 /maintenance.html; then location = /maintenance.html { internal; } and return 503; inside the server or location blocks you are taking down. Add add_header Retry-After 3600 always; because nginx drops add_header on error responses without always.
- Exclude your health check and monitoring paths only if you want to keep tracking the real backend during the work; otherwise let monitoring see the 503 and silence the alert instead.
- Keep maintenance short. A 503 that lasts days can lead search engines to treat the pages as gone.
How to monitor 503 errors (and maintenance windows)
503s are often brief: a burst during a deploy, a minute of overload, a scale-up that arrives late. A monitor with a long interval can miss them entirely, and one that alerts on a single failed check from a single location will page you for every blip. Choose an interval that matches how short an outage you care about, and confirm failures from more than one place.
- Monitor the endpoints that depend on the backends most likely to run out of capacity: login, search, checkout, the main API.
- Set a response time limit too. A server approaching overload gets slow before it starts refusing, so a latency alert often arrives before the first 503.
- Schedule maintenance in your monitoring tool instead of pausing monitors. In SutramX, a maintenance window keeps checks running but holds back down alerts, leaves the window out of the uptime percentage, and shows the work as planned maintenance on your status pages. If the service is still down when the window ends, the alert goes out then.
- Recurring work, such as a nightly backup that briefly takes the API offline, can be a daily or weekly maintenance window rather than a weekly argument about whether the alert was real.
- Watch for monitors that are blocked rather than down, and fix the allowlist rather than ignoring the alerts.
A 503 page says "unavailable" and nothing else. A status page that says what is affected, what still works and when the next update is due saves you a queue of support tickets. Post planned maintenance there in advance, too.
503 Service Unavailable FAQ
Is a 503 error permanent?
No. 503 means temporary by definition: overload, maintenance or no backend available right now. If a 503 lasts hours, something is stuck, such as a maintenance flag that was never removed or a health check that can never pass.
What is the difference between 503 and 502?
A 502 means a gateway tried the upstream and got an invalid answer or none. A 503 means the server or gateway decided it could not serve the request right now, because it is overloaded, in maintenance, rate limiting, or has no healthy backend to try.
Does a 503 hurt SEO?
A short 503 with Retry-After is the recommended way to take a site down and does not hurt rankings. Long-running 503s, or maintenance pages served with status 200, can.