A 502 Bad Gateway means a proxy, load balancer or CDN tried to reach the server behind it and got no valid response: the connection was refused, reset or answered with something that was not HTTP. The usual cause is an application process that crashed, is not listening, or closed a kept-alive connection the proxy reused.
What does 502 Bad Gateway mean?
HTTP 502 is defined as: a server acting as a gateway or proxy received an invalid response from the upstream server it contacted while trying to fulfil the request. A request to a typical site passes through a CDN, a load balancer and a reverse proxy such as nginx before it reaches the process that runs your code: Node.js, gunicorn, PHP-FPM, Puma or a Java servlet container.
Each of those hops can produce a 502 about the hop behind it. So the first job is never "fix the 502". It is "find out which hop said it, and what it saw when it looked upstream".
- 502 Bad GatewayThe proxy reached out and got a broken answer or none at all: connection refused, connection reset, the upstream closed the socket before sending headers, or it sent bytes that were not valid HTTP.
- 503 Service UnavailableSomeone deliberately answered "not now": no healthy backends, a rate limit, maintenance mode, or an overloaded server shedding load.
- 504 Gateway TimeoutThe proxy connected and waited, and the upstream did not answer within the proxy's timeout. The upstream may still be working on the request.
The difference between 502 and 504 is worth holding on to, because it splits the search space in half. A 502 says the upstream is unreachable or misbehaving. A 504 says the upstream is reachable but slow. The full list of codes, with what each one means for monitoring, is in the HTTP status code reference.
What causes a 502 Bad Gateway?
The upstream process is down or not listening
The most common cause by far. The application crashed, was killed by the kernel's out-of-memory killer, failed to start after a deploy, or is restarting in a loop. The proxy tries to connect, nothing is listening on the port or socket, the connection is refused, and the proxy returns 502.
A close cousin is the process that is running but listening somewhere else: on port 3001 instead of 3000 after a config change, or on 127.0.0.1 inside a container when the proxy connects from another container and needs it on 0.0.0.0.
The upstream died in the middle of the request
If the application crashes or a worker is killed while handling a request, the proxy sees the connection close before a response arrives. Gunicorn is the classic example: its default worker timeout is 30 seconds, and a request that runs longer gets its worker killed (the log says WORKER TIMEOUT), so nginx sees "upstream prematurely closed connection" and returns 502, not 504. The same pattern appears with PHP-FPM's request_terminate_timeout and with any container that is OOM-killed mid-request.
Keep-alive timeout mismatch
This one produces intermittent 502s that nobody can reproduce. Load balancers and proxies keep connections to the upstream open and reuse them. If the upstream closes an idle connection at the same moment the proxy sends a new request down it, the request fails with a reset and the client gets a 502.
The rule is simple: the upstream's keep-alive timeout must be longer than the proxy's idle timeout, so the proxy always closes first. The defaults get this wrong. An AWS Application Load Balancer keeps idle connections for 60 seconds by default; Node.js closes idle keep-alive connections after 5 seconds and gunicorn after 2. In Node, set server.keepAliveTimeout = 65000 and server.headersTimeout = 66000. In gunicorn, pass --keep-alive 65. Match the numbers to your own load balancer's idle timeout.
Proxy configuration errors
- Wrong upstream address or socket path, often after renaming a service or upgrading PHP (the socket moves from php8.1-fpm.sock to php8.3-fpm.sock and nginx still points at the old one).
- Permission denied on a Unix socket: nginx runs as www-data or nginx, and the socket is owned by a user it cannot write to.
- Response headers bigger than the proxy buffer. nginx logs "upstream sent too big header" and returns 502; large cookies or long redirect URLs are the usual trigger. Raise proxy_buffer_size (or fastcgi_buffer_size) and proxy_buffers.
- A stale DNS answer. nginx resolves upstream host names once, at start-up. If a container or managed service gets a new IP address, nginx keeps connecting to the old one until it is reloaded. Use a resolver directive with the upstream address in a variable if your upstreams move.
- Speaking the wrong protocol: proxying with http:// to a port that expects TLS, or the reverse.
Load balancer health and deploys
During a rolling deploy, a load balancer or Kubernetes ingress can keep sending traffic to an instance that has already started shutting down. Requests in flight are cut off and return 502. The fix is graceful shutdown: stop accepting new connections, finish in-flight requests, and in Kubernetes add a short preStop delay so the endpoint is removed from the load balancer before the process exits.
Cloudflare 502 versus an origin 502
If your site sits behind Cloudflare, look at who rendered the error page. A Cloudflare-branded page that shows your browser, Cloudflare and your host in a row, with the problem marked on the host, means Cloudflare could not get a valid answer from your origin. A plain, unbranded 502 page (often just "502 Bad Gateway" and "nginx") was generated by your origin's own proxy and passed through Cloudflare unchanged.
Cloudflare also has its own 52x codes that are more specific than 502: 520 (the origin returned an empty or unexpected response), 521 (the origin refused the connection), 522 (the connection to the origin timed out), 523 (the origin is unreachable), 524 (the origin accepted the connection but did not respond in time), 525 (the TLS handshake with the origin failed) and 526 (the origin certificate is invalid). If you see one of these, the problem is between Cloudflare and your origin, and the code tells you which part.
How to diagnose a 502, step by step
1. Confirm the scope
Is every request failing, or some? One URL or all of them? From everywhere, or one region? Request the page from several locations at once with the multi-region website checker. A 502 on every request points at a dead upstream or a bad config. An intermittent 502 points at keep-alive mismatches, a single bad instance behind a load balancer, or workers being killed under load.
2. Look at the response itself
Run curl -sv https://example.com/ -o /dev/null and read the response headers. The Server header and any vendor headers (cf-ray for Cloudflare, x-amz-cf-id for CloudFront, x-vercel-id for Vercel) tell you which layer produced the response. A cf-ray header with a Cloudflare-branded body means the edge generated it; a Server: nginx header with your own error body means your origin proxy did.
3. Bypass the CDN
Send the request straight to your origin while keeping the right host name and TLS server name: curl -sv --resolve example.com:443:203.0.113.10 https://example.com/ (replace the IP with your origin's). If the origin answers 200 directly, the problem is between the CDN and the origin: firewall rules that block the CDN's IP ranges, an SSL mode mismatch, or an origin certificate the CDN rejects. If the origin also returns 502, keep going inward.
4. Read the proxy's error log
This is where the answer usually is. On a typical nginx install, run tail -n 100 /var/log/nginx/error.log. The message names the exact failure:
- connect() failed (111: Connection refused) while connecting to upstreamNothing is listening at the upstream address. The application is down, or on a different port.
- connect() to unix:/run/php/... failed (2: No such file or directory)The socket path is wrong or PHP-FPM is not running.
- connect() to unix:... failed (13: Permission denied)The socket exists but nginx's user cannot use it. Fix listen.owner, listen.group and listen.mode in the PHP-FPM pool config.
- upstream prematurely closed connection while reading response headerThe upstream accepted the request and then hung up: a crash, a killed worker or an app-server timeout.
- recv() failed (104: Connection reset by peer)The upstream reset the connection, often a keep-alive race or a crash mid-request.
- upstream sent too big header while reading response headerResponse headers exceed the proxy buffer. Raise proxy_buffer_size or fastcgi_buffer_size.
- no live upstreams while connecting to upstreamnginx has marked every server in the upstream block as failed. Look at why each one failed first.
On an AWS Application Load Balancer, enable access logs and look at the target_status_code and error_reason fields; a 502 with a dash for the target status means the target never returned a valid response. Run nginx -t after any config change: a syntax error that blocks a reload can leave an old, wrong config in place.
5. Test the upstream directly
From the proxy host, call the application without the proxy: curl -sv http://127.0.0.1:3000/health. Then check that something is listening with ss -ltnp | grep 3000, and check the service itself with systemctl status yourapp and journalctl -u yourapp -n 200 (or docker ps -a and docker logs for containers, kubectl get pods and kubectl describe pod for Kubernetes, where a restart count that keeps climbing or a last state of OOMKilled is the giveaway).
6. Check memory and recent changes
Run dmesg -T | grep -i "killed process" to see whether the kernel's OOM killer took the process down. Then ask what changed: a deploy, a config edit, a package upgrade, a certificate rotation, a DNS change. Most 502s that appear suddenly on a stable system follow a change within the previous hour.
How to fix a 502 Bad Gateway
- Upstream downRestart it to restore service, then find out why it stopped. Make sure the service manager restarts it on failure (Restart=on-failure in systemd, a restart policy in Docker, liveness probes in Kubernetes) so the next crash costs seconds, not an hour.
- Workers killed by the app serverFind the slow request and make it faster, or move the work to a background job. Raising gunicorn's --timeout or PHP-FPM's request_terminate_timeout only hides the problem until the next slower request.
- Out of memoryReduce the number of workers, fix the leak, or give the host more memory. Set container memory limits deliberately rather than discovering them.
- Keep-alive mismatchMake the upstream keep-alive timeout longer than the load balancer or proxy idle timeout.
- Proxy misconfigurationCorrect the upstream address, socket path or permissions, raise header buffers if needed, run nginx -t, then reload with systemctl reload nginx.
- CDN to originAllow the CDN's published IP ranges in the origin firewall, make sure the origin serves a certificate the CDN accepts, and check that the CDN connects on the port and protocol the origin expects.
- DeploysAdd graceful shutdown and connection draining so instances leave the load balancer before they stop.
Proxies can retry a failed request on another upstream (proxy_next_upstream in nginx). That is useful for idempotent GET requests when one instance is bad, but retrying POST requests can charge a card twice. Fix the cause and keep retries for the requests that are safe to repeat.
How to monitor and alert on 502 errors
External monitoring catches a 502 the moment the gateway starts returning it, from outside your infrastructure, which matters because a crashed application cannot report its own crash.
- Monitor the pages and API routes users depend on, not only the home page. A 502 often hits one upstream (the API, the checkout service) while a cached home page keeps loading.
- Check through the CDN, as users do, and consider a second monitor on the origin directly so you can tell an edge problem from an origin problem at a glance.
- Confirm failures from more than one location before alerting, so a network blip near one checker does not page anyone, while a real 502 that every location sees does.
- Alert on 5xx responses explicitly. In SutramX, an HTTP monitor that does not get one of its expected status codes fails the check, and a 5xx is reported as an HTTP 5xx server error in the incident, with the status each region received.
On SutramX, every monitor is checked on its own interval (fastest by plan: 15 seconds on Pro, 30 seconds on Growth, 60 seconds on Starter and 3 minutes on Free), and an incident opens only when enough of its regions agree. Each incident explains itself: which regions saw the failure, what status they got, and whether an alert went out.
502 Bad Gateway FAQ
Is a 502 error my fault or the website's?
Almost always the website's. A 502 is generated between servers, after your request has already reached the site. Reloading after a minute sometimes works because a crashed process has restarted. Clearing your cache rarely helps, unless a stale cached error page is being served.
What is the difference between 502 and 504?
A 502 means the gateway got an invalid response or no response at all from the upstream: refused, reset or malformed. A 504 means the gateway connected and waited, but the upstream did not answer before the gateway's timeout. 502 suggests something is down; 504 suggests something is slow.