Guide · 10 min read

504 Gateway Timeout: Why Requests Time Out and How to Fix It

Every request to your site passes through a chain of timeouts, and a 504 means one of them ran out. The fix is rarely to raise that timeout. It is to find what was slow, and to make the chain fail in the right order.

Last updated

In short

A 504 Gateway Timeout means a proxy, load balancer or CDN connected to the server behind it but gave up waiting for a response. The upstream is usually reachable but too slow: a slow database query, a hung call to another API, or work that takes longer than the shortest timeout in the chain, such as nginx's 60-second proxy_read_timeout.

What does 504 Gateway Timeout mean?

HTTP 504 means a server acting as a gateway or proxy did not receive a timely response from the upstream server it needed to complete the request. Unlike a 502, the connection usually worked. The proxy sent the request and waited, and the upstream did not send its response headers before the proxy's timeout expired.

That has an uncomfortable consequence: the upstream may still be working. The database query may finish, the order may be placed, the email may be sent, all after the client has been told the request failed. A user who retries a timed-out checkout can end up with two orders. This is why 504s deserve more care than "raise the timeout".

The timeout chain: where 504s come from

A typical request crosses several hops, and each one has its own timeout. The shortest timeout in the chain is the one users hit. These are the defaults worth knowing:

  • nginxproxy_read_timeout defaults to 60 seconds, and applies between two successive reads from the upstream, not to the whole response. proxy_connect_timeout defaults to 60 seconds and cannot usefully exceed 75. For PHP-FPM the equivalent is fastcgi_read_timeout, for uWSGI uwsgi_read_timeout. The error log says "upstream timed out (110: Connection timed out) while reading response header from upstream".
  • AWS Application Load BalancerThe idle timeout defaults to 60 seconds and can be set from 1 to 4,000 seconds. If the target sends no data for longer than that, the ALB returns 504.
  • Amazon CloudFrontThe origin response timeout defaults to 30 seconds. When the origin does not answer in time, CloudFront returns 504.
  • Amazon API GatewayIntegrations time out at 29 seconds by default for REST APIs and at most 30 seconds for HTTP APIs, returning 504 with "Endpoint request timed out".
  • CloudflareOn Free, Pro and Business plans, Cloudflare waits 100 seconds for the origin to start responding. When that runs out it shows its own error 524, "A timeout occurred", rather than 504. Enterprise plans can raise the limit. A Cloudflare-branded 504 is different: it usually points at a problem inside Cloudflare or a 504 passed through from your origin.
  • Serverless platformsFunctions have a maximum duration. When a function runs past it, the platform's gateway returns 504; on Vercel the error is FUNCTION_INVOCATION_TIMEOUT.
502 or 504 depends on who gives up first

If the application server kills the slow request before the proxy's timeout (gunicorn's 30-second worker timeout, PHP-FPM's request_terminate_timeout), the proxy sees a closed connection and returns 502. If the proxy's timeout runs out first, it returns 504. The same slow query can produce either code, depending on which timeout is shorter.

What causes a 504 Gateway Timeout?

Slow database queries

The most common root cause. A query that took 50 ms on a small table takes 70 seconds once the table has grown and the query has no index to use. Lock contention does the same: a migration or a long transaction holds a lock, and every request that needs the row waits behind it. Connection pool exhaustion looks similar from outside: requests wait for a free connection instead of for the query.

Slow or hung calls to other services

A payment provider, a search cluster, an internal microservice or an AI API that suddenly takes 90 seconds will hold your request for 90 seconds. Many HTTP clients have no timeout by default, so a hung dependency holds the request until the proxy gives up. The 504 is your proxy reporting someone else's outage.

Long-running work inside a request

Report generation, CSV exports, image processing, bulk imports and large uploads that run synchronously can legitimately take minutes. Such a request will time out somewhere in the chain however fast your code is.

Saturation

When every worker is busy, new requests queue. Time in the queue counts against the proxy's timeout, so a server under load can produce 504s for requests that would take 100 ms if they were picked up immediately.

Network and firewall problems

A firewall or security group that silently drops packets, instead of rejecting them, makes connections hang until a timeout fires. This is common when an origin's firewall does not allow a CDN's or load balancer's address range, or after a network change between the proxy and the application.

How to diagnose a 504, step by step

1. Identify which layer timed out

Read the response headers and body with curl -sv https://example.com/slow-page -o /dev/null. A Cloudflare 524, a CloudFront error with an x-amz-cf-id header, an API Gateway "Endpoint request timed out" and nginx's own "504 Gateway Time-out" page each name a different hop. Then note how long the request took before failing: a failure at almost exactly 30, 60 or 100 seconds tells you which default timeout fired.

2. Measure where the time goes

curl breaks a request into phases. Run curl -o /dev/null -s -w "dns %{time_namelookup} connect %{time_connect} tls %{time_appconnect} first byte %{time_starttransfer} total %{time_total}\n" https://example.com/slow-page. If DNS, connect and TLS are fast and time to first byte is huge, the server is slow to produce a response. If connect itself hangs, the problem is the network or a firewall.

3. Log upstream timings in the proxy

Add timing variables to your nginx access log format: $request_time (total time nginx spent), $upstream_connect_time, $upstream_header_time (time until the upstream sent headers) and $upstream_response_time. A log_format line such as log_format timed '$remote_addr "$request" $status rt=$request_time uct=$upstream_connect_time uht=$upstream_header_time urt=$upstream_response_time'; and an access_log ... timed; directive turns every request into a measurement. Sort by uht to find the slow endpoints.

4. Check the database

On PostgreSQL, SELECT pid, now() - query_start AS runtime, state, wait_event_type, query FROM pg_stat_activity WHERE state <> 'idle' ORDER BY runtime DESC; shows what is running right now and what it is waiting on. On MySQL, SHOW FULL PROCESSLIST; does the same, and the slow query log (slow_query_log with long_query_time) records the offenders. Run EXPLAIN (ANALYZE in PostgreSQL) on the worst query to see whether it uses an index.

5. Check outbound calls

Use your application's tracing or logs to see how long each outbound call took. If one dependency accounts for most of the request time, check that provider's status page. If you have no tracing, time a call to the dependency directly from the server with the same curl timing command.

6. Check saturation and recent changes

Look at worker utilisation, CPU and queue length at the time of the timeouts. Then ask what changed: a deploy that added a query, a data migration, a traffic campaign, a dependency upgrade.

How to fix a 504 Gateway Timeout

  • Make the slow thing fastAdd the missing index, rewrite the query, paginate large results, cache expensive responses. This is the only fix that makes the problem go away instead of moving it.
  • Move long work out of the requestAccept the job, return 202 Accepted with a job ID, process it in a background worker, and let the client poll or receive a webhook when it is done. Exports, reports and imports belong here.
  • Set timeouts on every outbound callGive each HTTP client and database driver an explicit timeout shorter than your own request budget, and set a database statement timeout (statement_timeout in PostgreSQL, max_execution_time for MySQL SELECTs). A dependency that hangs should fail fast with an error you control.
  • Order the timeoutsEach inner layer should time out before the layer outside it: database statement timeout shorter than the app server timeout, shorter than the proxy timeout, shorter than the load balancer and CDN timeouts. Then the layer that knows the most, your application, is the one that reports the failure, with a useful error and a log line.
  • Raise a timeout only on purposeIf an endpoint genuinely needs longer, raise the timeout for that location only (proxy_read_timeout 300s; inside a single location block), and raise every layer in the chain that it passes through. Raising one layer just moves the 504 to the next one.
  • Add capacity for queueingIf requests are slow because they wait for a worker, add workers or instances, or reduce how long each request holds one.
  • Fix silent dropsMake sure firewalls and security groups allow the proxy's and CDN's addresses, and reject rather than drop where you control the rule, so failures are fast and visible.

How to monitor for 504s and slow responses

A 504 is the last stage of a slowdown that usually started much earlier. Response times creep up, then the slowest requests start crossing the timeout, then most of them do. Monitoring only for errors finds the problem at the last stage; monitoring latency finds it at the first.

  • Set a response-time threshold well below your shortest proxy timeout. If nginx gives up at 60 seconds, a check that fails at 5 or 10 seconds tells you about the problem while users are still getting responses.
  • Use two levels: a lower one that marks the endpoint as degraded without paging anyone, and a higher one that counts as a failure. SutramX HTTP monitors have both, Degraded above (ms) and Max response time (ms).
  • Set the monitor's own timeout deliberately. A SutramX check times out at the value you choose, from 1 to 60 seconds (never more than the interval minus 3 seconds), and records whether it timed out connecting, reading or overall.
  • Monitor the slow endpoints specifically: search, reports, checkout, anything with a heavy query behind it. The home page is often cached and fast while the rest of the site times out.
  • Watch the trend, not only the alert. Base thresholds on your measured p95 response time with headroom; the guide on reducing false alerts covers how much.
Timeouts from far away

Response time includes the network between the checker and your server. A threshold that is comfortable from a nearby region can be tight from another continent. SutramX records each region's own result and timings, and when a monitor uses more than one region, an incident opens only when enough of them agree. Free monitors check from 1 region, Starter from 3, Growth from 3 and Pro from all 3.

504 Gateway Timeout FAQ

Should I just increase the timeout?

Only for endpoints that genuinely need it, and only after checking why they are slow. Raising timeouts globally hides slow queries, ties up workers for longer and makes overload worse. For long jobs, a background queue is almost always better than a longer timeout.

Is a 504 error caused by my internet connection?

No. A 504 is generated between servers, after your request reached the site. A slow connection on your side causes the browser to time out with its own error, not a 504 page.

Why does Cloudflare show 524 instead of 504?

Cloudflare uses 524 when it connected to your origin but the origin did not start responding within its timeout, 100 seconds on Free, Pro and Business plans. It is Cloudflare's more specific version of a gateway timeout, and the fix is the same: find what is slow on the origin.

Next steps

Keep reading

Know it’s down before your customers do.

Start free — Free plan forever, no card required. Upgrade any time.