Why the average lies
Consider a hundred requests where ninety-five complete in 120ms and five take 3 seconds. The mean is around 264ms, which looks acceptable on a dashboard. But one in twenty of your users just waited three seconds, and they are the ones who will complain, abandon a cart, or churn.
Percentiles preserve that information. p95 answers a question the average cannot: how bad is it for the unlucky ones?
How the calculator works
Paste values one per line, or separated by commas, spaces or semicolons. The calculator drops anything that is not a non-negative number, sorts the rest and reports p50, p75, p90, p95 and p99 alongside the count, mean, minimum and maximum. For each percentile it finds the rank p/100 × (n − 1) in the sorted list and, when that rank falls between two samples, interpolates between them. Everything runs in your browser.
A worked example
The sample that loads with the page has 16 response times. Thirteen are between 119 and 142ms, and three are slow: 480, 1,850 and 2,100ms. The results:
- p50: 130msHalfway between the eighth and ninth values (129 and 131). This is what a typical request feels like.
- p75: about 138msStill inside the fast group.
- p90: 1,165msThe rank now falls between 480 and 1,850, so the slow requests take over.
- p95: 1,912.5ms and p99: 2,062.5msBoth sit between the two slowest values. With only 16 samples they are close to the maximum.
- Mean: about 381msA response time that not one of the 16 requests actually had.
Three slow requests out of 16 is almost one in five. The mean blends them into a number that looks merely sluggish, while p90 and above show that a real share of users waited over a second.
Which percentile to use
- p50 (median)The typical experience. Useful for spotting broad regressions that affect everyone.
- p95The usual choice for SLOs. Catches real pain without letting a handful of extreme outliers dominate your alerting.
- p99The tail. Important at scale — at a million requests a day, p99 is ten thousand slow experiences.
- maxAlmost always noise. One timeout, one garbage-collection pause, and your max is meaningless as a trend.
You cannot take the p95 of five servers and average them to get a fleet p95. That is a genuinely different number. Percentiles must be computed from the combined raw distribution, which is why aggregation strategy matters in monitoring tooling.
Setting a latency budget
Once you know your current p95, the practical next step is turning it into a threshold your monitoring enforces. Pick a number slightly above today's p95 — enough headroom that normal variance does not alarm, tight enough that a genuine regression trips it.
Then treat a breach as a failure rather than a note. A response that arrives after your client has given up did not succeed, whatever status code it carried.
Turning percentiles into monitor settings
SutramX HTTP and API monitors have two latency settings. Degraded above (ms) marks slower responses as Degraded without opening an incident, which suits a threshold a little above today's p95. Max response time (ms) fails the check, which suits the point at which your users or client libraries give up. Both are described in the HTTP monitor settings.
Over time, the Reliability page shows p95 and p99 for every monitor and probe region, calculated per region so one slow location is not hidden by fast ones, and flags a monitor whose latest response is well above its 14-day baseline. See tail latency and anomalies and API monitoring. For setting latency targets as SLOs, read SLA, SLI and SLO explained and monitoring APIs effectively.
Frequently asked questions
Which percentile method does the calculator use?
Linear interpolation between the two nearest ranks, the method used by Excel's PERCENTILE.INC and NumPy's default. If you check the numbers in a spreadsheet with PERCENTILE.INC, they will match.
Why is my p50 a number that is not in my data?
With an even number of samples there is no single middle value, so the median falls between the two middle values. The same interpolation applies to every percentile whose rank falls between two samples.
How many samples do I need?
Enough that the percentile is not just your slowest one or two requests. As a rule of thumb, p95 needs at least a few dozen samples and p99 at least a hundred. With fewer, treat high percentiles as a rough indication.
Does it have to be milliseconds?
No. The calculator works on any non-negative numbers, so seconds or bytes work as long as every value uses the same unit. Values that are not numbers, or are negative, are ignored and counted below the input.
Is my data uploaded anywhere?
No. Everything is calculated in your browser. Nothing you paste leaves your device.