How to Diagnose Memory Pressure on Linux
A Linux server is under memory pressure when processes are actually stalling to get memory, or the kernel is swapping heavily or killing processes, not when the “free” column looks low. To tell the difference, check four things in order: the available figure in free, the swap-in and swap-out columns in vmstat, the Pressure Stall Information (PSI) files in /proc/pressure, and the kernel log for out-of-memory kills. This guide shows each check with real output from a 1 GB production server, and what each result means.
It pairs with the guides on optimizing WordPress on a 1 GB server, sizing PHP-FPM workers and tuning MySQL and MariaDB: those articles tell you how to change memory use, and this one tells you whether you have a problem at all. The examples use Ubuntu 24.04.
Why low free memory is normal on Linux
Linux uses spare RAM as a cache for files it has read recently. That cache is not wasted: it makes the server faster, and the kernel gives it back the moment a program needs the memory. A healthy server that has been running for a while therefore shows almost no truly free memory. The number that matters is available, the kernel’s estimate of how much memory could be given to new programs without swapping.
Do not “free” memory by dropping caches. Commands that write to /proc/sys/vm/drop_caches throw away the cache you just benefited from, and the server is slower until it warms up again. They fix nothing.
Step 1: Read free and meminfo correctly
Start with the human-readable summary and the raw counters behind it:
free -h
grep -E 'MemTotal|MemAvailable|SwapTotal|SwapFree' /proc/meminfo
Here is what they printed on a 1 GB server that runs a Laravel application and a WordPress blog:
total used free shared buff/cache available
Mem: 911Mi 605Mi 100Mi 67Mi 443Mi 305Mi
Swap: 2.0Gi 130Mi 1.9Gi
MemTotal: 933332 kB
MemAvailable: 313200 kB
SwapTotal: 2097144 kB
SwapFree: 1963120 kB
Only 100 Mi is free, which looks alarming, but 443 Mi is cache and 305 Mi is available, about a third of all RAM. As a rough guide, available memory that stays below about 10 percent of the total on a small server deserves attention. That is a starting point for investigation, not a law, because what matters is whether anything is actually stalling, which the next steps check.
The shared column (67 Mi here) covers memory used by tmpfs filesystems and shared memory segments, such as a PHP OPcache. It is normally small.
Step 2: Watch swap activity with vmstat
Swap in use is not the same as swapping. The kernel moves pages that nothing has touched for a while out to swap so the RAM can serve as cache, and those pages can sit there indefinitely. What indicates trouble is continuous traffic in and out of swap. Sample it with vmstat, one line per second:
vmstat 1 5
procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------
r b swpd free buff cache si so bi bo in cs us sy id wa st gu
0 0 134024 103352 44076 410328 0 0 0 0 369 554 0 1 99 0 0 0
0 0 134024 103352 44076 410328 0 0 0 0 262 497 1 1 99 0 0 0
Ignore the first line of real output, because it shows averages since boot. The columns to watch are:
| Column | Meaning | Worry when |
|---|---|---|
si / so |
Swap-in and swap-out, in KB per second | Consistently nonzero under normal load |
b |
Processes blocked, usually waiting for disk | Persistently above zero |
wa |
CPU time waiting for I/O | Persistently high, often together with swapping |
st |
CPU time stolen by the hypervisor | Consistently above a few percent on a virtual server |
On this server swpd is 134,024 KB (about 131 MiB) but si and so are zero on every line, and the CPU is 99 to 100 percent idle. Those pages are simply parked. At an earlier snapshot the same server had about 71 MiB in swap, so the figure creeps up over time without any performance effect. Remember that this is a quiet moment, so repeat the sample while the site is busy.
The swappiness setting controls how eagerly the kernel swaps. This server uses a low value:
cat /proc/sys/vm/swappiness
10
Step 3: Read Pressure Stall Information (PSI)
PSI is the most direct measurement of memory pressure. Instead of inferring trouble from counters, the kernel records how much time processes actually spent waiting. Read all three resources:
cat /proc/pressure/memory /proc/pressure/cpu /proc/pressure/io
Each file has a some line and a full line:
someis the share of time at least one task was stalled waiting for the resource.fullis the share of time all non-idle tasks were stalled at once, meaning the machine did no useful work.avg10,avg60andavg300are percentages of wall-clock time over the last 10, 60 and 300 seconds.totalis the cumulative stall time in microseconds since boot.
This is the memory file from the same server:
some avg10=0.00 avg60=0.00 avg300=0.00 total=224193761
full avg10=0.00 avg60=0.00 avg300=0.00 total=214695691
All averages are zero, so nothing is stalling for memory right now. The totals are how you find out whether it ever did. Converted from microseconds:
| Resource | some total since boot |
full total since boot |
|---|---|---|
| Memory | 224 s (3.7 min) | 215 s (3.6 min) |
| CPU | 28,950 s (8.0 h) | 0 |
| I/O | 4,083 s (68 min) | 3,896 s (65 min) |
To judge what a total means, compare it with the uptime, which is the first number in /proc/uptime, in seconds. For example, three and a half minutes of stalls over a month of uptime would be a tiny fraction of one percent. Two lessons from this table:
- PSI tells you that stalls happened, not why. Here the memory total is small, but the I/O stall time is about eighteen times larger. Historically, this server has waited on its disk far more than it has waited on memory.
- The I/O and memory files are linked. If a server is swapping, the disk traffic shows up as I/O pressure. I/O pressure without any swap activity, as on this server at this moment, points at something else reading or writing the disk, such as backups, log writes or the database.
Tip. Watch pressure live while you load-test or run a heavy job: watch -n 2 cat /proc/pressure/memory. A memory some average that climbs above a few percent while si and so are active is real memory pressure.
If /proc/pressure does not exist, the kernel may have PSI disabled. Adding psi=1 to the kernel boot parameters enables it on kernels that support it (Linux 4.20 or newer).
Step 4: Find out who is using the memory
If the checks above show pressure, find the consumers. List the largest processes by resident memory:
ps -eo pid,rss,comm --sort=-rss | head -10
RSS counts shared pages in full for every process, so it overstates programs that share memory, PHP-FPM workers in particular. For a fairer per-process figure, read the proportional set size (PSS):
for p in $(pgrep -f 'php-fpm: pool www'); do echo -n "$p "; sudo grep '^Pss:' /proc/$p/smaps_rollup; done
On this server the database process used about 188 MB of resident memory, and each PHP-FPM worker used 31 to 35 MB by PSS, although RSS suggested 92 to 98 MB. The PHP-FPM guide explains why that difference matters when you set a worker limit.
Step 5: Look for out-of-memory kills
When the kernel cannot find memory, it kills a process, and the log records it. Search the kernel messages:
sudo journalctl -k --since "30 days ago" | grep -i "out of memory"
sudo dmesg -T | grep -i -E "out of memory|killed process"
A kill looks like Out of memory: Killed process 1234 (mysqld). On the server measured here, the first command found no matches over 30 days. Two caveats: kernel messages from before the last reboot are lost unless the journal is persistent (check with journalctl --list-boots), and a service that restarts automatically can hide a kill, so also search the service logs, for example sudo journalctl -u mysql | grep -i killed.
Step 6: Decide what the results mean
| What you see | What it means | What to do |
|---|---|---|
Low free, healthy available, no swap activity |
Normal caching | Nothing |
Swap in use, si and so at zero |
Idle pages parked in swap | Nothing; keep an eye on it |
si/so consistently active, memory PSI rising |
Real memory pressure | Lower PHP-FPM workers and database buffers, or add RAM |
| OOM kill lines in the kernel log | Severe pressure at that moment | Identify the victim and the trigger, limit concurrency, then add RAM |
| High I/O PSI with swap activity | Thrashing | Cut memory use first |
| High I/O PSI without swap activity | The disk is busy for another reason | Look at backups, logs and database writes |
| High CPU PSI with steal above a few percent | The hypervisor is limiting your CPU | Check the instance type and its CPU credit or burst limits |
Step 7: Keep a history so you can look back
PSI totals and vmstat show the present. To see what happened at 3 a.m., record metrics over time. The sysstat package is light and collects on a schedule:
sudo apt install sysstat
sudo sed -i 's/ENABLED="false"/ENABLED="true"/' /etc/default/sysstat
sudo systemctl enable --now sysstat
sudo systemctl restart sysstat
After it has collected for a while, use sar -r for memory, sar -W for swapping and sar -B for paging. Pair this with an external uptime monitor, so you find out about downtime from an alert instead of from a visitor.
When measuring says you need more memory
If available memory stays low, swap traffic is continuous under ordinary load, PSI memory averages keep rising, or the kernel log shows OOM kills, the workload no longer fits the machine. Tuning helps up to a point, and after that the answer is more RAM, moving the database to its own server, or moving to managed WordPress hosting. For sustained high traffic, a dedicated server removes the shared-resource ceiling entirely.
Frequently asked questions
Is it bad if my Linux server uses swap?
No. Swap in use just means the kernel moved some idle pages out of RAM. Constant swapping (continuous si and so activity) is the problem.
Should I disable swap on a small server?
No. Swap is a safety net that turns a memory spike into a brief slowdown instead of an OOM kill. Lowering vm.swappiness makes the kernel more reluctant to use it without removing the safety net.
What does “full” mean in PSI?
It is the share of time when every non-idle task was stalled at the same time, so the machine was doing no useful work. Even small sustained full values on memory are a serious sign.
Which single number should I watch?
The memory PSI some avg60 value. Near zero means healthy; a value that keeps climbing while swap traffic is active means the server needs less memory use or more RAM.
The short version: ignore the free column, read available, look for continuous swapping in vmstat, confirm with PSI and the kernel log, and only then decide whether to tune or add RAM.