{"id":3293,"date":"2026-09-30T10:10:17","date_gmt":"2026-09-30T10:10:17","guid":{"rendered":"https:\/\/siteharbour.com\/blog\/?p=3293"},"modified":"2026-09-30T10:10:17","modified_gmt":"2026-09-30T10:10:17","slug":"diagnose-memory-pressure-linux","status":"publish","type":"post","link":"https:\/\/siteharbour.com\/blog\/diagnose-memory-pressure-linux\/","title":{"rendered":"How to Diagnose Memory Pressure on Linux"},"content":{"rendered":"<p>A Linux server is under memory pressure when processes are actually stalling to get memory, or the kernel is swapping heavily or killing processes, not when the &#8220;free&#8221; column looks low. To tell the difference, check four things in order: the <code>available<\/code> figure in <code>free<\/code>, the swap-in and swap-out columns in <code>vmstat<\/code>, the Pressure Stall Information (PSI) files in <code>\/proc\/pressure<\/code>, and the kernel log for out-of-memory kills. This guide shows each check with real output from a 1 GB production server, and what each result means.<\/p>\n<p>It pairs with the guides on <a href=\"\/blog\/optimize-wordpress-1gb-ram-server\/\">optimizing WordPress on a 1 GB server<\/a>, <a href=\"\/blog\/optimize-php-fpm-workers-limited-ram\/\">sizing PHP-FPM workers<\/a> and <a href=\"\/blog\/optimize-mysql-mariadb-low-ram-wordpress\/\">tuning MySQL and MariaDB<\/a>: those articles tell you how to change memory use, and this one tells you whether you have a problem at all. The examples use Ubuntu 24.04.<\/p>\n<h2>Why low free memory is normal on Linux<\/h2>\n<p>Linux uses spare RAM as a cache for files it has read recently. That cache is not wasted: it makes the server faster, and the kernel gives it back the moment a program needs the memory. A healthy server that has been running for a while therefore shows almost no truly free memory. The number that matters is <code>available<\/code>, the kernel&#8217;s estimate of how much memory could be given to new programs without swapping.<\/p>\n<div class=\"callout callout-warning\">\n<p><strong>Do not &#8220;free&#8221; memory by dropping caches.<\/strong> Commands that write to <code>\/proc\/sys\/vm\/drop_caches<\/code> throw away the cache you just benefited from, and the server is slower until it warms up again. They fix nothing.<\/p>\n<\/div>\n<h2>Step 1: Read free and meminfo correctly<\/h2>\n<p>Start with the human-readable summary and the raw counters behind it:<\/p>\n<pre><code>free -h\ngrep -E 'MemTotal|MemAvailable|SwapTotal|SwapFree' \/proc\/meminfo<\/code><\/pre>\n<p>Here is what they printed on a 1 GB server that runs a Laravel application and a WordPress blog:<\/p>\n<pre><code>               total        used        free      shared  buff\/cache   available\nMem:           911Mi       605Mi       100Mi        67Mi       443Mi       305Mi\nSwap:          2.0Gi       130Mi       1.9Gi\n\nMemTotal:         933332 kB\nMemAvailable:     313200 kB\nSwapTotal:       2097144 kB\nSwapFree:        1963120 kB<\/code><\/pre>\n<p>Only 100 Mi is free, which looks alarming, but 443 Mi is cache and 305 Mi is available, about a third of all RAM. As a rough guide, available memory that stays below about 10 percent of the total on a small server deserves attention. That is a starting point for investigation, not a law, because what matters is whether anything is actually stalling, which the next steps check.<\/p>\n<p>The <code>shared<\/code> column (67 Mi here) covers memory used by tmpfs filesystems and shared memory segments, such as a PHP OPcache. It is normally small.<\/p>\n<h2>Step 2: Watch swap activity with vmstat<\/h2>\n<p>Swap in use is not the same as swapping. The kernel moves pages that nothing has touched for a while out to swap so the RAM can serve as cache, and those pages can sit there indefinitely. What indicates trouble is continuous traffic in and out of swap. Sample it with <code>vmstat<\/code>, one line per second:<\/p>\n<pre><code>vmstat 1 5<\/code><\/pre>\n<pre><code>procs -----------memory---------- ---swap-- -----io---- -system-- -------cpu-------\n r  b   swpd   free   buff  cache   si   so    bi    bo   in   cs us sy id wa st gu\n 0  0 134024 103352  44076 410328    0    0     0     0  369  554  0  1 99  0  0  0\n 0  0 134024 103352  44076 410328    0    0     0     0  262  497  1  1 99  0  0  0<\/code><\/pre>\n<p>Ignore the first line of real output, because it shows averages since boot. The columns to watch are:<\/p>\n<table>\n<thead>\n<tr>\n<th>Column<\/th>\n<th>Meaning<\/th>\n<th>Worry when<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>si<\/code> \/ <code>so<\/code><\/td>\n<td>Swap-in and swap-out, in KB per second<\/td>\n<td>Consistently nonzero under normal load<\/td>\n<\/tr>\n<tr>\n<td><code>b<\/code><\/td>\n<td>Processes blocked, usually waiting for disk<\/td>\n<td>Persistently above zero<\/td>\n<\/tr>\n<tr>\n<td><code>wa<\/code><\/td>\n<td>CPU time waiting for I\/O<\/td>\n<td>Persistently high, often together with swapping<\/td>\n<\/tr>\n<tr>\n<td><code>st<\/code><\/td>\n<td>CPU time stolen by the hypervisor<\/td>\n<td>Consistently above a few percent on a virtual server<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>On this server <code>swpd<\/code> is 134,024 KB (about 131 MiB) but <code>si<\/code> and <code>so<\/code> are zero on every line, and the CPU is 99 to 100 percent idle. Those pages are simply parked. At an earlier snapshot the same server had about 71 MiB in swap, so the figure creeps up over time without any performance effect. Remember that this is a quiet moment, so repeat the sample while the site is busy.<\/p>\n<p>The <code>swappiness<\/code> setting controls how eagerly the kernel swaps. This server uses a low value:<\/p>\n<pre><code>cat \/proc\/sys\/vm\/swappiness\n10<\/code><\/pre>\n<h2>Step 3: Read Pressure Stall Information (PSI)<\/h2>\n<p>PSI is the most direct measurement of memory pressure. Instead of inferring trouble from counters, the kernel records how much time processes actually spent waiting. Read all three resources:<\/p>\n<pre><code>cat \/proc\/pressure\/memory \/proc\/pressure\/cpu \/proc\/pressure\/io<\/code><\/pre>\n<p>Each file has a <code>some<\/code> line and a <code>full<\/code> line:<\/p>\n<ul>\n<li><code>some<\/code> is the share of time at least one task was stalled waiting for the resource.<\/li>\n<li><code>full<\/code> is the share of time all non-idle tasks were stalled at once, meaning the machine did no useful work.<\/li>\n<li><code>avg10<\/code>, <code>avg60<\/code> and <code>avg300<\/code> are percentages of wall-clock time over the last 10, 60 and 300 seconds.<\/li>\n<li><code>total<\/code> is the cumulative stall time in microseconds since boot.<\/li>\n<\/ul>\n<p>This is the memory file from the same server:<\/p>\n<pre><code>some avg10=0.00 avg60=0.00 avg300=0.00 total=224193761\nfull avg10=0.00 avg60=0.00 avg300=0.00 total=214695691<\/code><\/pre>\n<p>All averages are zero, so nothing is stalling for memory right now. The totals are how you find out whether it ever did. Converted from microseconds:<\/p>\n<table>\n<thead>\n<tr>\n<th>Resource<\/th>\n<th><code>some<\/code> total since boot<\/th>\n<th><code>full<\/code> total since boot<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Memory<\/td>\n<td>224 s (3.7 min)<\/td>\n<td>215 s (3.6 min)<\/td>\n<\/tr>\n<tr>\n<td>CPU<\/td>\n<td>28,950 s (8.0 h)<\/td>\n<td>0<\/td>\n<\/tr>\n<tr>\n<td>I\/O<\/td>\n<td>4,083 s (68 min)<\/td>\n<td>3,896 s (65 min)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>To judge what a total means, compare it with the uptime, which is the first number in <code>\/proc\/uptime<\/code>, in seconds. For example, three and a half minutes of stalls over a month of uptime would be a tiny fraction of one percent. Two lessons from this table:<\/p>\n<ul>\n<li><strong>PSI tells you that stalls happened, not why.<\/strong> Here the memory total is small, but the I\/O stall time is about eighteen times larger. Historically, this server has waited on its disk far more than it has waited on memory.<\/li>\n<li><strong>The I\/O and memory files are linked.<\/strong> If a server is swapping, the disk traffic shows up as I\/O pressure. I\/O pressure without any swap activity, as on this server at this moment, points at something else reading or writing the disk, such as backups, log writes or the database.<\/li>\n<\/ul>\n<div class=\"callout callout-tip\">\n<p><strong>Tip.<\/strong> Watch pressure live while you load-test or run a heavy job: <code>watch -n 2 cat \/proc\/pressure\/memory<\/code>. A memory <code>some<\/code> average that climbs above a few percent while <code>si<\/code> and <code>so<\/code> are active is real memory pressure.<\/p>\n<\/div>\n<p>If <code>\/proc\/pressure<\/code> does not exist, the kernel may have PSI disabled. Adding <code>psi=1<\/code> to the kernel boot parameters enables it on kernels that support it (Linux 4.20 or newer).<\/p>\n<h2>Step 4: Find out who is using the memory<\/h2>\n<p>If the checks above show pressure, find the consumers. List the largest processes by resident memory:<\/p>\n<pre><code>ps -eo pid,rss,comm --sort=-rss | head -10<\/code><\/pre>\n<p>RSS counts shared pages in full for every process, so it overstates programs that share memory, PHP-FPM workers in particular. For a fairer per-process figure, read the proportional set size (PSS):<\/p>\n<pre><code>for p in $(pgrep -f 'php-fpm: pool www'); do echo -n \"$p \"; sudo grep '^Pss:' \/proc\/$p\/smaps_rollup; done<\/code><\/pre>\n<p>On this server the database process used about 188 MB of resident memory, and each PHP-FPM worker used 31 to 35 MB by PSS, although RSS suggested 92 to 98 MB. The <a href=\"\/blog\/optimize-php-fpm-workers-limited-ram\/\">PHP-FPM guide<\/a> explains why that difference matters when you set a worker limit.<\/p>\n<h2>Step 5: Look for out-of-memory kills<\/h2>\n<p>When the kernel cannot find memory, it kills a process, and the log records it. Search the kernel messages:<\/p>\n<pre><code>sudo journalctl -k --since \"30 days ago\" | grep -i \"out of memory\"\nsudo dmesg -T | grep -i -E \"out of memory|killed process\"<\/code><\/pre>\n<p>A kill looks like <code>Out of memory: Killed process 1234 (mysqld)<\/code>. On the server measured here, the first command found no matches over 30 days. Two caveats: kernel messages from before the last reboot are lost unless the journal is persistent (check with <code>journalctl --list-boots<\/code>), and a service that restarts automatically can hide a kill, so also search the service logs, for example <code>sudo journalctl -u mysql | grep -i killed<\/code>.<\/p>\n<h2>Step 6: Decide what the results mean<\/h2>\n<table>\n<thead>\n<tr>\n<th>What you see<\/th>\n<th>What it means<\/th>\n<th>What to do<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Low <code>free<\/code>, healthy <code>available<\/code>, no swap activity<\/td>\n<td>Normal caching<\/td>\n<td>Nothing<\/td>\n<\/tr>\n<tr>\n<td>Swap in use, <code>si<\/code> and <code>so<\/code> at zero<\/td>\n<td>Idle pages parked in swap<\/td>\n<td>Nothing; keep an eye on it<\/td>\n<\/tr>\n<tr>\n<td><code>si<\/code>\/<code>so<\/code> consistently active, memory PSI rising<\/td>\n<td>Real memory pressure<\/td>\n<td>Lower PHP-FPM workers and database buffers, or add RAM<\/td>\n<\/tr>\n<tr>\n<td>OOM kill lines in the kernel log<\/td>\n<td>Severe pressure at that moment<\/td>\n<td>Identify the victim and the trigger, limit concurrency, then add RAM<\/td>\n<\/tr>\n<tr>\n<td>High I\/O PSI with swap activity<\/td>\n<td>Thrashing<\/td>\n<td>Cut memory use first<\/td>\n<\/tr>\n<tr>\n<td>High I\/O PSI without swap activity<\/td>\n<td>The disk is busy for another reason<\/td>\n<td>Look at backups, logs and database writes<\/td>\n<\/tr>\n<tr>\n<td>High CPU PSI with steal above a few percent<\/td>\n<td>The hypervisor is limiting your CPU<\/td>\n<td>Check the instance type and its CPU credit or burst limits<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Step 7: Keep a history so you can look back<\/h2>\n<p>PSI totals and <code>vmstat<\/code> show the present. To see what happened at 3 a.m., record metrics over time. The <code>sysstat<\/code> package is light and collects on a schedule:<\/p>\n<pre><code>sudo apt install sysstat\nsudo sed -i 's\/ENABLED=\"false\"\/ENABLED=\"true\"\/' \/etc\/default\/sysstat\nsudo systemctl enable --now sysstat\nsudo systemctl restart sysstat<\/code><\/pre>\n<p>After it has collected for a while, use <code>sar -r<\/code> for memory, <code>sar -W<\/code> for swapping and <code>sar -B<\/code> for paging. Pair this with an external uptime monitor, so you find out about downtime from an alert instead of from a visitor.<\/p>\n<h2>When measuring says you need more memory<\/h2>\n<p>If available memory stays low, swap traffic is continuous under ordinary load, PSI memory averages keep rising, or the kernel log shows OOM kills, the workload no longer fits the machine. Tuning helps up to a point, and after that the answer is more RAM, moving the database to its own server, or moving to <a href=\"\/wordpress\">managed WordPress hosting<\/a>. For sustained high traffic, a <a href=\"\/dedicated-server\">dedicated server<\/a> removes the shared-resource ceiling entirely.<\/p>\n<h2>Frequently asked questions<\/h2>\n<h3>Is it bad if my Linux server uses swap?<\/h3>\n<p>No. Swap in use just means the kernel moved some idle pages out of RAM. Constant swapping (continuous <code>si<\/code> and <code>so<\/code> activity) is the problem.<\/p>\n<h3>Should I disable swap on a small server?<\/h3>\n<p>No. Swap is a safety net that turns a memory spike into a brief slowdown instead of an OOM kill. Lowering <code>vm.swappiness<\/code> makes the kernel more reluctant to use it without removing the safety net.<\/p>\n<h3>What does &#8220;full&#8221; mean in PSI?<\/h3>\n<p>It is the share of time when every non-idle task was stalled at the same time, so the machine was doing no useful work. Even small sustained <code>full<\/code> values on memory are a serious sign.<\/p>\n<h3>Which single number should I watch?<\/h3>\n<p>The memory PSI <code>some avg60<\/code> value. Near zero means healthy; a value that keeps climbing while swap traffic is active means the server needs less memory use or more RAM.<\/p>\n<p>The short version: ignore the free column, read <code>available<\/code>, look for continuous swapping in <code>vmstat<\/code>, confirm with PSI and the kernel log, and only then decide whether to tune or add RAM.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>How to tell real memory pressure from normal Linux caching, with real free, vmstat and PSI output from a 1 GB production server.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[16],"tags":[31,22,32,34,24,33],"class_list":["post-3293","post","type-post","status-publish","format-standard","hentry","category-servers","tag-linux","tag-low-ram-server","tag-memory-pressure","tag-psi","tag-server-optimization","tag-vmstat"],"_links":{"self":[{"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/posts\/3293","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/comments?post=3293"}],"version-history":[{"count":1,"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/posts\/3293\/revisions"}],"predecessor-version":[{"id":3294,"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/posts\/3293\/revisions\/3294"}],"wp:attachment":[{"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/media?parent=3293"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/categories?post=3293"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/siteharbour.com\/blog\/wp-json\/wp\/v2\/tags?post=3293"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}