Your Agent Didn't Crash — It Was Killed. Diagnosing OOM on a VPS

Axel Grubba, September 30, 2026
Start selling digital products with Crevio
Crevio E-Commerce Platforms logo
Crevio
Sponsored
5.0
(1)
Free plan available
Crevio is an AI-powered platform that runs your business while you sleep. Describe what you want to se... Learn more about Crevio
Get an AI summary of this post on:

The signature is unmistakable once you’ve seen it: your agent is working, then it isn’t. The process is gone. The application log ends mid-sentence with no error, no stack trace, nothing. So you go looking for a bug in the agent.

There is no bug. The kernel killed it, and it didn’t tell your application because it doesn’t have to — the process was sent SIGKILL, which cannot be caught, handled or logged by the thing being killed.

The evidence lives somewhere else entirely.

Confirm it before you change anything

sudo dmesg -T | grep -i -E 'killed process|out of memory'

A line naming your process with a timestamp matching the disappearance is your answer. If nothing matches, stop — it wasn’t the OOM killer, and every fix below is wasted effort. Worth ruling out your provider too: a regional incident can be real while the status banner stays green. Go back to the application logs. If you’re not yet sure the cause is memory at all, work down the full decision tree for an agent that stopped responding — memory is step three of seven.

In Docker, the same event shows up as an exit code:

docker ps -a --format '{{.Names}}\t{{.Status}}'
docker inspect <container> --format '{{.State.OOMKilled}} {{.State.ExitCode}}'

Exit code 137 means SIGKILL — it’s the shell convention of 128 plus the signal number, and SIGKILL is 9. It doesn’t always mean OOM (anything sending SIGKILL produces it), which is why State.OOMKilled is the more precise check. true there is conclusive.

Four-step diagnostic flow: confirm with dmesg grep for killed process or out of memory and stop if there is no match; identify the victim, its RSS and the total_vm line, noting the victim is usually the largest process rather than the cause; contain by adding swap then capping the container then cutting concurrency, cheapest first; and resize only once the real peak is known, since resizing first hides the leak and is paid for monthly

Read what the kernel actually chose, and why

This is the part that changes how you fix it.

The kernel scores every process with a “badness” heuristic. From the kernel documentation: the score runs “from 0 (never kill) to 1000 (always kill),” and “a task using all allowed memory receives a badness score of 1000; using half receives 500.”

Read that again, because it’s the whole misdiagnosis. The victim is selected by how much memory it holds, not by what caused the shortage. A runaway script that allocates aggressively can push the system over the edge and the kernel will kill Ollama instead, because Ollama is sitting on 6GB of model weights and the script is holding 400MB. Your logs then show Ollama dying for no reason, and Ollama is entirely innocent.

So when you read the dmesg output, look at two things: the process that was killed, and the table of candidates above it with their RSS values. The second tells you where the memory actually went.

You can bias the choice with oom_score_adj, which the kernel adds to the badness score before picking. It ranges from -1000 to +1000, and -1000 is special: it “is equivalent to disabling oom killing entirely for that task since it will always report a badness score of 0.”

# protect a specific process from selection
sudo choom -n -1000 -p <pid>

Use that sparingly. Making a process unkillable doesn’t create memory — it just guarantees the kernel kills something else, possibly something you need more, and Docker’s own documentation warns this “can effectively bring the entire system down if the wrong process is killed.”

Fix it in order, cheapest first

The instinct is to resize the server. Do that last, because resizing before you know the real peak means paying for the leak every month instead of finding it.

1. Add swap

Most VPS images ship with no swap at all, which means there is no cushion whatsoever between “memory is tight” and “the kernel starts killing.” Swap doesn’t make you fast — it makes the failure gradual rather than sudden, which is usually the difference between a slow agent and a dead one.

sudo fallocate -l 4G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

Confirm with free -h and swapon --show. On NVMe a swapfile is far less painful than the folklore suggests, and 2–4GB on a small box buys real headroom.

One caveat: if your provider bills or throttles disk I/O aggressively, heavy swapping is visible in that bill. Swap is a safety net for spikes, not a substitute for RAM you genuinely need.

2. Cap the containers

An unbounded container will happily consume everything on the host. Docker’s --memory flag bounds it — the documented minimum is 6m — and the important consequence is that when the container exceeds its cap, the kernel kills something inside that container instead of picking a victim across the whole machine. The blast radius stops being “any process on the box.”

services:
  agent:
    mem_limit: 2g
    memswap_limit: 3g

--memory-swap (or memswap_limit) is total memory plus swap, and Docker notes it “only has meaning if --memory is also set.” Setting it higher than mem_limit is what actually permits the container to swap rather than die.

Resist --oom-kill-disable. It exists, but disabling OOM kills for a container that then grows without bound moves the problem to the host, where it’s worse.

3. Cut concurrency

Memory problems in agent workloads are usually concurrency problems wearing a disguise. Each of these multiplies:

  • Headless browser instances. A single headless Chromium wants 1–2GB to itself. Two parallel browser tasks on a 4GB box is the entire budget before the agent has done anything.
  • Default session caps you never set. Browser services ship with ceilings sized for a dedicated host — the Camofox server used by Hermes Agent defaults to 50 concurrent sessions, which permits several times the memory of the server people run it on. Check the default before you assume it’s sane for your box.
  • Parallel workflow executions. n8n and similar will run as many as you let them.
  • Loaded models. Each model held in memory is its full weight, not a fraction.

Halving concurrency is free and instant. It’s very often the actual fix.

4. Then, and only then, resize

Once you know the real peak — from the RSS values in the dmesg table, or from watching free -h during a representative task — you can size deliberately.

Published guidance for common agent workloads, as a starting point rather than a promise:

Workload Practical floor
Agent, text-only, hosted model API ~1–2GB
Agent with headless browser automation 8GB
Ollama, quantised 7–8B model 8GB
Ollama, quantised 13B model 16GB
n8n with a substantial workflow 4GB+

Those are floors for the component alone. Run two of them on one box and you add them together, plus the operating system.

Watch it before it happens again

The point of monitoring here is catching the trend, not the corpse:

# what is actually resident, largest first
ps -eo pid,comm,rss --sort=-rss | head -10

# live view including swap
free -h -s 5

If you want an alert rather than a habit, a cron job that greps dmesg for OOM lines and emails you is five minutes’ work and considerably better than discovering it days later.

When more RAM is genuinely the answer

Sometimes the diagnosis is simply that the box is too small — a 4GB server running an agent with browser automation is under-provisioned no matter how carefully you tune it.

If you’re going to pay for more memory, pay attention to what you get per pound. Hostinger’s KVM plans pair generous RAM with the easiest upgrade path, and its 8GB and 16GB tiers are the usual landing spots for agent workloads — our 16GB VPS comparison prices that tier across providers. Hetzner is the cheapest per gigabyte when its CX tier is in stock — worth checking before you commit, because that tier was showing as unavailable when we last looked, and the CPX line you’d fall back to is several times dearer. That assumes you’re comfortable with a plainer control panel, which is why it keeps winning our cheapest VPS index.

Before you buy, check the sizing guides rather than guessing twice: best VPS for Ollama covers model-size-to-RAM, and best VPS for OpenClaw covers the agent side.

How we checked this

The badness-score behaviour, the 0–1000 range and the oom_score_adj semantics — including that -1000 disables OOM killing for a task — are quoted from the Linux kernel documentation. The Docker memory-limit behaviour, the 6m minimum, the --memory-swap modifier semantics and the warning about killing the wrong process are quoted from Docker’s own resource-constraints documentation. Exit code 137 is the standard 128-plus-signal convention for SIGKILL, which is why we’ve paired it with State.OOMKilled rather than treating it as proof on its own.

What we did not do: we did not measure peak memory for these workloads ourselves. The brief for this article called for a benchmark table across idle agents, browser automation, n8n and Ollama, and the figures above are published guidance and documented model requirements rather than our measurements — which is why they’re presented as floors rather than precise numbers. Your peak depends on model quantisation, browser page weight, workflow shape and concurrency, and the only figure that matters for your box is the one in your own dmesg output.

That’s not a limitation you need to work around: the RSS table the kernel prints during an OOM kill is a better measurement of your actual workload than any benchmark we could publish, because it was taken on your machine under your load at the moment it mattered.

The two hosting links above are affiliate links. Nothing in the diagnosis or the first three fixes requires spending anything, which is deliberate — resizing is step four for a reason.

FAQ

What does exit code 137 mean in Docker?

The process received SIGKILL — 137 is 128 plus signal 9. It’s usually the OOM killer, but not always, since anything sending SIGKILL produces the same code. Confirm with docker inspect <container> --format '{{.State.OOMKilled}}'; true is conclusive.

How do I know if the OOM killer killed my process?

Run sudo dmesg -T | grep -i -E 'killed process|out of memory' and look for a line naming your process at the time it vanished. If there’s no match, it wasn’t an OOM kill and the cause is in your application logs.

Why did it kill the wrong process?

Because the kernel selects by memory footprint, not by blame. Its badness heuristic scores a task using all available memory at 1000 and half at 500, so the largest resident process is the likely victim even when a smaller one caused the shortage.

Will adding swap fix OOM kills?

It makes them much less likely by giving the kernel somewhere to go under pressure, and most VPS images ship with none. It won’t fix a genuine shortfall — if your workload needs 8GB and you have 2GB, swap converts a crash into severe slowness rather than solving it.

How do I stop one container taking down the whole server?

Set mem_limit (Docker’s --memory, minimum 6m). The container then gets a bounded allocation, and pressure inside it is resolved inside it rather than by the kernel choosing a victim anywhere on the host.

How much RAM does an AI agent actually need?

Roughly 1–2GB for a text-only agent calling a hosted model, but 8GB once headless browser automation is involved, since a single Chromium instance wants 1–2GB by itself. Self-hosting a quantised 7–8B model with Ollama is a separate 8GB on top.

Should I use --oom-kill-disable?

Almost never. It doesn’t create memory; it guarantees the kernel kills a different process instead. Docker warns that killing the wrong process can bring the whole system down, and an unkillable container that keeps growing is exactly how that happens.

Founder & Software Review Editor
Axel Grubba is the founder of Findstack, a B2B software comparison platform, with his background spanning management consulting and venture capital where he invested in software. Recently, Axel has developed a passion for coding and enjoys traveling when he is not building and improving Findstack.
Business Software Reviews SaaS Product Evaluation CRM Software
Subscribe, get software deals straight to your inbox.
Join 8,100+ other entrepreneurs staying up-to-date on all the latest deals.
Zero spam. Unsubscribe at any time.