Infrastructure · 07 Oct 2026 · 3 min read
Running an LLM on a 4-Core ARM Server
There is a specific kind of dishonesty in most local-LLM writing. Someone rents an A100 for an afternoon, runs a 70B model at 40 tokens per second, and concludes that self-hosting is solved. Meanwhile the machine most people actually own sits idle because nobody benchmarks it.
So I benchmarked mine. A four-core ARM server, 23GB of RAM, no GPU. The kind of thing you can run for free on a cloud free tier or leave humming in a cupboard.
The numbers, first
A 3B model at Q4 quantisation, four threads:
completion: 3 tokens in 1.18s → 2.54 tokens/second
That is slow. I want to be clear about that before anything else, because the rest of this only makes sense if you accept it. At 2.54 tokens/second, a 200-token answer takes about eighty seconds.
What that rules out, and what it doesn't
It rules out conversation. If you want to ask follow-up questions and wait for replies, a 3B model on four cores will make you hate it within an hour.
It does not rule out automation. And the distinction matters more than the raw number:
- Short outputs. A classification, a yes/no decision, a chosen action from a fixed list. Ten tokens, not two hundred. Four seconds.
- Latency tolerance. A job that runs at 3am does not care whether it took four seconds or forty.
- Background work. If the model is reasoning about a log file while you do something else, throughput is irrelevant.
The moment you stop treating it as a chatbot and start treating it as a slow function that returns a decision, the maths changes completely.
The 1.5B surprise
The most useful thing I learned: dropping from 3B to 1.5B roughly doubled throughput, and for a large class of tasks the quality difference was invisible.
If the task is "is this log line an error or noise" or "which of these six actions fits," a 1.5B model gets it right. If the task is "write a paragraph explaining this failure to a human," it doesn't. Knowing which of those you're asking for is most of the engineering.
My setup now uses the small model by default and escalates to the larger one only when the output needs to be read by a person.
Quantisation is not optional
Q4_K_M is the sweet spot. It roughly halves the memory footprint against Q8 with a quality drop that is genuinely hard to detect on task-shaped work. On a machine with 23GB total, that difference is the difference between a model that stays resident and one that gets swapped out.
The part nobody mentions: load time
The model takes about six seconds to become ready after the server starts. If you fire a request before then, you get a 503 and a vague "Loading model" error.
This is a small thing that wastes a large amount of time if you don't know it. Health endpoints that return 200 before the model is usable are worse than no health endpoint at all — poll for an actual readiness signal, not just a listening socket.
What this is actually good for
Concretely, on this hardware, I run:
- A watch process that reads service logs and decides whether something needs restarting
- A classifier that routes incoming text to the right handler
- An embedding model that indexes conversations for later search
None of these would benefit from a smarter model. All of them benefit from one that costs nothing to run and never rate-limits.
The uncomfortable conclusion is that most of what people want an LLM for locally is not intelligence. It is judgment applied at low stakes, in volume, for free. That turns out to be exactly what a small model on modest hardware is good at.