How far can I push a $7 server? - Installing an AI
This time, I wanted to see if I could run an LLM on my $7 server.
The server has 1 AMD vCPU, 1 GB of RAM and no GPU, so I wasn't expecting to run anything particularly big. But how small does the model actually need to be?
That's what I wanted to find out.
Constraints & experiment
I'm using the same DigitalOcean server from the previous experiments:
| Resource | Specification |
|---|---|
| Price | $7 / month |
| CPU | 1 AMD vCPU |
| RAM | 1 GB |
| Swap | 1 GB |
| Disk | 25 GB NVMe SSD |
| GPU | None |
The objective is also not to benchmark which model is the most intelligent. There are already proper benchmarks for that. What I want to measure is how these small models behave on this specific server as I move towards models with more parameters.
For each model, I'm mainly interested in:
- Peak memory used by the process.
- Minimum available RAM while the model is running.
- How much additional swap is used.
- Prompt processing speed.
- Generation speed.
- Whether the model is still practical to use.
For the models, I also want to add one restriction: their weights need to be publicly downloadable and released under the Apache 2.0 license.
Running the models
First, I needed something capable of running the models without adding too much overhead to an already very limited server.
I decided to use llama.cpp,
a lightweight runtime for running LLMs locally.
The models are stored in GGUF files, a format designed for
efficient inference and commonly used with quantized models.
Now I just need a model, but before choosing one, there is one more thing to consider: quantization.
Model weights can be stored using different levels of precision. Lowering that precision reduces the size of the model and the amount of memory required to run it, although going too far can also affect its quality.
This matters quite a lot on a server with only 1 GB of RAM. Even a 360M parameter model would require roughly 720 MB just for its weights at 16-bit precision. At 4 bits, the same number of parameters requires much less space.
For all the tests, I'll use Q4_K_M versions when available. It is a commonly used 4-bit quantization that provides a reasonable balance between model size and quality. More importantly for this experiment, using the same quantization keeps one more variable consistent while I increase the model size.
Which models are we testing and how?
I'll start with a small model and progressively increase the parameter count. All the models in the experiment have downloadable weights and are released under the Apache 2.0 license.
More parameters don't necessarily mean a better model or better performance. The goal here isn't to rank their intelligence, but to see how each one behaves under the same hardware constraints.
| Model | Parameters | License |
|---|---|---|
| SmolLM2-360M-Instruct | 360M | Apache 2.0 |
| Qwen2.5-0.5B-Instruct | 0.49B | Apache 2.0 |
| Qwen3-0.6B | 0.6B | Apache 2.0 |
| OLMo-2-0425-1B-Instruct | 1B | Apache 2.0 |
To keep the comparison as consistent as possible, I'll use Q4_K_M quantization, the same context size and the same prompt for the main tests. Qwen3 also supports a thinking mode, which I'll disable so that it behaves more like the other instruct models during the experiment.
I chose a simple database problem with a clear cause, solution and trade-off, so every model has the same short technical task to solve.
A web application becomes slow when its database grows from 10,000 to 5 million users.
The following query is frequently executed:
SELECT * FROM users WHERE email = 'user@example.com';
The email column has no index.
Explain:
1. Why the query may become slower as the table grows.
2. What change you would make to improve it.
3. One trade-off introduced by that change.
Keep your answer under 150 words.
And about measures, llama.cpp reports prompt processing and generation speed.
While each model is running, I'll also sample the process and system memory
every 500 milliseconds.
I'll track peak RSS, minimum available RAM and additional swap usage.
Together with generation speed, these numbers should show where the limits
of the server start to appear.
Additional swap means the difference between the swap already in use before starting the model and the highest swap usage observed during the test. The GGUF file size also isn't the same as runtime memory usage, so I'll measure the process and the system instead of estimating memory requirements from the file size.
Comparing the models
With all four models tested under the same conditions, the difference between them becomes much easier to see.
| Model | GGUF size | Peak RSS | Min. available RAM | Extra swap | Generation |
|---|---|---|---|---|---|
| SmolLM2 360M | ~271 MB | ~401 MiB | ~505 MiB | ~0.5 MiB | 18.8 tok/s |
| Qwen2.5 0.5B | ~398 MB | ~503 MiB | ~522 MiB | ~0.75 MiB | 15.7 tok/s |
| Qwen3 0.6B | ~484 MB | ~676 MiB | ~176 MiB | ~31.5 MiB | 19.9 tok/s |
| OLMo 2 1B | ~936 MB | ~739 MiB | ~10.8 MiB | ~451.8 MiB | 1.5 tok/s |
Memory pressure
Peak process memory and additional swap used during each test.
Generation speed
Tokens generated per second by each model.
SmolLM2 and Qwen2.5 both ran without meaningful additional swap usage or severe RAM pressure. Qwen3 at 600M is the first test where swap usage starts to increase, although generation performance remains good.
The big change happens with OLMo 2 1B. The model still runs, but available RAM falls to around 11 MiB, additional swap usage grows to more than 450 MiB, and generation speed drops to just 1.5 tokens per second.
Model size also doesn't translate directly into generation speed. Qwen3 was the fastest model in this test at 19.9 tokens per second, despite using considerably more memory than SmolLM2 and Qwen2.5.
What about the answers?
Performance is the main focus of this experiment, but tokens per second
aren't very useful if the model produces a bad answer. I didn't run a
proper quality benchmark, but the prompt gives us three simple things
to check: identify the table scan, recommend an index on
email, and mention a real trade-off of adding that index.
The answers were mixed. The smaller models understood parts of the problem but also introduced incorrect explanations. For this particular prompt, Qwen3 produced the best answer overall, although it still included a minor incorrect claim. OLMo correctly suggested the index, but its answer wasn't clearly better despite requiring far more resources.
This isn't enough to rank the models by quality, but it is enough to avoid treating generation speed as the only useful metric.
Conclusion
So, how far can I push a $7 server?
Far enough to run a 1B parameter language model on a single CPU core with around 1 GB of RAM. OLMo 2 1B loaded successfully, processed the prompt and generated an answer.
But running something and running it well are two different things. At 1B parameters, the server had almost no available RAM left, used more than 450 MiB of additional swap, and generation dropped to 1.5 tokens per second.
The smaller models were much more practical. The three models tested between 360M and 600M parameters generated around 16–20 tokens per second, and Qwen3 0.6B gave the best balance in this experiment: it was the fastest model tested, produced the best answer for this particular prompt, and still ran before memory pressure became severe.
The interesting part for me isn't that a 1B model can technically run on this server. It's finding where the useful limit is. More parameters don't automatically mean better performance or a better answer, and once the machine starts relying heavily on swap, the cost becomes very visible.
Among the models tested here, Qwen3 0.6B was the largest one that remained practical on this server. At 1B parameters, memory pressure and swap usage had already made the experience impractical.
Maybe a RAG system could help improve those answers without requiring a larger model. But would that fit on the same $7 server?