Building a local AI coding assistant is one thing. Making it fast enough for everyday development is another.
In Part 2, we focus on the practical techniques that can make a local AI coding assistant faster and more responsive—from CPU and context tuning to smarter code search and GPU offloading.
⚡ “It Works” Is Not the Same as “It Feels Fast”
The first response from a local model is exciting. The tenth slow response is not.
That is when the next question appears: How do we calculate and improve tokens per second?
1. Measure Before Changing Anything
A simple starting metric is generation throughput:
tokens / seconds = tokens per second
Example
If llama.cpp reports 120 generated tokens in 12 seconds, generation speed is 10 tokens/sec. If another setup produces 120 tokens in 8 seconds, it is running at 15 tokens/sec.
But agent performance is broader than generation speed. Prompt processing, context size, memory pressure, search time, and repeated iterations all matter.
| Parameter | What It Changes |
|---|---|
-t |
CPU thread count |
-c |
Context window |
-ngl |
GPU layer offloading |
| Q4_K_M | Model memory/quality trade-off |
2. CPU Threads: Don’t Guess
Try a small sequence such as:
-t 4 -t 6 -t 8
Benchmark the same workload. More threads do not automatically mean proportionally more tokens per second.
Example Benchmark — Copy and Test
Run the same model with different thread counts and compare the reported generation speed. Stop the server before starting the next test.
4 threads
./build/bin/llama-server \ -m /models/Qwen3-8B-Q4_K_M.gguf \ -c 4096 \ -t 4 \ -ngl 0 \ --host 127.0.0.1 \ --port 1111
6 threads
./build/bin/llama-server \ -m /models/Qwen3-8B-Q4_K_M.gguf \ -c 4096 \ -t 6 \ -ngl 0 \ --host 127.0.0.1 \ --port 1111
8 threads
./build/bin/llama-server \ -m /models/Qwen3-8B-Q4_K_M.gguf \ -c 4096 \ -t 8 \ -ngl 0 \ --host 127.0.0.1 \ --port 1111
Then send the same prompt for every run so that the comparison is meaningful:
curl http://127.0.0.1:1111/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen3-8B-Q4_K_M.gguf",
"messages": [
{
"role": "user",
"content": "Write a Python function that checks whether a string is a palindrome. Include type hints and three unit tests."
}
],
"temperature": 0,
"max_tokens": 256
}'
For example, if 4, 6, and 8 threads produce 8.5, 10.2, and 9.8 tokens/sec respectively, use -t 6. More threads are not automatically faster.
💡 The Interesting Part
We started trying to make the model faster. The bigger discovery was that we could make the agent do less work.
3. Context: Bigger Isn’t Automatically Better
A useful starting profile is:
| Machine | Starting Context | Starting CPU Threads |
|---|---|---|
| 16 GB | 4096 | 4–6 |
| 32 GB | 8192 | 6–8 |
Increase context only when the coding task actually needs it. Large context consumes memory and increases processing work.
Example
A small bug fix touching two files may work comfortably with a 4K context. A cross-module refactor that needs several interfaces, implementations, and tests may justify 8K. Increasing context for every task simply makes the model process more text than necessary.
4. The Biggest Speed Hack Wasn’t an LLM Parameter
A repository contains plenty of work that does not require intelligence.
rg "UserService" . rg "TODO|FIXME" src/
Let deterministic tools search. Let the model reason about the results.
Example — Search Before Asking the Model
Suppose you need every reference to UserService. From the repository root, run:
rg -n --hidden \ --glob '!node_modules/**' \ --glob '!.git/**' \ "UserService" .
Use the resulting file list to decide which files Aider actually needs. The model then spends its context on understanding and changing relevant code rather than searching the whole repository.
✅ The Division of Labor
- Search tools: find things.
- LLM: reason and edit.
- Compiler: compile.
- Tests: validate.
- Git: track changes.
5. MCP: Resist the Urge
One of the questions in the journey was whether multiple MCP servers were necessary for faster filesystem access and code changes.
For this Aider-only architecture, the answer was: not initially.
Start with the smallest useful stack:
Aider + llama-server + Git + ripgrep + shell
More tools are not automatically more capability. They also introduce more tool descriptions, choices, and context.
6. GPU Changes the Equation
If the machine has a supported GPU, -ngl becomes important. But there is no universal value that works for every GPU.
GPU VRAM ↓ Model size ↓ Quantization ↓ GPU offloading ↓ Benchmark
⚠️ Don’t Copy GPU Settings Blindly
An 8 GB GPU and a 24 GB GPU should not receive the same offloading configuration. Benchmark against the actual hardware.
Example
On a GPU with limited VRAM, attempting to offload too many layers can cause memory pressure or failure. A safer process is to start conservatively, measure tokens/sec and memory usage, then increase GPU offloading until performance stops improving reliably.
7. The Real Optimization
Before: search → huge context → LLM → edit → repeat After: fast search → relevant context → LLM → focused edit → test
🎯 Part 2 Outcome
The fastest local coding agent isn’t necessarily the one with the fastest model. It is the one that makes the model do the least unnecessary work.


