Building a Local AI Coding Assistant: From a 16/32 GB Machine to a Practical Coding Partner

Local AI Coding Assistant – Part 2: How to Make It Fast

Building a local AI coding assistant is one thing. Making it fast enough for everyday development is another.

In Part 2, we focus on the practical techniques that can make a local AI coding assistant faster and more responsive—from CPU and context tuning to smarter code search and GPU offloading.

⚡ “It Works” Is Not the Same as “It Feels Fast”

The first response from a local model is exciting. The tenth slow response is not.

That is when the next question appears: How do we calculate and improve tokens per second?

1. Measure Before Changing Anything

A simple starting metric is generation throughput:

tokens / seconds = tokens per second
Example

If llama.cpp reports 120 generated tokens in 12 seconds, generation speed is 10 tokens/sec. If another setup produces 120 tokens in 8 seconds, it is running at 15 tokens/sec.

But agent performance is broader than generation speed. Prompt processing, context size, memory pressure, search time, and repeated iterations all matter.

Parameter What It Changes
-t CPU thread count
-c Context window
-ngl GPU layer offloading
Q4_K_M Model memory/quality trade-off
2. CPU Threads: Don’t Guess

Try a small sequence such as:

-t 4
-t 6
-t 8

Benchmark the same workload. More threads do not automatically mean proportionally more tokens per second.

Example Benchmark — Copy and Test

Run the same model with different thread counts and compare the reported generation speed. Stop the server before starting the next test.

4 threads

./build/bin/llama-server \
  -m /models/Qwen3-8B-Q4_K_M.gguf \
  -c 4096 \
  -t 4 \
  -ngl 0 \
  --host 127.0.0.1 \
  --port 1111

6 threads

./build/bin/llama-server \
  -m /models/Qwen3-8B-Q4_K_M.gguf \
  -c 4096 \
  -t 6 \
  -ngl 0 \
  --host 127.0.0.1 \
  --port 1111

8 threads

./build/bin/llama-server \
  -m /models/Qwen3-8B-Q4_K_M.gguf \
  -c 4096 \
  -t 8 \
  -ngl 0 \
  --host 127.0.0.1 \
  --port 1111

Then send the same prompt for every run so that the comparison is meaningful:

curl http://127.0.0.1:1111/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3-8B-Q4_K_M.gguf",
    "messages": [
      {
        "role": "user",
        "content": "Write a Python function that checks whether a string is a palindrome. Include type hints and three unit tests."
      }
    ],
    "temperature": 0,
    "max_tokens": 256
  }'

For example, if 4, 6, and 8 threads produce 8.5, 10.2, and 9.8 tokens/sec respectively, use -t 6. More threads are not automatically faster.

💡 The Interesting Part

We started trying to make the model faster. The bigger discovery was that we could make the agent do less work.

3. Context: Bigger Isn’t Automatically Better

A useful starting profile is:

Machine Starting Context Starting CPU Threads
16 GB 4096 4–6
32 GB 8192 6–8

Increase context only when the coding task actually needs it. Large context consumes memory and increases processing work.

Example

A small bug fix touching two files may work comfortably with a 4K context. A cross-module refactor that needs several interfaces, implementations, and tests may justify 8K. Increasing context for every task simply makes the model process more text than necessary.

4. The Biggest Speed Hack Wasn’t an LLM Parameter

A repository contains plenty of work that does not require intelligence.

rg "UserService" .

rg "TODO|FIXME" src/

Let deterministic tools search. Let the model reason about the results.

Example — Search Before Asking the Model

Suppose you need every reference to UserService. From the repository root, run:

rg -n --hidden \
  --glob '!node_modules/**' \
  --glob '!.git/**' \
  "UserService" .

Use the resulting file list to decide which files Aider actually needs. The model then spends its context on understanding and changing relevant code rather than searching the whole repository.

✅ The Division of Labor
  • Search tools: find things.
  • LLM: reason and edit.
  • Compiler: compile.
  • Tests: validate.
  • Git: track changes.
5. MCP: Resist the Urge

One of the questions in the journey was whether multiple MCP servers were necessary for faster filesystem access and code changes.

For this Aider-only architecture, the answer was: not initially.

Start with the smallest useful stack:

Aider
+
llama-server
+
Git
+
ripgrep
+
shell

More tools are not automatically more capability. They also introduce more tool descriptions, choices, and context.

6. GPU Changes the Equation

If the machine has a supported GPU, -ngl becomes important. But there is no universal value that works for every GPU.

GPU VRAM
   ↓
Model size
   ↓
Quantization
   ↓
GPU offloading
   ↓
Benchmark
⚠️ Don’t Copy GPU Settings Blindly

An 8 GB GPU and a 24 GB GPU should not receive the same offloading configuration. Benchmark against the actual hardware.

Example

On a GPU with limited VRAM, attempting to offload too many layers can cause memory pressure or failure. A safer process is to start conservatively, measure tokens/sec and memory usage, then increase GPU offloading until performance stops improving reliably.

7. The Real Optimization
Before:

search → huge context → LLM → edit → repeat

After:

fast search → relevant context → LLM → focused edit → test
🎯 Part 2 Outcome

The fastest local coding agent isn’t necessarily the one with the fastest model. It is the one that makes the model do the least unnecessary work.


Series Navigation

← Part 1 — Make It Work

Part 3 — Make It Useful — Coming Soon

Leave a Reply

Your email address will not be published. Required fields are marked *