ARTICLE / 00208 CHAPTERS

Make local models usable before making them bigger.

Notes from trying Ollama, quantization, VRAM, and context—moving local AI from a downloaded model toward a usable work layer.

  • Local AI
  • Quantization
  • Setup
RESEARCH NOTE
ARTICLE / 002Technical research and development notes

01

Ollama was an entry, not the only source.

Once a model ran, I separated where it came from from how it ran.

I do not always get models through Ollama. I often choose a model or GGUF file directly from Hugging Face, then decide which runtime fits it. Ollama is useful because downloading, managing, and exposing a local API are easy entry points.

Once I cared about quantization formats, context, GPU memory, and how inference actually runs, I moved deeper into llama.cpp and GGUF. The model source, file format, runtime, and agent are separate decisions; getting a model to start is not the same as finishing work.

My main setup is a Windows desktop with an RTX 3080 Ti, 64GB of RAM, and an SSD. With models such as Qwen3.8-27B, I look at weights, KV cache, context, and the headroom needed by other applications. An SSD helps reads and writes; it does not add VRAM.

Easier setup exposed the next questions.CONCEPT DIAGRAM
  1. 01

    Prepare

    Get the model

  2. 02

    Run

    Files and tools

  3. 03

    Check

    Can work continue?

02

Fitting the file is not enough.

Budget weights, KV cache and runtime space.

I used to focus on file size. A compressed model close to VRAM capacity seemed usable, but execution also needs KV cache, workspace and buffers. Windows and other apps need room too.

Agent context accumulates files, diffs, tool descriptions and terminal output. With longer context and more sessions, I experienced rising memory use, desktop stalls, driver resets and Windows freezes. These happened in my environment; they are not universal outcomes.

I now leave headroom instead of opening the maximum context immediately. Confirming one stable task before adding context or concurrency is more practical for me.

Budget weights, KV cache and runtime space.CONCEPT DIAGRAM
STEPWeightsContextHeadroom
Model dataIncludedNot includedNot included
KV cache budgetNot includedIncludedNot included
Runtime and other appsNot includedNot includedIncluded

03

Bring Q4, Q8 and IQ back to the same task.

Size and usability need separate checks.

Q4_K_M, Q8 and IQ variants kept appearing in my GGUF research. I initially ranked them by bit count, then started checking format, runtime support and remaining memory.

Smaller IQ3 variants were appealing because larger models seemed closer to my hardware. But an agent still carries long context after weights shrink. A smaller download does not establish long-running stability.

I use llama.cpp / GGUF as a research baseline. This is not a completed quantization ranking. Next I want a fixed task and evidence that edits and checks finish, rather than just fluent answers.

Size
Weights and headroom
Support
Runtime compatibility
Finish
Complete the same task

04

More quantization names do not make a better setup.

Separate method, numeric format and runtime.

Researching AWQ, GPTQ and QAT showed me that quantization is not explained by a 4-bit label. The methods handle error differently, and the target runtime still has to support the resulting format.

I also compare original weights and GGUF files directly on Hugging Face. Unsloth’s Qwen3.8-27B-GGUF is a concrete route to a pre-quantized Qwen3.8-27B file, so I can think about file size, remaining memory, and execution instead of treating quantization as only a theory.

I do not treat Unsloth as a harness or claim that I trained and ranked the model. It is one way to obtain and compare a quantized file; after downloading it, an inference environment such as llama.cpp or vLLM still has to run it against a real task.

Method
AWQ / GPTQ / QAT
Format
FP8 / NVFP4
Execution
Check runtime and GPU

05

Watch the wait before the answer too.

Separate loading, prefill, TTFT and decode.

Prefill processes the prompt and context; decode generates output. TTFT is the wait for the first token. Fast output does not remove a long wait before it starts, which can interrupt my work.

After reading files or running tools, an agent brings new content into the next step. My next comparison should record loading, input processing, first output and generation separately, including longer context, rather than compressing everything into tokens per second.

Separate loading, prefill, TTFT and decode.CONCEPT DIAGRAM
  1. 01

    Load

    Prepare model

  2. 02

    Prefill

    Process input

  3. 03

    Decode

    Record TTFT and generation

06

NInfer and Strata widened the questions.

Keep research directions separate from measurements.

NInfer drew my attention to matching particular models, hardware and runtimes. Strata made me consider how a whole PC could share work when a model does not fit the GPU. My interest moved beyond choosing a model.

These are research directions. Project demonstrations are not results from my Windows desktop. Deployment tools also need to fit the hardware, so I am comparing Ollama, llama.cpp / GGUF, LM Studio, and vLLM around startup, APIs, context, tool calls, and stability. These are research configurations, not one finished deployment.

Keep research directions separate from measurements.CONCEPT DIAGRAM

NInfer

Study focused configurations

Strata

Study whole-PC resource use

07

Next, compare the same task through completion.

Move from speaking speed to the result.

I want to hold the repository, task and tests constant, then vary model, quantization, runtime or harness. Changing one main condition at a time should make differences easier to trace. This is my next testing plan.

Alongside time and memory, I want correct tool calls, understandable edits, tests actually run and a record of human intervention. A usable final result is closer to my daily needs than the amount of generated text.

Move from speaking speed to the result.CONCEPT DIAGRAM
  1. 01

    Fix

    Repository, task, tests

  2. 02

    Observe

    Time, resources, tools

  3. 03

    Accept

    Is the result usable?

08

Separate the model, runtime, service, and harness.

Draw the work boundary instead of ranking names.

I now separate local AI into four roles: the model file, the runtime that executes it, an API or service when other software needs to call it, and the harness that reads a project, uses tools, manages context and permissions, and returns work for acceptance. These are roles in a workflow, not products on one ranking.

My model sources are mixed. I often go directly to Hugging Face for original weights or GGUF; Ollama is an easier entry for downloading, managing, and exposing a local API; llama.cpp is where I study GGUF, quantization, and GPU execution more closely; and vLLM is an inference engine oriented toward serving and throughput. llama.cpp can expose a server API itself, so service is a deployment role rather than a mandatory extra product.

OpenCode and Codex are the tools I actually rely on for daily work. Pi, DeepSeek Harness, and MiniMax Code are exploratory tools I use to understand design differences. I look at how they connect models, context, and tools, but I do not present a trial as daily experience.

Model
The weights
Runtime
llama.cpp, GGUF, inference
Service
Ollama, vLLM, API
Harness
Context, tools, acceptance