CLI · runtime settings · named models
Get more out of a local model on the GPU you already have.
Tune runtime settings for your model and GPU, check whether the tested configuration stays on the GPU, and save a named variant. The saved settings fit the model and GPU Finetuna tested, under the load at test time. Weights are not trained or altered.
It sets num_ctx, num_batch, and num_gpu, then writes a Modelfile. The Modelfile records the configuration. Create the named Ollama variant from it, then run that model by name. Those settings change how much context and how many layers the host tries to keep in VRAM. They do not make the model reason better by themselves.
If part of the model spills to CPU, generation can run several times slower. Ollama picks a default context from detected VRAM, version, and any override. Check the CONTEXT column in ollama ps instead of assuming a 4K default from a 24 GB card label. Official docs use GiB bands, which are not the same as advertised GB.
npm · pnpm · Node 18+ · MIT
What it answers
Does this loaded run look GPU-resident? /api/ps compares size_vram to size. How much context still fits in that search? The largest window that stays on the GPU. Can you keep the settings? A named Modelfile variant.
A residency pass is for the model, context, and host load at check time. It is not a perpetual GPU guarantee. Re-run --verify after you change the model, context, concurrency, or host load.
--check and --dry-run never run ollama create. --verify loads an existing name and does not create a new one.
Leaving the incumbent is valid. The default context search picks the largest window that stayed on the GPU and was not measurably slower than the fastest candidate. A speed or batch search switches only when the win beats measured noise (median + spread).
Quick start
npm install -g finetuna
finetuna --check
finetuna
Non-interactive create: --model --name --ctx. Context fit-search: --auto-tune.
Flags: install and the GitHub README.
With ollanet
Finetuna runs on the machine that hosts Ollama. Each tool works alone. To find and chat with those models from another box, you can use ollanet.
Here: finetuna writes a named variant. There: ollanet scan then ollanet prompt. Same API.
What a run does
- Pick a source model and a new name
- Pick context / batch / GPU layers (presets through 128K)
- Write
Modelfile-finetunaand runollama create - Measure baseline speed; optionally search context that still fits
- Suggest a name like
gemma4-ctx32k-flash
Built by Catalyst Forge LLC.