Fit more context on your GPU — and keep it.
Finetuna is not weight fine-tuning. No LoRA, no training. It tunes runtime settings — num_ctx, num_batch, num_gpu — and saves them as a reusable named Ollama model.
If even a few layers spill to CPU, generation can slow down 5–10×. Most people never check whether their model is fully resident — and Ollama’s conservative defaults leave plenty of 24GB cards quietly running 4K context.
npm · pnpm · Node 18+ · MIT
Three questions
- Will it fit? Verified via
/api/ps(size_vram/size) — not a guess - How much context can I actually get? Largest window that stays on the GPU
- Can I keep it? A named Modelfile variant any client can
ollama runforever
Honest “no change needed” is a feature. Auto-tune only switches settings when the win beats measured noise (median + spread).
Quick start
npm install -g finetuna
finetuna --check
finetuna
Look-but-don’t-touch: --check / --dry-run. Non-interactive: --model --name --ctx. Optional fit-search: --auto-tune.
Full flags live on the install page and the GitHub README.
The ollanet loop
Finetuna runs on the machine that hosts Ollama. To discover and chat with those models from another box on your LAN, Tailscale, or VPN, use ollanet.
Here: finetuna → a tuned named variant. Elsewhere: ollanet scan → ollanet prompt. Same API, closed loop.
What a run does
- Choose a source model and a new name
- Pick context / batch / GPU layers (presets through 128K)
- Write
Modelfile-finetunaand runollama create - Measure baseline speed; optionally search context that still fits
- Suggest a self-documenting name like
gemma4-ctx32k-flash
Goal: an optimum that still fits in memory — not “always shrink context,” not “always claim a speedup.”
Built by Catalyst Forge LLC.