shoehorn
Quantize a GGUF model to fill available VRAM, then run it with llama.cpp
TLDR
SYNOPSIS
shoehorn fit path|owner/repo|url [-i imatrix] [fit-flags] [-o out.gguf] [--serve]shoehorn plan -m model.gguf [-i imatrix] [fit-flags]shoehorn quantize -m model.gguf [-i imatrix] [fit-flags] -o out.ggufshoehorn run -m model.gguf [--ctx N] [--kv type] [-- llama-server-args...]shoehorn vram
DESCRIPTION
shoehorn quantizes a BF16 (or F16/F32) GGUF so the weights fill a memory envelope instead of using a fixed preset such as Q4_K_M. It subtracts KV-cache and estimated compute-buffer cost from the target VRAM, then solves a per-tensor mixed-precision assignment (Lagrangian knapsack plus a greedy top-up) that minimizes imatrix-weighted error under that byte budget. The encoder is implemented in this project; output is standard GGUF v3 that any llama.cpp build can load.fit is the one-shot path: resolve a local file, Hugging Face repo id, or URL, download the BF16 GGUF into ~/.cache/shoehorn (resumable), pick up a published imatrix or generate one with llama-imatrix, solve and write <stem>-fit.gguf, and optionally exec llama-server. plan prints the mix only. quantize writes the file. run execs llama-server -m model -c ctx -ngl 99. vram prints the Metal recommendedMaxWorkingSetSize.On Apple Silicon the default envelope is the Metal working-set probe. Elsewhere (or to target a different machine) pass --budget or --target. Inference and imatrix generation are delegated to llama.cpp tools on PATH.
PARAMETERS
-m, --model path
Source GGUF (BF16/F16/F32, or an already-quantized file decoded in-process). Required for plan, quantize, and run.-i, --imatrix path
Importance matrix (legacy binary or GGUF llama-imatrix output). If omitted, fit may generate one; plan/quantize fall back to activation-agnostic weights and warn.-o, --output path
Output GGUF. quantize requires it. fit defaults to <model-stem>-fit.gguf in the current directory.--ctx N
Context length used for the KV budget and for run (default 8192).--budget size
Total memory envelope (18GiB, 800MB, 4.5G, or bytes). Overrides the Metal probe.--target size
Envelope for another Mac approximated as 74% of the given RAM. Conflicts with --budget.--kv type
KV cache type to budget and run with: f16 (default), q8_0, or q4_0.--reserve size
Safety margin subtracted from the envelope (default 512MiB, or 160MiB with --calibrate).--calibrate
After the first write, load the model in llama-cli, read real KV/compute sizes, re-solve, and rewrite reused tensors.--exact-errors
Score every row instead of a 128-row sample per tensor.--serve
With fit, exec llama-server on the written file.
COMMANDS
fit source
Fetch or open source, obtain an imatrix, quantize to the envelope, write the GGUF. --serve launches llama-server on the result.plan
Solve and print the per-tensor table, type rollup, budget utilization, and projected VRAM. Does not write a file.quantize
Same solve as plan, then encode and write -o out.gguf.run
Exec llama-server with full GPU offload. Arguments after -- are passed through.vram
Print the detected Metal device and usable working-set size. Prints no Metal device found when the probe is unavailable.
CAVEATS
The Metal VRAM probe is Apple Silicon only; without a device you must pass --budget (or --target). The metal crate is a build dependency, so a Linux cargo install may fail even though --budget is documented to work anywhere.fit refuses split GGUFs when fetching from Hugging Face. The compute-buffer term is a heuristic; --reserve and --calibrate absorb the error. Serving and automatic imatrix generation need llama-server / llama-imatrix / llama-cli on PATH. IQ1 formats are not implemented (floor is IQ2_XXS). token_embd.weight and output.weight are floored at 4-bit.
SEE ALSO
llama.cpp(1), llama-cli(1), auto-round(1), ollama(1)
