
Make any language model fit your machine's memory precisely
shoehorn is a tool that fits large language models into the exact memory budget of your hardware. Instead of relying on preset quantizations that either waste headroom or fail at load time, it starts from the amount of VRAM you actually have, subtracts what inference needs, and solves a per-tensor mixed-precision assignment that lands within a rounding error of the remainder—routinely using 99.99% of the budget, sometimes to the byte. This means you can run a model that would otherwise be too large, or run a given model with higher precision than standard quantizations would allow.
It works with GGUF models and uses llama.cpp as the inference backend. The local web app measures your machine, streams the fit, renders the budget as a tape measure, and shows a perplexity number indicating what the fit cost. It then provides a Chat button for immediate interaction. shoehorn is designed for anyone who wants to run open-source language models on their own hardware—whether a Mac with 8–128 GB of unified memory, a GPU with 8–32 GB of VRAM, or a custom workstation—and wants to maximize quality without manual trial and error.
The software is MIT-licensed, implemented from scratch in Rust (no llama.cpp code linked), and outputs standard GGUF v3 that any llama.cpp-compatible loader can use. It's available for macOS (Apple Silicon), Linux (x86-64, NVIDIA/AMD), and Windows (x86-64, NVIDIA).
Solves an assignment of quantization precisions per tensor to fit the exact memory budget, using up to 99.99% of available VRAM—sometimes to the byte.
Starts from the memory you actually have, subtracts inference overhead, and aims for the remainder, avoiding wasted headroom or load-time failures.
Measures your machine, streams the fit process, and renders the memory budget as a tape measure, with a perplexity score for the quality cost.
After fitting, jump straight into a chat interface with the model, eliminating manual setup steps.
Loads any GGUF model and produces standard GGUF v3 output, so it works with anything downstream of llama.cpp.
A web page scans Hugging Face's most-downloaded models and shows which ones fit your hardware, ranked by quality your memory affords.
Available for macOS (Apple Silicon), Linux (x86-64 with NVIDIA or AMD), and Windows (x86-64 with NVIDIA).
The quantizer is implemented from scratch in Rust, with no llama.cpp code linked, and is free to use and modify.
Standard presets use fixed quantization levels for the whole model, which often waste memory or fail to fit. shoehorn solves a per-tensor mixed-precision assignment that uses nearly all of your exact memory budget, maximizing quality.
Yes, shoehorn uses llama.cpp as the inference backend. The Homebrew installation automatically pulls in llama.cpp for you, but you need it on your PATH to run the model.
Any GGUF format model, which includes most open-source models on Hugging Face. The 'What fits your machine?' page scans the most-downloaded models and shows which ones fit your hardware.
shoehorn has native downloads for macOS (Apple Silicon), Linux (x86-64 with NVIDIA/AMD), and Windows (x86-64 with NVIDIA). It can also be built from source using cargo.
Yes, it is MIT-licensed. The quantizer is implemented in Rust from scratch, and it outputs standard GGUF v3.
On macOS, you can use Homebrew: 'brew install notactuallytreyanastasio/shoehorn/shoehorn'. Prebuilt binaries are available for Linux and Windows, or you can build from source via 'cargo install --path .'.