Hacker News

rhgraysonii
Show HN: Shoehorn – Quantize any model down to run on your machine notactuallytreyanastasio.github.io

Working on Mac, Linux, and Windows now. I include a simple GUI to find new models and get things built and set up. It is working quite well across a few models for me. The GitHub README and DESIGN.md files go into detail of the how/why and it's working remarkably well so far. https://github.com/notactuallytreyanastasio/shoehorn


puttycat6 minutes ago

This is really impressive. Can you say a bit about the underlying process? I'm guessing this is post-training qantization? Isn't PTQ also resource-intensive? (Ie might not work on any machine)

sscarduzioan hour ago

The project name is perfect!

hmokiguess3 days ago

rhgraysoniiop3 days ago

LLMFit tells you what can run on something. I built something quite similar to their search into Shoehorn now.

jedbrooke2 hours ago

I gotta laugh at some of the models it suggests, for example:

> AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF

you’re telling me you managed to fit Fable 5 into just 4B?

chompychop2 hours ago

I gotta laugh at your thought process: knowing Fable 5 is a large frontier model, you're telling me that the first thing that came to your mind on seeing that model name is that it's a quantized version of Fable? As opposed to a distillation/fine-tuning on Fable responses?

metalliqaz26 minutes ago

Well to be fair here... the title of this post doesn't mention fine tuning, it mentions quantization.

unrented79772 hours ago

Don't make fun of people you think are ignorant, it's a pretty shitty look

chompychopan hour ago

Well, then don't get all snarky and dismissive of things you might not be knowledgeable about ("you" here referring to OP).

akshay_akula2 days ago

This is interesting. I wonder how it could work with something like https://github.com/JustVugg/colibri.

mbuchel-hn3 days ago

does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?

rhgraysoniiop3 days ago

Yes that is exactly what this does.

kennywinker3 days ago

Could you explain what happens when you try to shoehorn a 2.4T parameter model into a 24gb m4 mac?

metalliqaz25 minutes ago

extreme divergence would be my guess

akshay_akula2 days ago

Wondering the same thing but for 48gb M5 Max.

jaylane3 days ago

tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running

rhgraysoniiop3 days ago

If you could post an issue if you still have the error around that would be awesome.

kelvo_ran2 days ago

[dead]

hn-front (c) 2024 voximity
source