LLMFit tells you what can run on something. I built something quite similar to their search into Shoehorn now.
puttycat 1 days ago [-]
This is really impressive. Can you say a bit about the underlying process? I'm guessing this is post-training qantization? Isn't PTQ also resource-intensive? (Ie might not work on any machine)
sscarduzio 1 days ago [-]
The project name is perfect!
mbuchel-hn 4 days ago [-]
does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?
rhgraysonii 4 days ago [-]
Yes that is exactly what this does.
kennywinker 4 days ago [-]
Could you explain what happens when you try to shoehorn a 2.4T parameter model into a 24gb m4 mac?
akshay_akula 3 days ago [-]
Wondering the same thing but for 48gb M5 Max.
metalliqaz 1 days ago [-]
extreme divergence would be my guess
jaylane 4 days ago [-]
tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running
rhgraysonii 4 days ago [-]
If you could post an issue if you still have the error around that would be awesome.
vancekai 11 hours ago [-]
[dead]
kelvo_ran 3 days ago [-]
[dead]
Rendered at 22:33:25 GMT+0000 (Coordinated Universal Time) with Vercel.
> AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF
you’re telling me you managed to fit Fable 5 into just 4B?
Qwen3 4b params distilled/trained with fable 5