243 points by matt_d 4 days ago | 41 comments | View on ycombinator
c7b 3 days ago |
infogulch 4 days ago |
If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
om8 4 days ago |
NooneAtAll3 4 days ago |
Who knew that if you actually look at information entropy you can pack stuff better!
yalok 3 days ago |
And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...
0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
wgd 4 days ago |
explainit2me 3 days ago |
kittikitti 3 days ago |
ant6n 3 days ago |
Why not just use an 8-bit LUT to encode the 256 most common ternary vectors with 6 components. That means of the possible 729 possible such vectors, you can only represent 256 different ones. You have to do more aggressive rounding, but at least the scheme is very simple to decompress and stream.
CodesInChaos 3 days ago |
Marchant_hq 3 days ago |
plqbfbv 4 days ago |
quillan_audits 3 days ago |
itsmeduncan 3 days ago |
locitra 3 days ago |
kadushka 4 days ago |
kadushka 4 days ago |
Kevcmk 4 days ago |
I honestly assumed that's how they already work. I have to admit that I even explained it like that to a friend. Why on earth wouldn't you design it like that from the start (talking about the adaptive, not the measure part; just sacrifice a few bits to clarify your encoding and save a ton of bits)?