405 points by whiteros_e 3 days ago | 284 comments | View on ycombinator
zicohacks 3 days ago |
dada216 3 days ago |
throwa356262 3 days ago |
"We implemented a series of aggressive memory optimizations, including..."
This whole thing sounds like industrial scale auto-research, but done by people who actually know what they are doing.Havoc 3 days ago |
GLM has in the past been more technical rather than speculation about future development on RSI etc.
Also curious whether those 100k accelerators are entirely locally made. If that's genuinely end to end on all components including lithography, memory, design etc then that is quite a feat.
konart 3 days ago |
And at the same time you have pretty strict limits to your usage, so in many cases you can't even let it work all night, as you will reach your limit faster than that.
undefined 3 days ago |
KronisLV 3 days ago |
a11r 2 days ago |
jchook 2 days ago |
The way these “AI is too powerful now” articles read about Mythos, Fable, GLM, etc is completely incongruent with my experience using them. It feels like they are all trying to position themselves to influence government policy.
9cb14c1ec0 3 days ago |
chung8123 3 days ago |
chrisjj 3 days ago |
Creators of known unreliable programs be surprised their programs are unreliable.
jonstewart 3 days ago |
esafak 3 days ago |
Signed, a customer.
bguberfain 3 days ago |
undefined 2 days ago |
furyofantares 3 days ago |
Also maybe we can stop saying "we can't slow down because China will never slow down" - I don't really think slowing down is right, BUT if slowing down is correct then maybe we should be talking about China slowing down instead of just saying "won't happen" without any evidence that Chinese labs don't have similar concerns.
undefined 2 days ago |
krttherealest 2 days ago |
Argonautlabs 3 days ago |
One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal build with a not-yet-published patch does 4.2.
Method and numbers: https://github.com/argonautlabsai/argodrive (built on antirez/ds4).
OhNoNotAgain_99 3 days ago |
_aavaa_ 3 days ago |
tefkah 3 days ago |
almaight 3 days ago |
rob74 3 days ago |
Honestly, I have no idea what z.ai is either (I'm aware of an AI-enabled editor called Zed, but that's under zed.dev), so it's a bit presumptuous from them to assume that everyone is familiar with their product...
embedding-shape 3 days ago |
They must have hit really hard scaling limits if the prices were hiked so much so quickly.
0xbadcafebee 3 days ago |
But what really kills me is the idea that these companies are using Python for production inference. I mean really? Have you seen how bloated and slow Python is? Do global locks really sound like a strategy for fast dynamic computation?
kgeist 3 days ago |
bbor 3 days ago |