520 points by meetpateltech about 12 hours ago | 441 comments | View on ycombinator
moojacob about 12 hours ago |
artdigital 17 minutes ago |
My SuperGrok subscription previously easily lasted me through the week even with mild coding through Grok Build. Now when I use the app 1-2 times a day to ask some questions, I’m almost running out by the end of the 7 days. It’s terrible.
I want to keep using Grok but logically it makes no sense for me to keep paying for it on the side when my quota just doesn’t last. I have also no desire to upgrade to Plus with these terrible limits, while previously I would have eaten up a $100/mo Grok plan. Rumors say SuperGrok got heavily nerfed with the SuperGrok Plus introduction, and that sounds about right to me.
I’m sure it’s a great model and I’d love to use it. I hope they get their plans under control and only only focus on Grok Bot.
mchusma about 6 hours ago |
4.7 is definitely slower & more expensive. It feels kind of like they really had it burn tokens to claw up the benchmarks. But it's not super clear to me whether it's above the line or not. A part of that is that it is so slow that i haven't been making fast progress today with benchmarking it.
Overall, it it gets above my intelligence line its a good release...but you can read the tea leaves and tell the Grok team thinks this was a miss.
vessenes about 11 hours ago |
simonw about 11 hours ago |
Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
joegibbs about 2 hours ago |
I told it to compose an image (putting headgear on top of a head) - kept getting it completely wrong, generating new headgear, getting that wrong and screwing up the scaling.
I told it to diagnose a webhook issue that was happening in production from a local environment and it kept giving me moronic answers like that environment variables weren't set (despite me telling it that the values WERE set in production).
kvirani about 4 hours ago |
Is the training data more valuable ? The training process ? The harness ?
I know they are all important but where are they (all the frontier labs) really pushing to get incremental gains?
saejox about 11 hours ago |
xAI missed its chance, Ball is on Anthropic's court.
sarjann about 7 hours ago |
Unless they produce the same token output on the face of it, it looks like they're trying to cover for 4.7 not having good model perf?
meerita about 11 hours ago |
jjcm about 8 hours ago |
Designs: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Astra's build: https://html.non.io/annui/
Grok's build: https://html.non.io/Annui-grok/
Additional prompt instructions: "Add scrolling clouds behind the statues. Dynamically light the statues based on mouse position. Use diffui to generate the normal maps/depth maps/roughness maps of the objects, and to separate out the assets on to different layers."
Overall I find these models are getting good at following image as a source of instructions, but their refinement of the output varies heavily between the models. Astra's final output feels more polished, has better visual contrast, and the animations between the pages are smoother. Grok also chose to light all of the background elements, which imo overcooks it a bit.
Still though, for the price it's a great starting point.
trentor about 10 hours ago |
notduckrabbit about 10 hours ago |
pampas about 5 hours ago |
Grok 4.7 is near the top of the board. A significant improvement over Grok 4.6 but still not as good as Gemini 3.8 Flash which is very cheap and fast too.
WarmWash about 11 hours ago |
dom96 about 11 hours ago |
johnfahey about 10 hours ago |
qwerpy about 10 hours ago |
Excited to try 4.7. I hope they fixed the "it's not X, it's Y" that showed up in 4.6.
GodelNumbering about 10 hours ago |
From their headline comparison:
Grok: $2/$6 per million
Fable: $10/$50 per million
What this doesn't say: Grok costs 0.50/M cache read, Fable $0.25/M cache read
Long running agentic workflows are dominated by cache reads.Just makes Grok sound deceptive, and more importantly, reliant on user's lack of understanding of costs aka predatory (which in turn is more infuriating)
shdtabasum about 10 hours ago |
epsteingpt about 2 hours ago |
ls1911 about 12 hours ago |
gslepak about 10 hours ago |
sbseitz about 3 hours ago |
swalsh about 10 hours ago |
maz1b about 11 hours ago |
6thbit about 11 hours ago |
sourcecodeplz about 9 hours ago |
Output tokens from Intelligence Index:
- grok 4.6 (xhigh): 97M (for 44 score)
- grok 4.7 (xhigh): 240M (for 46 score)
simonw about 11 hours ago |
alansaber about 9 hours ago |
AM1010101 about 11 hours ago |
XCSme about 4 hours ago |
https://aibenchy.com/model/x-ai-grok-4-7-medium/#showcase=dd...
karma_daemon about 3 hours ago |
oh_no about 9 hours ago |
c0rruptbytes about 10 hours ago |
jascha_eng about 9 hours ago |
sidgtm about 11 hours ago |
andsoitis about 11 hours ago |
MuffinFlavored about 11 hours ago |
Is there a metric for like... time taken when comparing these two? I see score and cost.
If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?
Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".
undefined about 10 hours ago |
Razengan about 6 hours ago |
thih9 about 10 hours ago |
But also Xai doesn’t seem to care about user experience and long term support.
inshard about 9 hours ago |
gaigalas about 9 hours ago |
brcmthrowaway about 9 hours ago |
Saline9515 about 11 hours ago |
It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
kristofferR about 12 hours ago |
usumgallu about 9 hours ago |
undefined about 11 hours ago |
Tsarp about 11 hours ago |
bluepeter about 11 hours ago |
nicolamanzini about 7 hours ago |
mempko about 9 hours ago |
felixgallo about 11 hours ago |
myko about 7 hours ago |
TylerJaacks about 10 hours ago |
johnnyApplePRNG about 9 hours ago |
eleventen about 11 hours ago |
toader about 12 hours ago |
jmward01 about 11 hours ago |
simianwords about 12 hours ago |
The personality is bland and it doesn’t work nearly as hard or even tries to help.
finnjohnsen2 about 8 hours ago |
Maybe I'm in some kind of bouble but I have never met or talked to anyone who has used Grok.
Invictus0 about 8 hours ago |
enraged_camel about 10 hours ago |
outside1234 about 8 hours ago |
BoumTAC about 10 hours ago |
zug_zug about 11 hours ago |
I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?
claaams about 2 hours ago |
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.