Hacker news

  • Top
  • New
  • Past
  • Ask
  • Show
  • Jobs

Step 5 Preview: Advancing the Pareto Frontier (https://www.stepfun.com)

139 points by nateb2022 2 days ago | 33 comments | View on ycombinator

BoppreH 1 day ago |

In their first demo video, to make a 3D render of the photo, the thinking trace gives away the game:

> Interesting! It turns out there's already an existing project here [...] The project is fully built [...]

I'm always astounded how little effort is put into checking the AI answers displayed in these announcements. Back when I paid more attention, I remember OpenAI's and Google's demos constantly showed their AIs giving wrong answers.

nh43215rgb 1 day ago |

  > Built on a sparse Mixture-of-Experts architecture, Step 5 Preview has 600B total parameters, with 27B active per token, and supports a 1M-token context window and vision input.

  > Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index.

  > The model will be released with open weights on October 15.
I guess being Chinese company they decided to skip version 4, while also giving impression to be on the similar iteration with leading companies (claude opus 5). I wonder if other Chinese labs like Kimi/Moonshot will follow suit.

bethekind 1 day ago |

> Without any Pokémon-specific optimization, Step 5 Preview has so far sustained progress for more than 3,000 turns and 6 million tokens of interaction. By turn 3,082, it had unlocked Cut, earned three Gym Badges, and defeated Lt. Surge. The run is now roughly one-third of the way through the main story.

Finally FireRed is being used as a benchmark again! I believe Astra can beat it in 18 hours. Not sure how that compares.

garo-pro 1 day ago |

IT's Artificial Analysis Index is the same as Kimi K3, which is about 4.6x bigger, and GLM 5.3, which is about 1.25x bigger. Pricing is $1/$2.70 i/o. Openweights on October 15.

ghoshbishakh 1 day ago |

Their posisitoning is nice. Instead of saying they are cheaper and a bit less performant (in terms of intelligence), they say they are best among the cheaper and a bit less performant ones.

segmondy 1 day ago |

The previous Step models were pretty decent, but unfortunately for them, their model reasons too much and too long and the Qwen/Kimi/DeepSeek/GLM have been stronger. Hopefully this doesn't reason too long to get to the answer. I welcome any open model, the more the merrier.

InsideOutSanta 1 day ago |

GLM-5.3 and Kimi K3 are just below where I can use them to completely replace frontier models. Oddly,* SWE-2 is there for me.

If this performs similarly in the real world, we're approaching a level of capability where for most devs, it only makes sense to pay for Anthropic or OpenAI subscriptions if they are heavily subsidized and actually cheaper than these alternative options.

* Oddly, because I perceived Devin as being kind of a joke before trying SWE-2.

Jacques2Marais 1 day ago |

wrs 1 day ago |

Based on the examples, It’s like the thing took a writing class from Claude, but isn’t quite as smart — kind of worst of both worlds. The “Interactive Reporting” one is particularly terrible. A huge amount of waffly padding around a thesis that may or may not exist. So if the goal is to make long fancy reports that nobody will read, then, yes, very cost-effective.

“the question it raises matters more than the answer” ???

“So ‘the river drifts from cool to warm’ is not a figure of speech.” Oh really?

conception 1 day ago |

Huh wonder why they skipped 4?

dofm 1 day ago |

"Pareto frontier" really is the new "web-scale".

just60sec about 6 hours ago |

[flagged]

cboyardee 1 day ago |

[dead]

derliebej 1 day ago |

How about adding a contested historical facts benchmark?