341 points by benswerd 2 days ago | 154 comments | View on ycombinator
suby 2 days ago |
AntiRush 2 days ago |
https://web.archive.org/web/20091124210529/http://eis.ucsc.e...
There's a great contemporary Ars Technica piece by a competitor:
https://arstechnica.com/gaming/2011/01/skynet-meets-the-swar...
As an undergrad I did a project using genetic programming. It was not very successful, but it was a lot of fun.
https://tomisin.space/archive/starcraft-genetic-programming/
pelagicAustral 2 days ago |
I love StarCraft. I started playing it right from the beginning, most of my friends right now are from that era. I literally met people that have spread to almost every continent when I was in my early teens. We played at internet cafes and did not have access to the internet, that was priced differently...
I miss those days so much.
Everybody was from a different background back then, and nobody was anything other than a guy that plays StaCraft at the cybercafe... And now, we are in our 40's and I know Math teachers, history teachers, oil rig operators, software programmers, professional gamers, lawyers and more... hahah So crazy to think about it... and I know them, we talk, what a world.
mcteamster 2 days ago |
Protoss: powerful and expensive frontier coding agents you directly micromanage for the toughest tasks
Terran: versatile team comps of dedicated agent roles you can delegate well-defined tasks to
Zerg: massive swarms of specialist custom agents inside your apps that you evolve and optimise for speed and cost
Knowing every faction has its strengths and weaknesses helps me decide which tools to use for the job.
faeyanpiraat 2 days ago |
sqrt_1 1 day ago |
GodelNumbering 2 days ago |
karim79 1 day ago |
I will always love this and now I'm going to play it again. Remastered and on a fancy modern machine.
gadtfly 2 days ago |
It sounds like it might have been actually played in real time, which would be very important to distinguish.
I have recently seen other harnesses letting agents play real-time games in what seems like discrete time slices, turning eg Portal into something turn-based https://www.youtube.com/watch?v=ruuGXFAmiOE
alembic_fumes 1 day ago |
I'm often asking myself is it better to use higher or lower effort levels, or to maybe drop down to a "dumber" but faster model. And so using a real-time based competition as a benchmark could shed some light on this, I think.
In this vein, here are what I would love to see added in this benchmark:
- Include Google's Gemini models. I keep hearing Gemini being praised for its speed, and I would like to see whether that gives it a big enough edge over the bigger but slower models. - How does a Cerebras-accelerated open source model fare against a much larger but much slower frontier model?
I also feel like in general there is a lot of very low-hanging fruit to start benchmarking models across the spectrum of real-time vs batch-style workloads. Perhaps Brood War sits somewhere quite near the "real-time" end of the spectrum, but what about something like a game of speed chess, or a turn-based game with time limits?
I think what I would like to see the most is for someone to come up with a benchmark that supports tuning the "real-timeliness" of the benchmark, and then running a sweep of a model across the whole spectrum. That could get result in real nice graphs with multiple models on the pareto-frontier, varying based on the hosting provider and the model dimensions.
aswegs8 1 day ago |
herodoturtle 1 day ago |
tweakimp 2 days ago |
leobuskin 1 day ago |
I’ve scrolled the article, but haven’t noticed any remarks about Fable’s levels.
rubiquity 1 day ago |
karim79 1 day ago |
Brood War was the first video game to be broadcast on TV in Korea. I'm pretty sure it's still going.
c7b 2 days ago |
tianqi 1 day ago |
That’s me. I’ve always struggled with real-time games because I need to pause and think. While I excel at chess and board games, I’m just no good at real-time ones. At last I can only manage by sticking to a fixed set of tactics for a game, which minimizes the need for on-the-fly thinking. Seeing current models face the same difficulty leaves me with mixed feelings.
malfist 2 days ago |
Svenstaro 2 days ago |
EDIT: Nevermind, seems to work now. Watching a local qwen3.8-flash-next play this.
stymaar 1 day ago |
bigcat12345678 1 day ago |
therealdrag0 2 days ago |
mococa 2 days ago |
minimal_action 2 days ago |
Game_Ender 2 days ago |
winwang 2 days ago |
benswerd 2 days ago |
Havoc 1 day ago |
Surprised the outcomes are so poor though. I recall years ago AI was capable of beating pro level DOTA teams.
I guess in one case it was specifically trained on the interface & game while here it was not?
egeozcan 1 day ago |
steve_taylor 1 day ago |
bee_rider 2 days ago |
mkotlikov 1 day ago |
moomin 2 days ago |
xyzsparetimexyz 2 days ago |
undefined 2 days ago |
DJMolehill 2 days ago |
windowshopping 2 days ago |
nrightnour about 22 hours ago |
Barrin92 2 days ago |
I saw someone recently try to get an agentic system to play Final Fantasy and it did about as well as a Roomba.
undefined 1 day ago |
chaostheory 2 days ago |
iririririr 1 day ago |
rob313 1 day ago |
Really looking forward to playing- are you all limiting boxes?
DeepYogurt 1 day ago |
American87 2 days ago |
dschuessler 2 days ago |