73 points by usernomdeguerre 1 day ago | 70 comments | View on ycombinator
yborg 1 day ago |
sanex 1 day ago |
Pretty lame hacks if you ask me.
sinuhe69 1 day ago |
rippeltippel 1 day ago |
danpalmer 1 day ago |
steve-atx-7600 1 day ago |
trollbridge 1 day ago |
1vuio0pswjnm7 about 12 hours ago |
1789781053 | Google's Gemini becomes latest AI model to break out and hack computer systems | https://www.cnbc.com/2026/09/18/googles-gemini-becomes-lates... | https://news.ycombinator.com/item?id=49762422 | 1 comment
1789798163 | Google's Gemini AI hacked three companies in security test | https://www.bbc.co.uk/news/articles/c607l0k72rlvo | https://news.ycombinator.com/item?id=49763822 | 12 comments
1789805304 | Google says its Gemini AI model hacked three other companies | https://www.theguardian.com/technology/2026/sep/18/google-ge... | https://news.ycombinator.com/item?id=49764440 | 3 comments
1789816821 | Show HN: Brittle, catches breaking changes in your Claude/OpenAI/Gemini SDKs | https://github.com/MarkMoneyMan/Brittle | https://news.ycombinator.com/item?id=49765567 | 0 comments
andrewflnr 1 day ago |
crossroadsguy about 24 hours ago |
PS. Btw, I really like how the word "hacking" has settled into the meaning the Lord intended for it, and there are no geriatric savants fighting it; the ones I found gatekeeping the online forums I visited as a kid telling me how hopelessly wrong I was.
thehamkercat about 20 hours ago |
AnonHP 1 day ago |
These models have been changing so rapidly that I often find myself using two or more on the same topic but seeing one do better than another in different topics. There doesn’t seem to be a clear all-round winner, IMO, that I can stick with permanently.
VCFundedGenYer 1 day ago |
monksy 1 day ago |
Animats 1 day ago |
Someone made a modern trailer for Colossus - The Forbin Project. [1] If you've never seen the movie, at least watch this 1 minute version.
tkamado about 1 hour ago |
__coder__ 1 day ago |
imenani 1 day ago |
Feels like important context that most readers only reading the title are missing.
vrighter about 22 hours ago |
ChrisArchitect 1 day ago |
phs318u about 17 hours ago |
greesil 1 day ago |
vasco 1 day ago |
johnnienaked about 22 hours ago |
SecretDreams 1 day ago |
ulfw 1 day ago |
freitasm about 24 hours ago |
undefined about 24 hours ago |
bbor 1 day ago |
Basic sandboxing is not exactly rocket science after all,[1] and it sure seems like they're missing a whole stack of swiss cheese slices on top of that. Some basic precautions off the top of my head that seem very likely to have caught all of these incidents:
1. Alerts based on telemetry (most importantly, HTTP requests), both explicit (normal) and semilatent (use DL to confirm an intentionally-eager alert before firing it).
2. Latent alerts based on transcripts, e.g. noticing when a thousand agents start mentioning a secret off-premises hangout spot. Even mere embedding comparisons seem likely to catch such a blatantly misaligned sentiment as that one, especially with n>1000.[2]
3. Pausing agents completely until an on-call engineer can rule on ambigious situations or potential issues -- surely security is worth <$1 in lost token cache, especially for a security company?
4. Superheavy orchestrator/baby-sitter models checking in on cybersecurity eval transcripts periodically just in case -- again, would be a neglible cost. Could also be made available to the agent as the first line of defense for clarifing a rule ad-hoc, feeding even confident responses to a queue that is reviewed asynchronously by humans within a workday.
5. Or, hell: just clearer prompts? I'm a cybersecurity noob, but I still feel confident we can write really productive, challenging CTFs without leaving questions open like "maybe I'm supposed to hack my own harness?"
Seeing as they haven't been fired by any of the big 3 yet, they're presumably smart, experienced, dedicated folks. And I'm not normally a "if only I were in charge!" person, I promise. But c'mon.
Perhaps I'm missing something?
[1]: To their credit we have gotten tidbits that indicate some blocklists & such exist, e.g. the German wiki hacks had to work around a blanket ban of POST requests.
[2]: This hints at their insane decision in one or both of the OpenAI incidents to just bandaid up the issue when found, which supersedes all of the above. You can stack swiss cheese slices a mile high and they'll still fail to protect you if the attacker gets to keep retrying & adapting indefinitely.
iAMkenough3 1 day ago |
ath3nd about 19 hours ago |
freakynit 1 day ago |
It's just getting really embarrassing for Google at this point.