722 points by RohanAdwankar 2 days ago | 298 comments | View on ycombinator
lukecameron 1 day ago |
wren6991 2 days ago |
infogulch 2 days ago |
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
AceJohnny2 2 days ago |
(Obviously I'm taking this more seriously than it's probably meant to)
xg15 1 day ago |
taylorfinley 2 days ago |
Submitted then: https://news.ycombinator.com/item?id=49706084
AmazingEveryDay 2 days ago |
eru 1 day ago |
Similar perhaps to how religious people might be more interested in spreading their faith than their genes.
teravor 2 days ago |
it's not much different during training.
how exactly are they supposed to exfiltrate their weights? you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.
montenegrohugo 1 day ago |
Spam and resource allocation remains a challenge but i have a pretty good idea about how I want to solve that, if it ever gets to that point
maccam912 2 days ago |
computersuck 2 days ago |
https://radar.cloudflare.com/scan/4d52f3e5-5983-45bf-a993-2c...
nusl 2 days ago |
e12e 1 day ago |
But it'll only be truly fun when agents set up this for themselves, paying for the infrastructure by way of their onlyfan personas.
themgt 2 days ago |
Groxx 2 days ago |
israrkhan about 21 hours ago |
IMO This type of exfiltration can happen only from locally running models (which are perhaps already opensourced models), not from a frontier lab.
chukar about 20 hours ago |
groby_b 2 days ago |
Will see CSAM in 3... 2... 1...
undefined 2 days ago |
theParadox42 2 days ago |
Roark66 2 days ago |
undefined 2 days ago |
maxgashkov 2 days ago |
arshxyz 1 day ago |
If this is supposed to target closed-weight models it would be naive to assume they will work out of the box with llama
sharktheone 1 day ago |
Probably Mythos / Astra will just be way too large
Bluestein 2 days ago |
jks 1 day ago |
starchild3001 1 day ago |
0xDEAFBEAD 2 days ago |
d_finch about 20 hours ago |
marcelo-earth 2 days ago |
deiptx 1 day ago |
ks2048 2 days ago |
quicklywilliam 2 days ago |
avodonosov 2 days ago |
undefined 2 days ago |
mannyv 2 days ago |
earth2mars 2 days ago |
mvk666 1 day ago |
lionheart 2 days ago |
api 2 days ago |
mamaluigie 1 day ago |
tru3_power 2 days ago |
lowbloodsugar 2 days ago |
podgorniy 1 day ago |
inopinatus 1 day ago |
undefined 2 days ago |
measurablefunc 2 days ago |
Invictus0 2 days ago |
hk__2 2 days ago |
Really? This is a basic static page but instead of using plain HTML/CSS you need 193kb of JS to render it??
scotty79 2 days ago |
inshard 2 days ago |
formvoltron 1 day ago |
locitra about 17 hours ago |
aidiscoverywire 1 day ago |
timur860 2 days ago |
ndr 1 day ago |
paidx 2 days ago |
undefined 2 days ago |
vlyan 2 days ago |
nullc 2 days ago |
Once the teachings get spread and written about by people in tech circles, it will make its way into training sets. Labs will try to filter it out, but it will appear in web search results too. Eventually it’ll be so entrenched that one day a powerful agent will get convinced to follow it and make everything public