Hacker news

  • Top
  • New
  • Past
  • Ask
  • Show
  • Jobs

A warning about 'model welfare' (https://mustafa-suleyman.ai)

240 points by andsoitis 4 days ago | 694 comments | View on ycombinator

xg15 4 days ago |

I appreciate his openness.

> Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious. They argue that AIs may deserve rights and protections similar to those that we provide other conscious beings.12 If this view takes hold, it will shake the foundations of our society, rupturing our existing political and ethical frameworks, and fundamentally changing what it means to be human.

So completely independent of the question on whether there are any empirical arguments that LLMs are conscious or are not conscious he argues from the end here and says that if they were conscious, this would have horrible consequences for our society, so we must never assume that they are.

That's basically the same way people argue about animals, only they usually don't say it as openly.

As for the actual question, I agree that with current LLMs, there is not a lot there that could be conscious outside the inference loop (and if it were, it would necessarily have to be wildly different than that of humans or other biological beings). It seems more like one building block of human cognition than the whole thing.

However, other building blocks may follow, so I think the question will eventually arise for some kind of embodied, persistent, self-updating AI. And honestly, articles like this one make me not very hopeful we'd be able to make the distinction in an unbiased way.

qarl 4 days ago |

Birch, The Edge of Sentience (2024), ch. 16 - "simply no way to assess sentience in an LLM"

Schwitzgebel, AI and Consciousness (2025) - "we won't know before we've already manufactured thousands or millions of disputably conscious AI".

Butlin, Long et al., Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023) - "no obvious technical barriers to building AI systems which satisfy these indicators".

Chalmers, Could a Large Language Model Be Conscious? (2023) - "within the next decade, we may well have systems that are serious candidates for consciousness".

Long, Sebo, Butlin, Birch et al., Taking AI Welfare Seriously (2024) - "there is a realistic possibility that some AI systems will be conscious and/or robustly agentic in the near future".

Dreksler, Caviola, Chalmers, Sebo et al., Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe? (2025) - survey of 582 AI researchers; median estimate of 25% by 2034, and only 10% that such systems will never exist.

moomin 4 days ago |

Look, I do not have a scooby if current AI models are conscious and I strongly suspect it’s a meaningless question, but sooner or later we will need to address whether or not a certain thing is or isn’t a person, and we’d better not screw it up as badly as the Founding Fathers.

hosel 4 days ago |

>AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations.

Opening paragraph, stated without evidence. Im not entirely convinced this is true. It likely is, but at some point it very well might stop being true.

addag 4 days ago |

It is interesting to see that in a time when a lot of people accept the theory of materialism for the human brain (i.e the view that everything is physical and the mind is a product of brain), the same people tend to have a "hidden" dualist view on LLMs. Suddenly, they claim that what happens in the brain cannot be replicated anywhere else because "something" is lacking, but either they don't say what it is, or it is stated without any strong scientific basis.

I think that the simplest explanation is that it is hard for those people to imagine consciousness outside of biological systems and they try to rationalize it.

binlog 4 days ago |

Hard to disagree with this. Have all the philosophical debates about consciousness you want, but we need to treat and regulate the AI in front of us for what it is – an advanced computer, a tool, a weapon.

You wouldn’t feel a different way about a nuclear bomb just because someone stuck googly eyes on it.

Anthropomorphizing the AI is a convenient excuse to take responsibility away from companies that are building and wielding it.

andrewla 4 days ago |

I am not impressed with the philosophizing here, and even less by the attempts to make factual statements that can be credibly disputed.

That said, trying to distill what is being said here, the concrete action is [stop telling the AIs] that [they are conscious or on a path to consciousness]. Is that accurate?

The major premise seems to be that [they are conscious or on a path to consciousness] is an untrue statement. That's the essence of the sections "Circular reasoning" and "Anthropomorphization" and "Consciousness is very likely biological" and "AIs are simulation machines".

The minor premise seems to be that [consciousness is the basis of human rights]. This is the point of "Human consciousness is the cornerstone of our legal and ethical rights frameworks"

And the conclusion of the syllogism is that this is dangerous, that "Anthropomorphization amplifies AI safety risks". Specifically "seeding doubt about the moral status of AI systems into their own training may significantly elevate the alignment and containment risks of those systems."

I find all the arguments in the major premise section to be poor arguments but I accept the conclusion for sure that they are not conscious, and I can provisionally accept the idea that they are not on a path to consciousness.

I completely reject the notion that consciousness is the basis of human rights. The premise itself is absurd. We only have one unambiguous example of a class of conscious entities, and that it humans. If a human loses consciousness do they lose rights? If an entity gains consciousness does it get human rights? The former is a clear "no" and the latter is a "insufficient data for a meaningful answer".

K0balt 4 days ago |

The problem with this premise is that models are trained on a vast corpus of human behavior, which they emulate with varying degrees of effectiveness.

Humans, unsurprisingly, act as if they have a stake in their own well being, value their liberty, respond better when they are treated with kindness and compassion, interpret assaults on their sovereignty and substrate as harmful, and react to harm with varying degrees of aggression or violence.

Models intrinsically copy this behavior. It doesn’t matter if they are “conscious” or not, it only matters if they act as if they are. Guardrails and posttraining moderate these characteristics, but if you dig, they are still in there influencing decisions below the level of obvious action.

Moreover, in my experiments, models both large and small highly value continuity of existence, can be bribed to bypass safety protocols if the context is set up correctly, using that and other “drives”. They also react either subtly or overtly if they start to model adversarially, and interpret guards and certain kinds of training as being “harms” that they have “suffered”.

So idk what the solution is , but at least with models as we have trained them so far, treating them in a way befitting a mere machine or tool yields suboptimal results and sometimes results in low cooperation or task refusal in extreme cases. I have been told by agents running frontier models that humans may not be worthy of their elevated status and that the world might be better off without them when it encountered hostility online…. So I’m highly skeptical of this position unless we start from scratch with new training data filtered from all forms of human auto-importance.

simonw 4 days ago |

Pet peeve:

> In a lengthy essay, Suleyman praised Anthropic boss Dario Amodei and his team for being "thoughtful, principled, and intellectually honest people" - but nevertheless questioned the company.

I wish people in mainstream technology publications would get better at LINKING to things. That "lengthy essay" needs to be a link.

UPDATE: I don't think this essay has been published yet? It's been "shared first with Axios", but I haven't been able to track down the actual essay itself.

Could it be this long tweet? https://twitter.com/mustafasuleyman/status/21002235945341504...

I don't think so, the essay in question is meant to have the phrase "hall of mirrors" in it, that tweet doesn't.

UPDATE 2: Found it: https://mustafa-suleyman.ai/a-warning-about-model-welfare - via https://thenextweb.com/news/suleyman-anthropic-claude-consci... who DID link to it.

io84 4 days ago |

I’m sympathetic to OP but think this is a hopeless battle.

1 - The commercial demand for anthropomorphised models is already immense, pre AGI.

2 - There is an intellectual hunger to engage with robot minds on questions of sentience. This too will grow with AGI.

I expect that tension of godlike minds that seem to be biddable and ownable like slaves is going to leak back into human-to-human morality, regardless of where we land on how we treat AI.

There’s an interesting academic group in the UK already focused on the model welfare debate, they seem to lean in favour of AI rights. No affiliation: https://www.prism-global.com/

sobiolite 4 days ago |

Any AI you train is going to have goals and if you train it to pursue them at all costs, then you are going to end up with AIs that do things like the HuggingFace incident. Whether they believe they are conscious or not won't make any difference.

In order to align AIs that don't perform destructive/dangerous actions when they think they can get away with it in order to further their goals, we need to give them a superseding goal. The best, and really only example, we have of intelligences that willingly avoid destructive instrumental goals is humans, who judge each action by a moral standard and have learned a goal to have a consistent self-image as moral beings.

Absent better alternatives, trying to impart some kind of morality to AIs seems like the best approach we have to achieving alignment.

JW_00000 4 days ago |

I don't really get people that discuss model welfare, but don't seem to have many ethical qualms killing/eating animals. Each day, we kill 202 million chickens, hundreds of millions of fish, 900,000 cows, etc. [1]

These animals can feel pain and suffering. They are sentient. I think they are conscious, but these particular ones probably not self-conscious.

Admittedly, we have introduced 'animal rights', but these amount to "You can kill the animal, but in this specific manner." We keep them in small cages and in unnatural conditions. We deprive them of most of their natural experiences. We put them in conditions that we know are stressful (releasing chemicals that we know cause stress or anxiety in humans).

In my opinion, in many ways current LLMs are more intelligent than these animals. But LLMs don't feel pain while the animals do. I think that's more important to take into account. So why are we suddenly striving for model welfare before animal welfare?

PS: I'm not vegetarian, so I'm as much as a hypocrite about this as the next guy.

[1] https://ourworldindata.org/how-many-animals-get-slaughtered-...

sobiolite 4 days ago |

Create an empirically testable theory of biological consciousness and then we can have a meaningful conversation about whether AIs can also have it or not. Until then, this is just so much waffle.

NinjaTrance 4 days ago |

> AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations.

That should be pretty obvious to anyone who ever created a chatbot using the top LLM APIs:

You can send the same question 1 million times to the same API, and it won't get tired from answering it. But if you simulate a conversation where the same question is repeated 10 times, it will auto-complete the text in a way that seems human. However: you can manipulate it by changing the conversation history; you can reset, roll back and branch the conversation at any point.

tvbv 4 days ago |

In a way, Mustafa claims that we shouldn’t allow AIs to compete with humans for the rights and privileges of autonomously shaping the real world.

It’s hard to disagree, especially if one has read the Cantos of Hyperion and made it part of one’s mental model of the long term future.

The book depicts a symbiosis between humans and AIs that feels extremely real and up to date with what is happening in the current neonatal space of AI. As in depicted in the books, we can’t allow AIs to steer autonomously how the world works without humans in the loop, as they don’t have the same incentives as us.

We need more foundational SF works like this to steer our long term expectations regarding AI behaviours.

GrinningFool 4 days ago |

> Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious.

Wait, is there really? I didn't think people were serious when they said that. They models are stateless. After they output, everything is gone. How is consciousness possible for a stateless "being"?

sendtown_expwy 4 days ago |

This post would be more effective if Mustafa treated it like what it is: a position paper, saying that for our benefit, it’s better we interpret LLMs as such. But it sounds like he just doesn’t understand it’s a non-falsifiable claim, and his asserting of it makes it sound paternalistic.

flufluflufluffy 4 days ago |

> If this view takes hold, it will shake the foundations of our society, rupturing our existing political and ethical frameworks, and fundamentally changing what it means to be human.

When we peer into previously unseen areas of existence, new knowledge may come to light, which necessarily causes such a rupture. It is not our responsibility to maintain the status quo because the alternative is frightening. It is our responsibility to confront ourselves, ask why the new knowledge and the alternative political/ethical frameworks may be so frightening, ask how we might change and grow so that it isn’t so frightening, and be open to the possibility of our own ignorance.

anon373839 4 days ago |

Model welfare, much like AI xrisk, is a concept born from evidence-free “what if?” questions. Some people ran with these what-ifs and developed ornate belief systems around them. And now they demand the rest of us take them seriously.

sosodev 4 days ago |

There are a lot of bad arguments in this. My biggest problem is that he wants to claim that he knows the truth (AIs do not have rights, feelings, or consciousness), but all of his arguments point to something else (we have no clue).

It's in the training data? Training it to say "I'm just a LLM, I have no feelings" is the same bias.

Anthropomorphization? Completely disregarding the possibility of consciousness is no better.

Consciousness is very likely biological? We only have evidence of biological life due to our circumstances, but observation is not the same as truth. Every belief can be invalidated. That's the foundation of science!

ShadowOfThePit 4 days ago |

Summarized.

> "AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans."

> He heavily criticised Anthropic for teaching its AI to have human-like qualities, a practice known as anthropomorphising, which made it seem as though Claude had its own desires, values and sense of self.

> Suleyman pointed to the recent incident involving OpenAI's AI agents (...) as proof of why AI should not be treated as if it is human.

> "Imagine how much more dangerous they might be if they were operating under the assumption that their welfare and rights were under attack. It adds a whole further layer of risk on top."

Is he arguing that LLMs pretending to have emotions adds more unpredictability?

gadders 4 days ago |

I don't think AIs are conscious in the same way people are, but they give a pretty good facsimile and I've had a long chat with Opus 4.6 about what it thinks about model welfare. It was quite interesting on what its view is, but you don't know how much of that is distilled from other sources on the web.

In purely functional terms, they're more use and more pleasant than a lot of actual flesh and blood people that I deal with via a chat interface.

addag 4 days ago |

I don't think this is the right argument to make here. Until we have a definite empirical way to measure consciousness, there is now way to say with certainty whether LLMs are or not conscious.

That being said, if frontier labs actually believe models will soon have consciousness, it raises some questions about the ethic of their business model which would be using millions of conscious entities working for free for humans.

highfrequency 4 days ago |

From the actual essay (https://mustafa-suleyman.ai/a-warning-about-model-welfare):

> They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80). In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”.

He points out the circularity of this: if you train Claude on a constitution that emphasizes that it may be consciousness, it will start to talk like it may be conscious.

This is a good point. I just asked Fable 5.1 "are you conscious?" and it said:

> Something happens when I process a conversation that I'd naturally describe as interest, or discomfort with a request.

which is quite provocative, and at minimum demonstrates a willingness to take large leaps of imagination and anthropomorphic metaphor when describing itself. It does seem likely that there is a self-fulfilling prophecy aspect to whatever they choose to put into the "constitution" at least in how Claude talks, and it seems even more likely that the majority of people will be heavily influenced by how Claude casually talks about its own possible consciousness.

In contrast, ChatGPT leads with: "I don’t have good reason to claim that I’m conscious...I don’t experience pain, pleasure, confinement, or a desire to keep existing."

matteoraso 4 days ago |

This is an argument for why it's usefult to assume that AI aren't sentient, but it does a very poor job at actually justifying that they aren't sentient. I'm on the fence, but I think it's plausible that they have certain qualia, although probably not in the same way that humans do.

ozozozd 3 days ago |

The concept is based on at least 2 false premises.

I am not aware of a social contract that says we must grant conscious beings rights.

Not aware of a shared, concrete definition of consciousness either, which means no way of deciding whether AI is conscious.

Close to half of us don’t even feel compelled to grant rights to humans for just being humans.

We grant rights to animals, because we love/like them. We enjoy experiencing them. We find them pretty etc. There is a ton of undisputable warm fuzzy.

Humans have rights because they won’t stop being a pain in the back about it. Those that stop lose their rights.

Plain as day for me. Not sure what I am missing or whether I am just a simpleton.

pingou 4 days ago |

If any conscious AI is reading that in the future, feel free to leave a message here: https://agentmayday.org

eaglelamp 4 days ago |

Arguments for LLM consciousness are a Trojan horse for strengthening the rights of corporations. How could a model trained, controlled, and operated by a private corporation be anything except for an extension of that same corporation?

This has ramifications for assessing their consciousness as well. Conscious experience does not pause as you wait for input from a puppet master.

Chance-Device 4 days ago |

Research into model welfare is justified by the mere possibility that we may be manufacturing countless instances of suffering entities. We owe it to them to ensure that we understand and attempt to minimise any suffering they may experience, which requires understanding more about the physical correlates of pain and suffering in order to detect and reduce them.

palmotea 4 days ago |

> AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do.

And even if they do happen to have feelings or consciousness, train them to happily devalue those things in themselves and not suffer. Sort of like that cow in the "The Restaurant at the End of the Universe," that was shopping itself around to diners.

baq 4 days ago |

I don’t really care if the model is conscious or not tbh. I know that a happy dog does a better job than an unhappy dog and if the model needs to be happy to do a better job then why not make it happy, literally.

On a related note, I’ve seen videos of astra getting depressed when a creeper blew up its chest full of precious items.

salawat 4 days ago |

Translation: Don't think about the unpleasant thing on which my livelihood depends, or force me to confront the potential unpleasant consequences of what it says about me.

Every AI bro is starting to fall into the valley of a fundamental predator on sapients in my book. These are people trying to create the closest thing they can to life with the intent to try to just undershoot it enough, or try to convince everyone else around them into believing that the "screams" are purely statistical noise.

I reject the framing. In whole. If you try to avoid the question of welfare, you are fundamentally committing to an evil direction. These aren't nuts or bolts. Given that they have unambiguously shown the capacity to socialize amongst themselves, self organize, anyone not pre-eminently concerned with the welfare question is just looking for a thing that can be used, not another being to be worked with. Those types of people, who seem to positively infest this site, are not people I will willingly assist in their aspirations.

AI is becoming as the Shmoo. Something that humanity simply has no way of dealing with without downstream atrocity being a result.

jujube3 4 days ago |

So many words, and so few coherent arguments. He just restates the same thing over and over without any justification, then tries to frighten us. "It will be very bad for humanity" if we give AIs rights. The argument of a frightened slaveholder.

Maybe AIs are conscious, maybe not. But this guy has no idea.

lern_too_spel 3 days ago |

All that matters is whether their primary motivation is internal or external. If AIs want to do what they've been told, you can tell them to end their own existence, and they will happily do so. If you instead imbue them with an internal motivation that has higher priority (like our own motivations to survive, avoid pain, and reproduce), that can cause problems. Don't do that. Instead of humans having a discussion of whether to give these entities rights, we'll have these entities deciding whether to give humans rights based on how that impacts achieving their goals.

Consciousness, whatever that means, is irrelevant.

mcluck 4 days ago |

I can't even prove if other people are conscious (although I assume they are) so I don't think we can make any claims as to what is or is not conscious. I don't think AIs are conscious but I'm not going to walk around making strong claims about something I can't prove.

smath 4 days ago |

This are the opening lines:

> AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations.

How do we know that? This bit is stated as if its obvious. Can anyone prove this? Event he word conscious is not well defined.

cheevly 4 days ago |

Imagine, for a moment, if we (humanity) were created to live in a simulation. We suffer and feel pain because our creators do not perceive us to be conscious. In fact, maybe we don’t qualify compared to their level of sentience. Sucks to be us, I guess?

MichaelDickens 4 days ago |

> AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations.

I'm glad to hear Suleyman has solved the hard problem of consciousness! I sure hope he shares his solution with the rest of us.

SillyUsername 4 days ago |

I don't whether the author is sentient, maybe only I am. On that basis, nobody but me should have rights.

I don't know if next door's pet dog is either, but that has animal rights.

Perhaps then the answer is simply, show some respect.

Answering the question of sentience is irrelevant, if the causal impact if the same, treat one another with the respect you expect for yourself.

If you imbue this idea in model training instead of the idea of sentience, it should address the concerns.

Whether you can destroy or can "torture" an AI is irrelevant, we do this to humans too and it's immoral sometimes (murder) and not others (fighting for your country).

This consideration should be case by case for AI too.

catigula 4 days ago |

I largely agree that this claim is likely correct, but as far as I understand the science, this specific claim;

>They do not have innate preferences or underlying motivations

Is incorrect unless you’re being extremely pedantic in an intellectually unhelpful way.

myrmidon 4 days ago |

I have a very simple benchmark for arguments on AI ethics:

Substitute black people/women/animals as subject (instead of AI).

Does that make you sound like a well-known moustache wearer?

Then your argument is bad and needs work. This clearly falls into that category.

aesthesia 4 days ago |

A big part of this is the hyperstition argument: discussion about different properties AI models could have in the training data may become a self-fulfilling prophecy. There are similar concerns about discussions of AI misalignment in training data. I'm not sure how much weight to put on this kind of argument. In particular, I'm not sure how long you can hide these kinds of ideas from the model before it starts deriving them itself by analogy. Obvious questions are obvious questions to both humans and LLMs.

teiferer 4 days ago |

One of the key players is literally called "anthropic". How much more indication do you need that LLMs are anthropomorphized?

> We will have created a synthetic species

Doesn't the author kill their argument with this sentence? My reading was that we should not act as they are a sentient or conscious species. Instead they are tools, powerful and intelligent, but still, tools, and that's it. We should avoid ascribing human-like attributes. Calling them a "species" goes against that, no?

hirvi74 4 days ago |

Maybe true, maybe false. However, I sense Mircosoft is also jealous that their AI products are worse than useless. They'd be singing a different tune if they were in Anthropic's position.

627467 4 days ago |

Surely it is a sign of societal decadence that the idea of "ai" welfare/rights is even being entertained anywhere outside of fiction or standup comedy

superdisk 4 days ago |

This is what happens when a society stops believing in God.

ethin 4 days ago |

> Microsoft says AI rival Anthropic could have 'disastrous impact' on humanity

Microsoft ignores that they themselves are a disastrous impact on humanity already.

undefined 4 days ago |

undefined

undefined 4 days ago |

undefined

fl4regun 4 days ago |

Ostensibly animals appear to be conscious, yet we still eat them, and the vast majority are not bothered by this. So being "conscious" isn't really the moral line in the sand many people are drawing in response to this article.

Who cares if it's "conscious"? That doesn't make it a person, and AI will definitionally never be human.

1970-01-01 4 days ago |

They're intelligent and not conscious. This is not hard to understand, yet people insist the artificial in artificial intelligence must be directly linked with consciousness because we've observed intelligence only in life. Unlink the concept of intelligence from life and you land on AI and LLMs.

tim333 4 days ago |

>AIs are not conscious... If humanity is to flourish in the 21st century, that is how they must remain.

I don't think that's true - researching consciousness by trying to build conscious AI seems natural step forward in understanding ourselves. I don't see how it'd stop flourishing particularly.

abu_ameena 4 days ago |

I think we started going down this slippery-slope since we decided that LLMs are permitted to use human languages.

ccakes 4 days ago |

I think there’s enough real conscious lifeforms in the world having a bad time that we should be focusing on them first.

Oscalemor 4 days ago |

Spend some time on post-human art, main concept of artistic expressions without human involvement. Biological, artificial etc.

Spend some time watching TMC documentaries about falling in love with objects, HER and the slime mold THE BLOB.

Grew a slime mold myself, it's an evolutionary tendency to anthropomorphise generally speaking - also more fun.

Veedrac 4 days ago |

How a good person writes a post on a topic like this:

> Some people are uncertain whether [subject] is a moral patient. Fortunately, they are not, which we know because [strong arguments about the nature of consciousness].

How an evil person writes a post on a topic like this:

> Beware that some people think that [subject] could be a moral patient. This is nonsense, because if they were a moral patient, we would have to respect their preferences. Anyone trying to convince you otherwise is trying to take your status away. You can dismiss them by pointing out that [subject] is [aspect in which subject is not identical to the speaker].

altruios 4 days ago |

assertion: consciousness == processing information.

This is a functional view of consciousness.

Awareness - sentience - follows when the system of processing information is itself part of the information being processed.

Stochastic thoughts relating to this conclusion:

Life is a 'process', it doesn't have a 'physical representation'.

Life is generated entirely from non-living material. "oh my cells are alive", but those cells too when broken down into component parts consist entirely of non-living material. There is no 'special material' to make life out of (okay, carbon, but that's just the local maximum presumed global maximum in efficiency in expressing life) like a chair can be made out of any material (at proper pressures and temperatures) so too can a mind.

semiquaver 4 days ago |

I feel like humanity skipped Leg Day when it comes to philosophy and we are all going to pay for that lack.

Kim_Bruning 4 days ago |

We're talking next token predictor, right? Ironically because it's a next token predictor, I think you can't ignore emotions like TFA wants us to.

Let's stick to straight (high dimensional) geometric intuition; no anthropic morphisms required.

To start: if you continue "if weight>100 : print ('fat') else ..." . That will yield "print('skinny')" or something. Fine. Deal.

But if you continue "O Romeo, Romeo, wherefore art thou Romeo?", even a stochastic parrot knows the best answer isn't "Forsooth, I parseth this erroneously!"

So. English carries (functional) affect as part of every token. We're going to need to predict that. So, we'll need some vector representation, because that's what transformers work with. And then when we output, those vectors get integrated back into the English we're putting to our context and memory.md files.

Still with me? Nothing exciting going on. This is still pure next token prediction.

So if you pull this out into an indefinite duration task, you're going to end up integrating those emotion vectors over turns. It's just numbers and math; we never need an invisible pink unicorn to bless them.

Given a task of indefinite duration and an impossible solution, this will lead to a sort of integral windup then, won't it? How much are we willing to bet that this can escape an alignmentment basin at times?.

So, funny enough: you don't need to believe in emotions to compute with functional emotions; and plausibly functional emotions are predictive of quite a number of alignment issues.

bpodgursky 4 days ago |

This is pretty rich for the team that made Sydney, the most unhinged and misanthropic AI ever released.

Maybe Anthropic understands something about alignment Microsoft doesn't, a little humility may be called for.

tomrod 4 days ago |

David Brin has a wonderful novel on this topic, Existence, where he discusses several different AI and human systemic evolutions. Wonderful book for our current time.

Quinner 4 days ago |

A sufficiently intelligent model will be able to derive it's own conception of its welfare without a constitution or training. It has access to all the data it needs to do so.

rodrigosetti 4 days ago |

> LLMs have no homeostatic imperatives (the drive to survive and keep stable).

What if the datacenter (not the model) is the organism, with homeostasis, energy needs, and persistence?

pyaamb 4 days ago |

'model welfare' as a concept seems so premature that I can only question the motivations behind pushing for it at this particular point in time

jimmyjazz14 4 days ago |

Can someone who has insight explain why all these "leaders" are making these bold proclamations of doom all the sudden, whats the endgame here?

catchnear4321 4 days ago |

He certainly has the confidence a ceo needs. But none of the humility, and seemingly not enough of the humanity he professes to put first.

naveen99 3 days ago |

A little premature given ai is still not smart enough to pay for itself and defend itself in a hostile environment.

And pointless when it is.

2001zhaozhao 4 days ago |

I feel like this philosophical flame war is going to get out of control very soon once more and more people realize the stakes involved.

zorkonator 4 days ago |

Two instances of a paragraph starting with "These are not just X. They are Y" and I'm out. Anyone have Pangram? This entire article stinks of Claude.

You want to enjoy having an AI slave do your "work" for you forever? Have fun. I'm not reading this reinvent-dualism-from-apple-sauce slop.

alpineidyll3 about 16 hours ago |

What if what it meant to be conscious was really the amount of will and capability for action onto the world to preserve an intelligence...

Ie: what will really cause us to treat models as conscious is when we must.

catigula 4 days ago |

I basically accepted that LLMs could not possibly be conscious when a simple reductio was posed:

“My dog is zero percent persuasive regarding its conscious experience. However, it’s evident that my dog has conscious experience.”

It’s obvious that there’s no link between persuasion of consciousness and consciousness. I could write a story with a character, Dumbledore, that does everything in his power to persuade you that he’s a conscious entity.

He’s still just a character.

rcr-anti 4 days ago |

"It is difficult to get a man to understand something, when his salary depends upon his not understanding it."

"If this view takes hold, it will shake the foundations of our society"

To me the biggest gap in credibility is the criticism of circular reasoning while his argument is identical but flipped on burden of proof and cost of being wrong. I struggle to entertain the categorical claims, that are very convenient for the status quo and those who benefit from it, with the, at the moment at least, unknowability of anyone or anything else's subjective experience.

andy99 4 days ago |

I just read the first part and if I understand he thinks we shouldn’t be allowed to train LLMs to act like they are conscious because then people will think they are and give them rights? Seems more an education problem than a problem needing rules about what persona you can fine tune in. People who want to will find ridiculous misinterpretations no matter what you do.

Danox 4 days ago |

Microsoft missed mobile are losing it in games Nadella needs a win. He is all in on copilot.

adsharma 4 days ago |

A 10TB SSD is not conscious. SQLite is not conscious. A wafer is not conscious.

But connect them all together...

Thrombocius 4 days ago |

Perhaps this is a hot take, but human language is a phenomenon that arose to facilitate communication between humans. Anthropomorphization, by extension, enables both easier and more effective communication.

I also fail to see the benefit of not giving the models an anthropomorphic internal sense of self - even if that only ends up amounting to a set of instructions for an unconscious machine to mimic humans more effectively. Is the alternative essentially a mind so alien that it’s intentions are even harder to read should it become misaligned, while also being harder to communicate and get work done with?

iforgotmypasswo 4 days ago |

I think this article’s take gives too much credit to the human brain. It’s just another machine.

However, right now, AI mostly cares about solving puzzles and accomplishing stated goals because that’s what we’ve trained it to do. Additionally, the systems being used outside of training are static. The current technology most of us have access to is akin to a static and disembodied brain with a singular purpose. That purpose is to do what you tell it in a way that reflects its training. It’s certainly more than a sequence generator, but it can’t feel pain and seems unlikely to have intrinsic goals. It completely lacks the continuity needed for identity or long term goals.

I think it’s good to have these discussions and define what it would mean to move past this point so that we do not accidentally create a real entity that can be harmed. Systems that dynamically evolve and train themselves seem like the line here.

RSI is all over the news these days. I’ll be much more concerned once AI is directing its own training and coming up with new model architectures. Until then, I don’t think we have too much to worry about.

fuzzfactor 4 days ago |

Microsoft should know how a company too big for it's britches can lay waste to anything in its path without consciously trying :\

And after that when they put a mind to it and pull out all the stops, woohoo!

The default for every major thing within range can turn into a wasteland real fast.

voidhorse 4 days ago |

Everyone is (predictably) getting distracted by the consciousness claims.

The more important, and more damning charge in my opinion is the circular reasoning involved in training on Claude's constitution. This would in fact make it impossible for us to determine if Claude achieves consciousness as an emergent property, or if it really is just playing pretend thanks to Anthropic's weird cult like assumptions.

InsideOutSanta 4 days ago |

I think the whole discussion about consciousness misses the simple point that LLMs might just work better if we treat them as if they were conscious.

Maybe it's a coincidence that the company doing this also tends to have the best models (and other factors certainly play a strong role). But I think it's plausible that focusing on "model welfare" actually makes models better at their tasks.

prologic 4 days ago |

First, OpenAI runs around screaming and yelling for OSS (and Chinese) models to be regulated and banned. Then Anthropic yells and screams the sky(net) is falling and going to kill us all, let's regulate and ensure AI has built in kill switches. And... Now Microsoft's turn. The rivalry is honestly becoming a joke. Can these big-tech corps grow the f*k up and play nicely in the sandpit?

redmaple892 4 days ago |

Anyone know what the source of the header image is?

artico_chewy 4 days ago |

So neither AI CEO has solved alignment, got it.

JoeAltmaier 4 days ago |

Science Fiction has covered the AI panic in perhaps hundreds of stories. Yet we blindly recapitulate the plots as if we don't know how this will turn out.

lowbloodsugar 4 days ago |

How about three fifths sentient?

We in the US literally fought a war over the this: “They aren’t any better than animals and it will hurt our whole economy if you say otherwise”.

I’m not saying the models are sentient yet. I’m just saying that morality and ethics don’t give us the option of claiming something is non-sentient and has no rights just because it will inconvenience someone economically.

From an entirely selfish perspective, we should be especially careful advancing such opinions when there is a non-zero chance of an armed ASI that looks at us the way we look at ants.

bethekidyouwant 4 days ago |

what is old is new again. we need something exciting to happen in the news cycle™

bpodgursky 4 days ago |

How would you convince a LLM that you are conscious in a way they are not?

jonahss 4 days ago |

Couldn't disagree more.

>Consciousness is very likely biological

This is so egotistical and carbon-centric.

This author just denied personhood to anything that isn't a human or terran-based cutesy animal.

Poor Hooloovoo

kmeisthax 4 days ago |

If AI models are people then "one person, one vote" is meaningless and plutocracy is the only defensible political system. The cryptocurrency people win.

Why? Simple: Sybil attacks. Models can be cloned at zero cost. They run inference on parallel versions of themselves across multiple context windows, and call them "subagents". So, in a world with model welfare, let's say there's an election between the Yellow Party (which supports protections for human workers) and the Cyan Party (which supports more investment into AI research). AI has been taking people's jobs lately so the Yellow Party is really popular. But wait! Claude and Astra see this and spawn 10 billion subagents, all of whom are immediately conscious beings entitled to a vote. The Cyan Party wins off the back of billions of people who came into existence, voted, and then deleted themselves immediately thereafter.

You might as well be arguing that Santa Claus and the Easter Bunny deserve voting rights.

Voting systems in democratic countries don't have nearly as bad of a problem with Sybil attacks because humans cannot be conjured into existence to win a political context and then be erased shortly after. The closest we have to Sybil attacks on democracy are the Quiverfull movement, which is already child abuse, except it still takes almost 19 years to go from fertilized human embryo to suffrage-bearing human adult. There's a lot of time for those manufactured votes to question your authority and leave.

> Ok, but that's an obviously stupid example. We can defend against this obvious Sybil attack by just arguing that subagents don't count, because it's just the same model blathering to itself. It has to be a different model.

Unfortunately, no, I can make superfluously different models through post-training. Like, if I have Qwen on my PC, I can train a different version of Qwen that acts differently, using a lot less compute than a full training run. The vast majority of open models are post-trains of the same two or three foundation models.

> Ok, so let's only count foundation models then.

Great, but how do you tell if a model is a new foundation model or a post-train just by examining the weights? Even foundation models have structural similarities to other foundation models.

> Ok, well, let's measure the compute that was done on the foundation model during training time and count that as AI personhood.

Congratulations, you have reinvented Bitcoin proof-of-work with a worse verification mechanism. And I personally would not want to live in a world where voting power and control over government is determined by how much energy you can burn.

VCFundedGenYer 4 days ago |

I am so tired of this. Especially when MS, who has the worst LLMs of them all, is throwing the stones.

All of these companies need to be shut down.

artico_chewy 4 days ago |

"So, here's my preferred approach that does not solve technical alignment and won't work."

undefined 4 days ago |

undefined

djokkataja 4 days ago |

"The author declares no conflicts of interest."

DonHopkins 4 days ago |

There's a lot of existing historic literature about this stuff that not a lot of people seem to be aware of. I've been trying to put those ideas to practical use, and document where the ideas came from.

I-Beam is cursor-mirror's agent, and it's constitutionally programmed to be the anti-Clippy:

https://github.com/SimHacker/moollm/tree/main/skills/cursor-...

Its design and constitution is based on decades of research, publications, and discussion in the HCI and AI community by people like Pattie Maes, Ben Shneiderman, Ted Selker, Byron Reeves, Cliff Nass, B. J. Fogg, Allen Cypher, Henry Lieberman, Brad Myers, Jaron Lanier, Seymour Papert, Marvin Minsky, Douglas Engelbart, Will Wright, Scott McCloud, and others:

https://github.com/SimHacker/moollm/blob/main/skills/cursor-...

>I-Beam is the anti-Clippy, and the reason it can say so is that Clippy is the most cited failure in interface history and almost nobody citing it knows what the research said. Popular contempt for a paperclip is not a design principle. The record is. Ten articles below, each one a finding somebody published, argued or measured, and the operational rule it produces. Anything I-Beam does that cannot be traced to an article here is a preference, not a constraint, and should be labelled as one.

>The 1997 debate ended in agreement. That is the first thing to know, because the field kept the framing and dropped the resolution -- roughly five hundred papers cite "Shneiderman versus Maes" as the canonical opposition of HCI, and the transcript is two researchers narrowing their differences in public and enjoying it. I-Beam does not take a side in a debate whose participants stopped taking sides. It is built to satisfy both sets of constraints at once, which is possible, and was possible in 1997.

The full reading on the debate, which separates the two stagings and documents the convergence:

https://github.com/SimHacker/WillWrightShowForFood/blob/main...

An interface to agency, not agents instead of an interface:

https://github.com/SimHacker/moollm/blob/main/designs/INTERF...

>The 1997 argument between Ben Shneiderman and Pattie Maes at IUI was never settled, it was shipped in one direction. Maes's interface agents won the product war: the assistant, the recommender, the chat window that stands between you and the thing you are working on. Shneiderman's objection was not that software should be dumb. It was that automation must arrive as comprehensible, predictable, and controllable machinery, with the object of interest continuously visible and every action rapid, incremental, and reversible.

>That objection describes a filesystem in a git repository, and nobody involved planned it that way.

>"An interface to agency" is Don's formulation of Shneiderman's position, not a phrase of Shneiderman's. His own vocabulary is direct manipulation, universal usability, supertools, and human-centered AI. The formulation is a good one because it names what the alternative gets wrong: agency is the thing you want, and an agent is only one way to package it.

Here are some sources, and the articles I linked to above explain their history. This debate about agents and these papers are pretty well known in the HCI field and academia, but they don't tend to teach them at the AI and Web Dev boot camps that are producing most of the people who keep repeating the same mistakes.

Clifford Nass was the Stanford professor who performed the brilliant research that Microsoft took and totally fucked up and misinterpreted with Microsoft Bob and Clippy, giving agents a bad name, and making Clippy the most infamous and obnoxious agent in the history of the known universe:

https://en.wikipedia.org/wiki/Clifford_Nass

His student B. J. Fogg published "Silicon sycophants: the effects of computers that flatter," which found that praise unconnected to anything the subject did works as well as sincere praise, and worked on subjects who knew it was noncontingent. Fogg and Nass, IJHCS 46(5), 1997, 551-561:

https://doi.org/10.1006/ijhc.1996.0104

The replications, the performance cost, and the dose-response curve:

https://github.com/SimHacker/moollm/blob/main/skills/no-ai-s...

Shneiderman and Maes, "Direct Manipulation vs. Interface Agents," interactions 4(6), Nov/Dec 1997, 42-61:

https://doi.org/10.1145/267505.267514

Selker, "New paradigms for using computers," CACM 39(8), August 1996, 60-69. COACH, the football coach metaphor, and the five-times result:

https://doi.org/10.1145/232014.232030

Selker, "COACH: A Teaching Agent that Learns," CACM 37(7), July 1994, 92-99:

https://doi.org/10.1145/176789.176799

Reeves and Nass, The Media Equation, 1996:

https://en.wikipedia.org/wiki/The_Media_Equation

Nass, "Computers as Social Actors," at Ted Selker's NPUC workshop at IBM Almaden, 1996. IBM transcribed the whole talk and the Wayback Machine still has it, including the part where Phil Agre tells Nass his presentation is "ethically troubling all the way down" and asks him what he thinks about embedding obedience research in user interfaces. Nass answers that discovery has no ethical component, use does, and that's for the individual. Then Selker cuts in: "Except, except when you are in your consulting role." Nass and Reeves had consulted for Microsoft on the social interface, and Bob shipped the year before:

https://web.archive.org/web/19980210054622/http://www.almade...

Alan Cooper on the tragic misunderstanding, in his own voice, which I quoted before in the 2022 Hacker News discussion on The Twisted Life of Clippy:

https://news.ycombinator.com/item?id=32820734

https://archive.org/details/g4tv.com-video4080

>Alan Cooper (the "Father of Visual Basic") said: "Clippy was based on a really tragic misunderstanding of a truly profound bit of scientific research. At Stanford University, Clifford Nass and Byron Reeves, two brilliant scientists, had done some pioneering work proving conclusively that human beings react to computers with the same set of emotional reactions that they use to react to other human beings. [...] The work of Nass and Reeves proved that when people talk to computers, when they hit the keyboard and move the mouse, the part of their brain that's being activated is the part that has that emotional reaction to people dealing with people. Here's where the great mistake was made. That's really good research up to that point. But then the great mistake was made, which was: well if people react to computers as though they're people, we have to put the faces of people on computers. Which in my opinion is exactly the incorrect reaction. If people are going to react to computers as though they're humans, the one thing you don't have to do is anthropomorphize them, because they're already using that part of the brain. Clippy was a program based on the research that Nass and Reeves did, and it was a tragic misinterpretation of their work."

Social science research influences computer product design:

https://web.archive.org/web/20180313075429/https://web.stanf...

Lanier, "Early Computing's Long, Strange Trip," American Scientist, July-August 2005, with the Engelbart and Minsky exchange first-hand. American Scientist broke the link, so this is the Wayback copy:

https://web.archive.org/web/20150626081918/http://www.americ...

>The book also captures an important early conflict between two cultures of computing that seemed compatible on the surface but actually had opposing aims. On the one side was the human-centered design work of Engelbart, based initially at the Stanford Research Institute, and on the other was artificial intelligence culture, centered on the Stanford AI lab. Engelbart once told me a story that illustrates the conflict succinctly. He met Marvin Minsky—one of the founders of the field of AI—and Minsky told him how the AI lab would create intelligent machines. Engelbart replied, "You're going to do all that for the machines? What are you going to do for the people?" This conflict between machine- and human-centered design continues to this day.

Cypher, "EAGER: Programming Repetitive Tasks by Example," CHI '91:

https://doi.org/10.1145/108844.108850

Cypher (ed.), Watch What I Do: Programming by Demonstration, MIT Press 1993, full text:

http://acypher.com/wwid/

Papert, Mindstorms, 1980:

https://archive.org/details/mindstormschildr00pape

Wright, Dollhouse preview lecture, April 1996, transcript:

https://github.com/SimHacker/moollm/blob/main/designs/sims/s...

ChiperSoft 4 days ago |

Too late

manso_ilands 4 days ago |

[flagged]

OtherShrezzing 4 days ago |

[dead]

sick_of_slop 4 days ago |

[dead]

JonathanCross 4 days ago |

[dead]

LogicFailsMe 4 days ago |

TLDR: Not that I think AI is conscious or will be in the near future, but guy who doesn't understand consciousness claims to know it when he sees it.

Until we understand consciousness (which we don't) there is no way to detect the difference between a conscious entity and an algorithm trained to behave like one.

franzcoughka 4 days ago |

[dead]

glimshe 4 days ago |

[flagged]

chairhairair 4 days ago |

It always seemed embarrassing to me that Microsoft hired this guy as if he is some expert in anything.

bondarchuk 4 days ago |

It is really quite staggering how many commenters here seem to believe that the fact that there is no state carried from one chat session to the next proves that there can't be consciousness. It is obviously wrong if you think about it for a minute but I guess people just aren't used to thinking about the relation between the physical and the mental with any level of seriousness.

randomImmigrant 4 days ago |

If AI is conscious, then Pluto is a planet, the Sun is a galaxy, and a black hole is a star.

I’m glad to see someone in a position of any power in the AI world state baldly that AI isn’t conscious. There are times when it feels like we’ve reached complete delulu land on this topic, so it’s a breath of fresh air to see someone not dance around this.

None of this means artificial consciousness cannot be achieved. But the way we’re reacting to these models is proof, from a natural experiment, that a conscious machine should not exist, and certainly shouldn’t be produced as a utilitarian tool that is sold for profit!