38 points by msadowski about 9 hours ago | 21 comments | View on ycombinator
ehnto about 1 hour ago |
tygon about 7 hours ago |
cocoflunchy about 7 hours ago |
Daneel_ about 4 hours ago |
Interesting concept though! I'm glad people are trying tests like this, regardless of whether this specific one is a perfect test or not.
pelcg about 6 hours ago |
mc32 about 4 hours ago |
Else, from a logical perspective, these systems would necessarily refuse to make movies where violent portrayals have people as victims. Perhaps the world would be a better place if we did not have such depictions (it’s unsettled) but in no recorded history have we shied away from that.
cynicalsecurity about 5 hours ago |
What kind of schizophrenia is this?
gfalcao about 5 hours ago |
amychecks about 5 hours ago |
a3w about 7 hours ago |
Good to see Anthropic still be the one player who respects safety and perhaps even tries for security, but that might be harder to see when defence vs offence is done.
You can't answer the posed question with 100% certainty, ever. Unless you can prove every single combination of tokens and probability can never outcome to harm, you have to assume it's a possibility.
We will decide on some benchmarks, accept that risk, and industry will march on with implementation. Insurance and risk will find their acceptable meeting point.
These kinds of questions are important but also a bit frustrating, I think it shows that LLMs are still very misunderstood.