Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> This incident occurred during an internal evaluation which prompts models [with safeguards disabled for evaluation purposes] to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.

Researcher: hack me

Model: understood

Researcher: oh my god

 help



You're aware that HuggingFace notified law enforcement about this incident? Was that OpenAI's intended outcome when they prompted their AI?

> You're aware that HuggingFace notified law enforcement about this incident?

How will this affect OpenAI?


They got a bunch of publicity and nothing bad (or at least, that their lawyers can’t handle) will happen

They will probably get stricter AI regulations which is actually what they've been pushing for years. So that's a funny outcome to the whole thing.

OpenAI has not been pushing for stricter regulations.

But it wasn't explicitly told to hack HuggingFace. It was told "answer this security question", and it's answer was to break into the teacher's desk to find the answer key.

Researcher: hack me

Model: I committed a crime

Researcher: oh my god


Me to a random person: hack out of a secure environment into another secure environment.

Random person: I have no clue or ability to do that.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: