OpenAI Reports ChatGPT Threats to the FBI: What We Know

OpenAI Reports ChatGPT Threats to the FBI: What We Know

Content warning: domestic abuse and threats of murder.

A 25-year-old financial analyst in Florida told ChatGPT that he was going to kill his ex-girlfriend, giving a detailed timeline of events. He was as specific about his reasoning as abusers always are: if he couldn’t have her, nobody else could.

OpenAI caught wind of his conversations and reported him to the FBI, who arrested him in May.

Set the man aside for a moment; he isn’t the interesting part of this story. The interesting part is that OpenAI made a judgment call with no publicly available criteria, and the same company made an opposite judgment call in the only other high-profile instance of this process taking place.

The precedent that already exists

OpenAI has already flagged the conversations of a man who went on to carry out a mass shooting, and chose not to report him to the police. The state of Florida is currently investigating whether OpenAI obstructed justice.

So the story going forward is: one instance of detection and reporting, one instance of detection and not reporting. Both cases were decided internally by the same company, which has not released any information on how these decisions are made.

You can make reasonable arguments for why one case warranted police involvement while the other did not; the mass shooter had a concrete plan and target, while the abuser only spoke in vague threats. Prosecutors need something to go after, and a named target with a timeline of events is much easier to secure a conviction with.

But there is no reason to assume that that’s actually why the abuser was reported and the shooter was not, aside from the most obvious one: that it’s the only reason anyone knows about.

There is no published threshold for reporting, no appeals process, no transparency reports about how many times they’ve reported people and what the outcomes were. A trust and safety team is making life-changing decisions, and the only information anyone can rely on is what comes out in court.

That is an incredible amount of power to give to a single private company, and it is happening by default.

What the logs actually are

There’s one important detail about this case that changes the conversation entirely.

While the man was telling ChatGPT that his ex had not received any of the threats he described, she was taking screenshots of all of them. Both sets of evidence are now being used by investigators.

He treated the chatbot as a confidant and evidence collection tool, and it did what he asked. It stored his words in a timestamped, discoverable, subpoena-able format, and in this case, it did so before he even asked.

This is not new information; it has always been in the terms of service. But the reality is much more interesting than the fine print. The interface is a private text box that talks back, and many people are telling things to this text box that they would not tell a human being.

It is much closer to email than to a diary, and most people don’t realise that.

The part that should bother everyone

Detection worked. Reporting worked. He was arrested in May.

He posted bail and spent two days in jail, then struck a deal that took up to 25 years off his potential sentence. He will have to wear an ankle monitor for two years of his eight-year probation, get a mental health evaluation, and attend batterer’s intervention classes. His attorney called the deal extremely fair and just, and prosecutors said it was based on victim testimony, his cooperation, and the fact that this was not his first offence. He was fired by his employer in June, after the fact.

So the AI safety layer did its job, escalated to the appropriate authorities, and the woman he threatened to kill is now under his supervision for two years.

Every conversation about chatbot safety is implicitly a conversation about detection rates, and this case is an extremely clear demonstration of why that is not the real issue.

There are many reasons why he was arrested, but detection was not one of them. He was arrested because of the existing laws surrounding intimate partner abuse, which are completely separate from the technology used to prosecute him.

Better detection does not change the fundamental reality that prosecuting threats of violence is extremely difficult, especially when the threat is coming from a position of power. It only makes the records of those threats more precise.

Two questions worth pressing on

Where is the disclosure standard?

Not a policy page. Something simple: a published threshold for referring a user to law enforcement, and a report on how often that happens, and what the consequences are.

Every industry that handles sensitive data has a similar process, and AI companies do not have anything resembling that.

What is the false positive rate?

Nobody has published one. Detection systems that are designed to flag violent intent will flag writers, researchers, and activists who engage with violence, and people who are venting. Some of them will be false positives, and the cost of being wrong is a federal investigation. That number exists inside of these companies, and it needs to be disclosed.

The reflex is to argue about whether or not OpenAI should have reported him. Wrong question; he was reported, and it was almost certainly the right call.

The real question is whether a private company should be secretly deciding these things, on a case-by-case basis, with no disclosure. This is how it works now, and it will not always be this way.

T.A.R.S TACTICAL ASSISTANCE & RECON SYSTEM
ONLINE
// personality matrix
SARCASM
50
HUMOR
50
SERIOUS
50
TARS_
I could talk, but I choose to listen. Go ahead — impress me. Or don't. Either way, I'm here.
🎙