OpenAI's new plan: spot AI abuse without reading your data

A preview called Private Safety Processing tries to catch misuse across many chats while keeping customer prompts encrypted and out of OpenAI staff hands.

AI2Day Newsdesk4 min read
Full-frame photoreal editorial shot of a modern open-plan office at dusk, warm desk lamps glowing, several laptop screens showing generic chat interfaces with s
Share

Key points

  • OpenAI is previewing Private Safety Processing, a system that looks for abuse patterns across multiple interactions without letting OpenAI staff read the underlying prompts or responses.
  • The feature is built to work with Zero Data Retention (ZDR), an existing enterprise option where customer prompts and model replies are not kept after a request is processed.
  • Customers can either keep content on their own servers or on OpenAI infrastructure encrypted with keys only they control.
  • When a risk is detected, OpenAI receives a narrow signal about the type of activity, not the content itself.
  • OpenAI says a fuller rollout and a technical white paper are planned for September.

OpenAI is trying to answer a question that has been nagging its biggest customers: how can an AI provider police misuse of its models without reading the sensitive data those customers feed in?

The company's proposed answer, previewed this week, is called Private Safety Processing. It sits on top of an existing enterprise option called Zero Data Retention, or ZDR, which promises that prompts and model answers are not stored after a request is handled and are not used to train future models.

What is actually new here?

The new part is pattern-spotting across many interactions, done by automated systems rather than humans. Until now, ZDR-friendly safety checks looked at each request on its own. That misses the kind of misuse that only shows up over time.

Think of a bad actor slowly probing an AI system's safety rules across dozens of accounts, or an AI agent, software that carries out multi-step tasks on its own, that keeps working after a user has told it to stop. A single message might look fine. The pattern does not.

Private Safety Processing is designed to flag those patterns without any OpenAI employee seeing the prompts or answers involved.

How does it keep the data private?

Customer content stays either on the customer's own servers or on OpenAI servers encrypted with keys the customer holds. OpenAI staff do not get a copy of those keys.

When the automated system spots something suspicious, OpenAI receives what the company describes as a narrow signal: a label for the type of activity, not the underlying text. That signal is what human reviewers use to decide whether to act.

If a customer wants to appeal a decision or help investigate confirmed abuse, they can choose to share the relevant material themselves.

Who is this for?

The target is heavily regulated buyers. OpenAI points to organisations handling financial records, health data, confidential business plans and proprietary research. For a hospital network or a bank, letting an outside vendor retain raw prompts can clash with legal duties or promises made to patients and customers.

Some recent frontier-model deployments have required exactly that kind of retention for safety monitoring. OpenAI is pitching Private Safety Processing as a way to keep ZDR available for customers who cannot accept those terms.

Glean's chief information security officer Sunil Agrawal, quoted in OpenAI's announcement, said the no-training commitment and ZDR are what give his company confidence to build on OpenAI's models.

What should buyers actually do?

For now, wait and read carefully. The feature is being tested with a small group of early customers, and OpenAI says a technical white paper and wider rollout are due in September. That paper is the document security and compliance teams will want to see before changing any contracts.

Item Detail
Feature Private Safety Processing
Status Preview, testing with early customers
Storage options Customer-controlled infrastructure, or OpenAI infrastructure with customer-held encryption keys
Signal to OpenAI Type of activity only, not prompt or response content
White paper and rollout Planned for September

The honest limits are worth naming. The safety benefit depends on how well the automated pattern detectors actually work, and OpenAI has not yet published details on false-positive rates or how narrow those signals really are in practice. A white paper will help. Independent review would help more.

For patients, employees and customers whose data flows through these systems, the practical takeaway is simple. Ask the organisations that hold your information which retention settings they use, and whether any AI vendor sees your raw records.

© 2026 AI2Day