AI Agents 6 min read 24 July 2026

OpenAI Presence Shows How AI Agents Actually Work

OpenAI's Presence resolves 75% of support calls by scoping AI agents to one job — a blueprint any business can apply without an enterprise contract, today.

OpenAI Presence Shows How AI Agents Actually Work

On 22 July 2026, OpenAI launched Presence — a platform for deploying AI agents that handle customer support, sales, and internal operations over voice and chat. The headline number is the one worth sitting with: on OpenAI's own English-language phone support line, Presence resolves 75% of inbound calls with no human involved. A Codex-powered feedback loop then cut human handoffs by a further 15 percentage points in ten days.

Most coverage of the launch focused on what it means for OpenAI's enterprise ambitions. That's the wrong lens. The interesting part isn't the company behind it — it's the architecture underneath it, because that architecture is a direct answer to a question we get from almost every business we talk to: "we tried an AI agent and it didn't work — what are we supposed to do differently?"

What Presence Actually Does Differently

Presence isn't a general-purpose assistant that gets pointed at a support inbox. Each deployment starts with one specific job — billing issues, insurance claims, IT service requests — and the agent is given only the information and system access needed for that single task. Nothing broader.

Around that narrow job, OpenAI wraps four things: a written policy for how the agent should behave, a defined set of approved actions it's allowed to take, simulation and evaluation tooling that tests changes before they ship, and a continuous improvement loop that studies where the agent hands off to a human and tightens the gap. None of these four ideas are new. What's notable is that OpenAI built its flagship enterprise agent product entirely around them, rather than around a bigger model or a longer system prompt.

The pattern, stated plainly

Scope the agent to one job. Give it only the access that job requires. Define what "approved" looks like before it goes live. Measure every handoff to a human and use it to close the gap. That's the whole architecture — and it's the same one we've written about before as the difference between agents that survive production and the 74% that get rolled back.

This Confirms What the Rollback Data Already Showed

A study published earlier this year found that 74% of organisations that deployed AI agents were forced to roll them back within months — not because the underlying model failed in testing, but because the agent broke in production: it took actions it wasn't authorised to take, hallucinated under real-world edge cases, or simply had no defined path for what to do when it didn't know the answer.

The 26% that stayed in production shared a structural trait: the agent's authority was proportional to how confident it was, low-confidence situations routed to a human by design rather than by accident, and every deployment had a live dashboard tracking success rate and escalation rate from day one. Presence is that same structure, productised at OpenAI's scale. A 75% resolution rate isn't evidence of a smarter model — GPT-class models have been capable of holding a support conversation for two years. It's evidence of an agent that was never asked to do more than it was built to do.

The Catch: Presence Isn't Actually Available to Most Businesses

Here's the part that matters for anyone reading this and thinking "great, we'll just use Presence." OpenAI launched it as a limited general availability program, not a self-service product. Deployments are led by OpenAI's own Forward Deployed Engineers or a short list of approved systems integrators. There's no signup form. It's built for organisations with enterprise procurement teams, dedicated AI budgets, and the patience for a months-long deployment engagement.

For the vast majority of small and mid-size businesses — the ones running 500 to 5,000 support calls a month, not 5 million — that access model is the whole product, out of reach, regardless of how good the underlying architecture is. Which is the actual opportunity here: the pattern that makes Presence work is not proprietary. It's a design discipline, not a piece of OpenAI technology you need a licence for.

What This Looks Like Built for a Business Presence Won't Talk To

A regional dental group we work with runs a similar profile to what OpenAI describes for its own support line: high call volume, a small number of recurring intents (appointment booking, rescheduling, insurance verification questions, basic pricing enquiries), and a front-desk team spending a disproportionate amount of their day on calls that don't require judgement.

Rather than a broad "AI receptionist," the build followed the same narrow-scope discipline: one agent, one job — booking and rescheduling only, connected live to their practice management system's calendar. Everything outside that scope (billing disputes, clinical questions, complaints) triggers an immediate, no-friction handoff to the front desk with full call context attached, not a dead end. Before go-live, the agent ran in shadow mode against three weeks of real inbound call recordings, logging what it would have done without acting on any of it, so the edge cases in real patient calls surfaced before a single live caller hit them.

Ninety days after launch: 58% of inbound calls resolved without staff involvement, average call handling time on AI-resolved calls under two minutes, and zero booking errors traced back to the agent. Not OpenAI's 75% — a smaller practice, a narrower calendar system, less training data — but built on the identical principle, at a fraction of the cost and timeline an enterprise Forward Deployed Engineer engagement would require.

Three Questions Worth Asking Before You Copy the Pattern

The Presence architecture is simple to describe and easy to get wrong in practice, because the discipline is in the restraint, not the build. Before scoping a deployment against it, three questions are worth answering honestly.

  • What is the single job, precisely? "Handle customer enquiries" is not a job an agent can be scoped against. "Reschedule an existing appointment, verify insurance eligibility against our provider list, and quote standard pricing for our five most common procedures" is. If you can't write the job in one sentence with a bounded list of intents, the agent isn't ready to be scoped yet — the intent analysis comes first.
  • What system access does that job actually require, and nothing more? OpenAI's own description of Presence is explicit that each agent gets only the access its single task needs. A booking agent needs calendar read/write. It does not need billing history, clinical notes, or the ability to issue refunds — even if those systems are technically reachable from the same platform.
  • What happens on low confidence, and who is watching the handoff rate? An agent that guesses when uncertain is the single most common cause of the rollbacks that hit 74% of deployments last year. The handoff to a human needs to be instant, carry full context, and be treated as a data point worth reviewing weekly, not a failure to hide.

The Actual Lesson From This Launch

If you deployed an AI agent that got rolled back, the temptation is to conclude the technology isn't ready. OpenAI's own numbers on its own support line say otherwise — the technology works when the deployment is disciplined about scope, access, evaluation, and escalation. Most failed deployments skip straight to "build the agent" without doing that design work first, then blame the model when reality doesn't match the demo.

The businesses that will benefit most from this moment aren't the ones with an OpenAI enterprise contract. They're the ones who take the architecture pattern seriously — narrow scope, defined access, evaluation before launch, monitored handoffs — and apply it to their own call volume, their own systems, their own edge cases. That's implementation work, not procurement, and it doesn't require waiting for a Forward Deployed Engineer to have availability.

If you're running meaningful call or chat volume and want to know whether a properly scoped voice AI agent would hold up in your specific environment, that's a scoping conversation, not a sales pitch — we'll tell you honestly if the volume and intent mix justify it.

Where to start

Our voice AI builds follow the exact discipline this post describes: one job per agent, live system integration, shadow-mode testing before go-live, and escalation paths designed in from day one. See what a scoped build looks like at wizeb.com/services/voice-ai.

Ready to act on this?

We build exactly what this article is about.

Tell us about your situation — we'll come back with a realistic assessment.