Why a one-size-fits-all approach doesn't work with voice AI models
The newest voice models, such as OpenAI's GPT-Live-1, are getting better fast. The real challenge is choosing the right one for the calls customers actually make.

The best voice experiences are custom-built
Every voice AI demo sounds great through a laptop speaker. Wideband audio, a quiet room, a presenter who doesn’t interrupt. That same model then goes on a production phone line, audio is squeezed down to 8 kHz, the network adds another 100 to 200 milliseconds, and a real caller reads out a 22-character serial number with four consecutive zeros and cuts in halfway through the answer.
That gap is the primary driver behind voice remaining as the channel most enterprises haven't automated yet. The difficulty is not in getting a voice model to sound good in a demo; it is making it work when a caller interrupts, reads a long identifier, shifts between languages, or needs the agent to take action in their account. That's why we stopped trusting published benchmarks, including those measured through a vendor's own app, and built a way to test voice models in the ways customers experience them: on a phone call. The newest model to go through that test is OpenAI’s GPT-Live-1.
How Maven approaches the model choice
Maven uses a cluster of cutting-edge models to power customers’ voice experiences. In order to provide the best experience for each customer's unique use case: account lookups, complex identifiers, multi-step diagnostics, interruptions, linguistic requirements, deployment specifications, and requests that require an action, a model that works flawlessly for one customer may have marginal results for another. So, we test models to identify the best one for each customer’s individual use cases and their unique criteria. We use those results to decide which model should run the live conversation, where another model should help, and which safeguards belong in the experience.
Model selection is only the first decision. The same evaluation framework also tells us where to add confirmation for complex identifiers, guard against dead air in a conversation, or route a more difficult task to a stronger reasoning layer. The goal is not to pick a single model to serve all use cases. It is to deliver the best experience possible for our customers’ customers.
Evaluating the best models in the Arena
The Voice Arena is Maven’s repeatable model evaluation framework. It puts candidate models through real outbound calls on the public phone network, then scores what happened from the customer’s point of view. We borrow the discipline of blind, pairwise evaluation and adapt it for telephony.
Every evaluation starts with the same scorecard:
- Conversational fluency: how natural and clear the agent sounds.
- Interruption recovery: whether the agent stops, understands the new request, and responds to it.
- Identifier capture: whether it can hear and confirm the strings callers actually read aloud.
- Time to useful answer: the wait time for a substantial response, not just an initial acknowledgment.
- Call reliability: whether the agent completes the interaction without missing requests or avoidable handoffs.
We then add the scenarios that matter to that specific customer’s voice experience. A company whose support line handles complex product serials requires a different stress test from one that deals with refunds, appointments, or billing questions. The framework remains consistent, while the scorecards are relevant to the specific deployment.
Where OpenAI’s GPT-Live-1 stood out
GPT-Live-1 was particularly strong in an aspect that most people notice first: turn-taking. The model is able to process speech continuously, which lets it respond when a caller speaks over the agent instead of waiting for a separate turn detector.
We put OpenAI’s newest voice model to the test for one of our Fortune 50 customers with millions of calls annually against their support and sales use cases. Their evaluation was tailored to models’ ability to work with complex product identifiers and long diagnostics, and GPT-Live-1 led the models we tested on several key measures:
- 90% exact readback on their service tags, the highest rate in the test.
- 100% interruption recovery; it addressed the new question after every interruption.
- 1,257 milliseconds median response to the new request, the fastest on the test bench.
- 1.7 seconds median answer time, with 81% of answers arriving within two seconds.
The model did not lead every scorecard. Exact readback on long serials was 60%, with repeated zeros its main weakness. And 38% of answers made callers wait more than five seconds for something substantive after an initial “Sure” response. Those are not reasons to dismiss the model. They are the kinds of production seams the Arena is designed to expose.
GPT-Live-1 is a better lead than one model for every call
In this case, GPT-Live-1 is a compelling choice to lead the conversation: it is fast, responsive, and unusually reliable when a caller interrupts. For identifier-heavy or reasoning-intensive moments, the system may need another layer. That’s why they chose Maven. We build the voice experience around the evidence, rather than asking every customer to inherit one model’s tradeoffs.
Help us raise the bar for voice
Customers: bring us the calls that would break a demo: complex identifiers, repeated interruptions, multi-step tasks, or anything that cannot end in a deflection. We will evaluate each model against the conditions your callers actually face.
Model providers: if you want to see how your model holds up on real enterprise calls, we want to put it through the Arena. The better the field gets, the better the voice experience becomes for every customer.
Don’t be Shy.
Make the first move.
Request a free personalized demo.



