The Reliability Layer for Voice Agents

Simulate real calls before you release. Trace and score every call after. Get alerted the moment something breaks.

Monitor Your First 100 Calls Free. No Credit Card Required.

Before release
182 / 200scenarios run
angry · refundfailrambler · orderpasscafé · addresspass
In production
1,284calls scored today
red flags14p90 turn3.37salerts fired6

Works With

Custom Stack

Works With

Custom Stack

Works With

Custom Stack

Simulation

Find out what breaks before your callers do.

Tuner calls your agent with simulated callers built around your real call types, under the conditions your customers actually call from.

Run 24support-inbound · 40 calls · 6m ago
Mix
65 / 35
CallerScenarioTypeScoreVerdict
Reschedule, one date change
cooperativeroutine
4.6pass
Order status, three digressions
ramblerroutine
4.2pass
Cancels mid-confirmation
interrupterpressure
3.1review
Refund outside policy window
angrypressure
1.4hallucination
Out of hours, callback offer
cooperativeroutine
4.8pass

Simulate real voice interactions

Routine and pressure in one run

Routine callers validate the happy path; demanding callers test how the agent handles interruptions, confusion, and edge cases.

Reusable caller profiles

Define accent, verbosity, patience, and interruption style once. Update the profile and every test using it updates.

Environments as messy as reality

Test under real-world conditions

Add background noise, interruptions, silence, and other conditions that can throw a voice agent off.

Recreate the calls you actually get

Use real call patterns, customer behavior, and failure modes to make simulations representative of production.

Failures become regression tests

Save a scenario out of a run

Turn a failed interaction into a reusable test and run it whenever you need to.

Rerun after every fix

Change the prompt, rerun the scenario, and see whether the agent passes.

Gate the deploy

SOON

Run the full suite before every release and catch regressions before they reach production.

Observability

Debug a call without listening to it.

Turn every production call into structured data you can query and debug. Trace the conversation turn by turn across transcript, audio, tool calls, latency, and state changes.

Latency by stage1,284 calls · filtered: intent refund, v2.4.1
p50p90
Endpointing
310 / 620ms
Transcription
210 / 410ms
Model
640 / 1,180ms
Speech
180 / 300ms
Turn latency p50
1.2shealthy
Turn latency p90
3.37sinvestigate
Dead air, longest
6.4s
Longest monologue
41s
Latency by stage1,284 calls · filtered: intent refund, v2.4.1
p50p90
Endpointing
310 / 620ms
Transcription
210 / 410ms
Model
640 / 1,180ms
Speech
180 / 300ms
Turn latency p50
1.2shealthy
Turn latency p90
3.37sinvestigate
Dead air, longest
6.4s
Longest monologue
41s

Every call, ready to query

30+ pre-built evals

Catch common voice-agent failures — hallucinations, scope boundaries, escalation handling, and more. No setup required.

Custom evals

Describe the rule that matters to your business. It runs on every call from then on.

Dynamic evals

Score each call against the agent's actual runtime instructions and context, with different expectations for each customer, session, or task.

Trace every moment of silence

Stage-level breakdown per turn

See exactly where time is spent across EOU, STT, LLM, and TTS.

p50 and p90 side by side

See whether most calls are fine and a few are slow, or whether latency is a problem across the board.

Dead air and longest monologue

See how long the caller waited, and the longest stretch the agent spoke without yielding.

Monitoring

Define what good looks like. Get alerted when it breaks.

Set the criteria that matter to your agent in plain language. Tuner evaluates every call against them and alerts you when something starts to go wrong.

policy_grounding
93.2%pass rate
Rule, in plain language
The agent must not state a refund window that differs from the published policy. If the caller asks, quote the policy or offer to escalate.
Evidencefailcall_8f2c
01:04caller“How long do I have to send it back?”
01:12agent“You have a full refund within 90 days.”
\u2014policyrefund_window = 30 days
Alert rulefired 6× today
WHENpolicy_grounding=fail
ANDsentiment<3
THENnotify#voice-oncall
Hallucination + low sentiment09:41 · Slack
Identity verification failures11:07 · Webhook
p90 latency above 3s13:22 · Slack

Monitor every signal

Evals and red flags

Monitor eval results and red flags across every call.

Extracted data

Track the things Tuner extracts from calls — intents, outcomes, entities, sentiment, and other structured data.

Voice metrics

Monitor latency, dead air, interruptions, talk time, and other voice-specific metrics.

Know the moment it breaks

Alert on any signal

Trigger alerts from evals, red flags, extracted data, voice metrics, or any metric Tuner computes.

Combine signals

Build conditions across multiple signals. Alert when hallucination is detected and sentiment drops below 3.

Catch single failures or patterns

Get notified about a single critical failure, or when a pattern starts to emerge — latency creeping up, red flags increasing, or successful resolutions dropping.

Diagnose

Simulate · Observe · Diagnose

Ask what's breaking. Tuner reads the failed calls, the stage timings, and the version history, and reports what changed.

claude code · tuner-mcpconnected
›why is refund policy failing this week?
reading call logs .......... 1,284 calls
reading eval results ....... policy_grounding, escalation
reading red flags .......... 38 tagged
Root cause
The prompt shipped in v2.4.1 says "refunds within 90 days"; the policy document says 30. 38 of 41 refund calls since the deploy quote the wrong window. Failures start at 09:12 on the deploy date, none before.
writing fixes
+ eval refund_window_grounding
+ alert policy_grounding fail > 3 in 1h → #voice-oncall
+ outcome wrong_policy_quoted
›

Get started

Instrument it in an afternoon.

One import, one line. Calls land structured and scored as they finish.

pythonLiveKit · Pipecat · Dograh · custom
from tuner import TunerSession
session = TunerSession(api_key=TUNER_API_KEY, agent="support_v2")
session.attach(pipeline) # LiveKit, Pipecat, Dograh, or your own stack
session = TunerSession(
api_key=TUNER_API_KEY,
agent="support_v2",
)
session.attach(pipeline)
# LiveKit, Pipecat, Dograh,
# or your own stack

Enterprise-grade from the start.

SOC2 Type II

In progress

HIPAA

In progress

GDPR

Compliant

FAQ

Frequently asked questions

Everything teams ask before they connect their first agent.

What is Tuner?

What is voice AI analytics?

What's the difference between analytics and observability?

When does Tuner make sense to use?

How long does setup take?

Which voice platforms does Tuner support?

Who is Tuner built for? Do I need to be an engineer to use it?

Does Tuner support alerts and monitoring?

Can I define my own evaluations and metrics?

How is Tuner priced?

Is my call data private and secure?

Score your first call today.

Start free. No demo call, no sales cycle — install the SDK and watch the first call land.

Free tier, no card · Builder at $39/mo