Fewer steps synthesize faster; more sound better. The real value comes from the server as soon as the call connects.
Runs a full turn โ retrieval, prompt, model, speech โ with the same trace as a spoken question.
Spoken commands
Conversation
Nothing yet. Press the microphone and speak.
Every step the pipeline took, in order, with the data behind it. Click a step to see what it actually sent or received.
No turns yet. Ask something and the whole chain appears here โ what was heard, what was looked up, what was sent to the model, and what came back.
Last turn
- Time to first audio
- โ
- ASR
- โ
- Agent
- โ
- TTS first chunk
- โ
- Realtime factor
- โ
- Turns
- 0
Session
- Median TTFA
- โ
- Audio in / out
- โ / โ
- Barge-ins
- 0
Everything the agent is allowed to answer from. Search it here to see exactly which records a question would pull in โ before you ask it.