scenarios: list. A scenario with a turns: key is scripted; one with a persona: key is a simulated scenario, where an LLM plays the user instead. This page covers the full scripted format. If you haven’t run a scenario yet, start with the quickstart.
Use a scripted scenario when you want exact control over the user’s side. You know what the user says on every turn, so you can assert exactly what the agent must do: call this tool with these arguments, say this, answer within this budget, recover from this interruption. The input is the same every run, so a failure is easy to reproduce and fix. Use a simulation when you want to check a goal instead, and let the caller adapt to the agent.
Anatomy of a scenario
user:) and lists the events expected in response (expect:). Expected events must arrive in the order listed, but the agent may emit other events in between, so you don’t have to enumerate everything it does.
The scenario runs as multi_turn/multi_turn, its file’s name and then its own. A file can hold several scenarios, and any scenario key at the top of the file, like the judge: above, is the default for all of them. See Several scenarios in one file.
The rest of this page is in four parts:
Configuration
The shared
user: and judge: blocks, plus the scripted-only context:
and stop_on_failure:.User turns
Drive each turn with an utterance, keypresses, an image, or timing.
Events
The semantic events the agent emits, and what each one means.
Assertions
Check an event’s content or timing with
eval:, text_contains:, and more.Configuration
The file format, theuser: and judge: blocks, factory:, !include, running scenarios back to back, and the disconnect path are shared with simulated scenarios and documented once, on Scenario Configuration. A scenario with none of those blocks runs entirely in text mode with the default judge, which is the fastest way to start.
Two fields are scripted-only, since a simulation has no scripted turns to seed or to score. Like any scenario key, both may sit on the scenario or at the top of the file as a default for every scenario in it:
Seeding the context with context:
By default the harness leaves the bot’s LLM context alone: whatever the bot sets up for itself (for example, a system prompt added in its connect handler) is what the scenario runs against. Provide context: to replace that with messages of your own, which lets a scenario start mid-conversation:
LLMMessagesUpdateFrame that replaces the bot’s context wholesale. Omit context: and the harness sends nothing, leaving the bot’s own context in place.
Scoring every turn
By default the first turn with a failed assertion ends the scenario, since a conversation that has gone wrong rarely tells you much about the turns after it. Setstop_on_failure: false when the turns are independent and you want a score across all of them, for example when benchmarking intent classification over a list of utterances:
User turns
Each turn drives the agent by speaking (auser: utterance) or pressing keys (a dtmf: sequence); the two are mutually exclusive. A turn can also register an image:, or be observation-only with no input. send_after: controls when the input is sent.
Utterances with user:
Each turn’s user: field is the user’s utterance for that turn, a plain string. You write it the same way in both modes; whether it’s delivered as text or synthesized into real speech is set once by the user: block, not per turn.
A turn without a user: field is observation-only: the harness just waits for the expected events. This is how you test agent-first behavior like an on-connect greeting:
Playing audio files with audio:
In audio mode, a turn can play a recording instead of synthesizing its user: text. The audio: field (a path relative to the scenario file) names the audio file to stream to the agent:
soundfile reads works: WAV, MP3, FLAC, OGG. Multi-channel audio is downmixed to mono. user: is required alongside audio: and gives what the recording says, since the judge and text_contains see it as the turn’s input.
A scenario whose spoken turns all name an audio: file needs no user.speech: block, since nothing is synthesized.
DTMF keypresses with dtmf:
Instead of a user: utterance, a turn can press phone keypad keys with dtmf:. The two are mutually exclusive: a turn either speaks or presses keys. This drives keypad menus (IVR) and any agent that reacts to telephony tones:
InputDTMFFrame, the same path a telephony transport’s keypress takes, regardless of the scenario’s user:/judge: modality. Valid characters are the keypad entries 0-9, *, and #; any other character is a parse error.
A bot running a DTMFAggregator accumulates the keys and flushes them into a DTMF: ... transcription, which (with the default transcription-based turn-start strategy) drives a full user turn: user_started_speaking, user_transcription, user_stopped_speaking, and the agent’s response. So a dtmf: turn can assert on user_transcription and response just like a spoken turn.
The aggregator flushes either on the # terminator or on its idle timeout. To exercise the idle-timeout path, omit the # and pace the keys with a time-based send_after::
expect: is optional on a dtmf: turn: omit it for a turn that only presses keys, with the assertion living on a later turn.
Vision with image:
A turn may register an image with image: (a path relative to the scenario file). When a vision agent requests a user image during the turn, the eval transport serves it:
Scheduling with send_after:
By default, a turn is sent once the agent has finished speaking, like a caller who waits for the sentence to end. This ensures a reply the previous turn was satisfied with early is never talked over. A turn may override this with send_after:, which controls when the input (its user: utterance or dtmf: keypresses) is sent relative to a prior event or after a plain delay. Anchoring it to an event is how you script barge-in tests:
event: anchor is optional. A bare send_after: { delay_ms: 500 } is a pure time delay measured from the previous turn’s send, with no event to wait on. This is handy for pacing turns by time rather than off a bot event (for example, spacing out DTMF keypresses to exercise an aggregator’s idle-timeout flush):
send_after: with no event: and a zero delay_ms is rejected as a no-op: give it an event:, a positive delay_ms, or both.
Events
Scenarios assert on a small set of semantic events, mapped from the RTVI messages the agent emits:Assertions
Each entry inexpect: names an event and, optionally, asserts on its content or timing.
Semantic judging with eval:
The eval: field is a natural-language criterion that the event’s text must satisfy, decided by the judge LLM:
eval: only makes sense on the agent’s text output (response, llm_response, tts_response).
Substring checks with text_contains:
For exact content, text_contains: does a substring check, ignoring whitespace differences:
response, llm_response, and tts_response, the harness accumulates successive segments and re-checks on each new segment until the check passes. On user_transcription, it accumulates the STT’s final transcription segments within the turn, so an STT that finalizes an utterance in pieces can still satisfy a phrase that spans them.
text_excludes: is the mirror: the event’s text must not hold the substring. It fails with the kind text_present. Its use is a marker the LLM let slip into its reply, which would reach the user:
Asserting nothing arrives with absent:
absent: true inverts an expectation: no event of that type may arrive before its within_ms budget expires. It matches on the event type alone, so it can’t be combined with text_contains:, eval:, or calls:. Set within_ms explicitly, since the default budget makes the quiet window a full minute. A response that continues the reply an earlier expectation matched is not a new one; only a reply the agent began after that match counts. This is the check for a duplicate reply, or for an agent that should hold its turn:
Latency budgets with within_ms:
within_ms: bounds how long after the turn’s user send the event may arrive. All of a turn’s expectations share that one anchor:
--timeout), so timing is only asserted when you ask for it.
Because every deadline is measured from the send, time spent matching earlier expectations counts against later ones. In the example above, if llm_started arrives at 1.5 seconds, the response (with the default 60 second budget) has 58.5 seconds left, and a turn that stalls completely fails within a single budget rather than one per expectation.
Function calls
Afunction_call expectation asserts that the turn invoked one or more tools. List the expected calls under calls:; they’re matched by name in any order, and the expectation passes once all are found:
args is a subset check: every listed key/value must be present in the call’s arguments, and extra arguments are ignored. A single expected call can use the name:/args: shorthand directly on the expectation, and a bare function_call with neither just asserts that some call happened.
Arguments take part in the matching, so the turn is satisfied by any call matching both the name and the arguments. A call the model gets wrong and immediately repeats correctly still passes. When nothing matches, the failure names the arguments that did arrive.
Judging a call with eval:
args: matches verbatim, which is no use for an argument the model phrases in its own words. A function_call expectation can carry an eval: instead, or as well: each call the expectation matches is put to the judge LLM by name and arguments, over the conversation so far, under a judge prompt of its own. A rejected call fails the turn with the kind judge_no:
function_call_stopped carries no arguments, only how the call ended, so it takes no eval:.
function_call_stopped takes the same calls: shape and reports a call ending, which is how a scenario asserts that a cancellable tool was actually stopped:
Turn-completion markers with llm_marker
An agent that filters incomplete user turns has its LLM open every response with a marker saying whether the user’s turn was complete. The llm_marker event reports the marker the LLM produced when the response ends, and marker: names its meaning: complete (the turn was finished and the agent answers), short (the user was cut off and the agent waits), long (the user asked for time), or incomplete for either of the last two. It checks the marker’s meaning, not its text, so a scenario holds whatever marker characters the agent configured. A bare llm_marker asserts only that the agent read a marker:
marker_first: asserts that nothing comes before the marker, markers: how many markers the text holds, and text_after: whether text follows the first marker, which a complete turn should have and an incomplete one should not:
text_excludes: on the reply. In audio mode, anchor the follow-up turn on vad_user_stopped_speaking rather than user_stopped_speaking, since turn detection defers the latter while the turn is held open.
A run’s result records what each expectation matched, the marker an llm_marker saw included, so a passed run keeps the marker it read. See results.jsonl.
Several scenarios in one file
A file’sscenarios: list can hold several scripted scenarios, and any key at the top of the file is the default for all of them. That suits testing one behavior through many short conversations: the file sets the judge: and the context: once, and each scenario is a conversation of a few turns. Each runs on its own, against its own bot, as turn_completion/short_answer and turn_completion/cutoff here, and a suite’s -s turn_completion runs them all:
turn_completion.yaml
turns:, context:, and stop_on_failure:, and the shared user:, judge:, and trigger_disconnect:. A scenario that sets the same key replaces the whole value, so a scenario’s own context: is written out in full, never added to the file’s.
Sharing turns: is for a file whose scenarios hold the same conversation and differ in one thing only. With turns: at the top and one scenario per judge or modality, every judge sees exactly the same conversation, which is how Pipecat’s interruption scenario runs in text and in audio:
interruption.yaml
Next steps
Simulated Scenarios
Hand the user’s side to an LLM with a persona and a goal, and judge the
whole conversation.
The Eval Loop
Let a coding assistant write agent code, run evals, and iterate
automatically until the agent is better.