Never Stall: A Conversational Call Engine for Filmed Branching Narrative
A filmed character calls you and the story branches on what you say out loud. An engineering report on the route-or-clarify call engine behind it.
Engineering report, published 1 October 2026. Describes the conversational call engine shipped in Treezy Play on 12 July 2026 (backend and Android) and 15 July 2026 (iOS 1.6).
Abstract
Treezy Play is a mobile app in which a filmed character telephones the viewer and the story branches on what the viewer says out loud. Because the branch point is a video call, the interaction carries a constraint menu-driven branching video does not: a filmed clip runs on a clock and cannot wait politely for a decision. Playback must never stall, and it must never look like the character stopped listening.
Speech is transcribed on the device and endpointed locally into complete utterances. Each finished utterance, with the running turn history, is posted to one authenticated backend endpoint, which asks a small language model for exactly one forced tool call returning either route, meaning the utterance clearly selects one of the scene's filmed branches, or clarify, meaning the character says one short line and keeps listening. Routing is declared in the prompt to be a mechanical judgment that overrides the character's personality, because a character given free rein argues with clear answers instead of routing them. Clarifications are budgeted and escalate to a forced route, and every failure path drops to the pre-existing on-device classifier, so the scene always ends on a branch.
1. The problem
Treezy Play plays filmed, branching stories on iOS and Android. The branch point is not a menu. The character calls you, you answer, and you talk. The app's one feature film, The Seminal Works of Dr. Tiffany Belle, is built around a protagonist who phones the viewer repeatedly, and every call forks the story.
As originally built, before I took over hands-on engineering, the mechanic worked the way voice input usually works in a video product: an on-device text classifier returned positive or negative, a literal string match ran against branch labels written for a menu, and a timed default fired when the clip ended. That fails in two ways that look identical to a viewer. People do not say the label; they say "I guess so", or "hold on, who is this", and a binary verdict maps some of that correctly and the rest arbitrarily. People also say nothing at all, and silence produces no classification, so the character appears to have lost interest in her own question. Either way the clip runs out and the default takes the branch, teaching the viewer who spoke that speaking does not matter.
The constraint underneath both failures is worth stating plainly, because every design choice below descends from it. A filmed scene is a fixed-duration asset, with no idle animation to loop. You can hold the last beat of a clip for a moment, and that is all. The system therefore has one non-negotiable invariant: playback must never stall. A better understanding of the viewer's speech is worth nothing if the cost of obtaining it is a frozen frame.
2. Architecture
A turn has three stages. The device captures speech and decides when an utterance is finished, the backend decides what it means, and the device acts, either by cutting to a branch or by showing a caption and continuing to listen. The backend holds no session state: every request carries the scene identifier, the transcript, the turn history and the clarifications already spent, so each turn is independently replayable. The two mobile engines are twins with the same states, nudge ceiling and wire format.
2.1 Transcription on the device, offline preferred
Audio never leaves the phone. On iOS we set requiresOnDeviceRecognition on SFSpeechRecognizer from the recognizer's own supportsOnDeviceRecognition, so local recognition is used wherever the device offers it. On Android we set EXTRA_PREFER_OFFLINE on the recognizer intent, which is best effort: offline coverage depends on a language pack the operating system may not have downloaded, and the recognizer silently falls back to network recognition when it is missing. We document that asymmetry rather than claim parity.
Endpointing is local too. The iOS recognizer accumulates the running transcript and arms a timer for 0.9 seconds of silence following speech; when it fires, the text is delivered as one finished utterance and recognition restarts. On Android the platform recognizer's end-of-speech and no-match callbacks play the same role. What crosses the network is text, not audio: a privacy, bandwidth and latency property at once.
Silence is treated as an utterance rather than an absence. On iOS a 7 second timer, re-armed whenever speech begins, submits an empty transcript; on Android the recognizer's no-match and speech-timeout errors publish an empty string into the same path. The backend reads "" as silence and the character teases the viewer into answering, which is what a person on a video call does when the other side goes quiet.
2.2 The request
Each turn posts {sceneId, transcript, history, nudgeCount} to POST /api/interaction/converse with a 6 second client timeout, deliberately shorter than the patience of the scene. Identity is the signed-in viewer's Firebase ID token in an Authorization: Bearer header, verified server-side first, because the endpoint spends money on every turn and anonymous traffic is better rejected at the door than rate-limited later. The controller truncates history to the last twelve turns and each turn's text to 400 characters, so a client cannot inflate the prompt.
2.3 One forced tool call
The backend loads the scene and its branch choices, builds a single prompt, and calls a small model, claude-haiku-4-5-20251001, with exactly one tool named decide and tool_choice set to force it. The schema is the entire contract: decision, an enum of route or clarify; choiceId, required when routing; reply, required when clarifying, at most 18 words; and confidence from 0 to 1. max_tokens is 300.
There is no prose to parse, no JSON to repair and no second round trip. A response with no tool_use block for decide is a failure rather than a puzzle, and so is a route to a choiceId the scene does not contain: the service returns an error instead of guessing at the nearest label. Being wrong in a way the client can detect beats being plausibly right.
2.4 Mechanical routing, separated from personality
The prompt has two parts that must not blend: a distilled character canon block derived from the film's story bible, which is itself derived from the shooting script, and a set of decision rules that state in the prompt that they override the canon.
They have to. This character is written to flip from flirtatious to furious and to argue, so given the canon alone she treats a clear "no" as an opening position and negotiates, and the story never forks. The rules say that routing is a mechanical judgment about whether the utterance selects one of the branch options, that natural agreement and clear refusal both route, that she never argues with a clear answer because the story does that, and that her personality colours only the wording of clarifications. The canon block is kept compact, because it rides on every turn.
2.5 Nudge escalation, then the default
Every clarify increments a nudge counter the client owns and sends back, and the counter changes behaviour twice. In the prompt, once two clarifications have been spent, the instruction flips from preferring clarification over a low-confidence route to routing to whichever option best fits the conversation so far. The character stops asking and commits, and the canon gives her a way to do that in voice, by blaming the connection rather than the viewer. In the client, a third clarification ends the negotiation: the engine releases its hold on the scene and lets the pre-existing behaviour finish the job, which is the on-device classifier's warm pick or the timed default. The conversation improves on the old path and never replaces it.
2.6 Failure is a first-class path
The failure path is short, total and quiet. Missing sign-in, a token error, a non-200 status, a missing tool block, an unparseable body, an unknown choice identifier, a timeout, a dropped connection: all converge on one callback, which turns conversational mode off, un-ducks the film audio, resumes listening and hands the scene back to the on-device classifier. The viewer sees no error, only a film that keeps playing.
On iOS the feature also sits behind a Remote Config flag, conversational_calls_enabled, so it can be switched off for the entire install base without a release. Android has no such flag and relies on the failure path alone, an asymmetry rather than a design, listed in the limitations.
2.7 Text-only clarifications
Clarification lines appear as captions styled like a video call's, never as synthesised speech. The reason is a consent rule the project set for itself: no performer's voice is cloned, a stock synthetic voice was rejected as character-breaking, and the text-to-speech path that had been built was removed. Rendering reuses the viewer's own caption preferences, the same size, colour and background model that governs the film's closed captions, with one addition: the call caption keeps a floor of translucency behind the text even when film captions are set to a transparent background, so the line stays legible over moving video. A constraint that began as an ethics decision produced an interface cheaper, quieter and more accessible than the synthesised alternative.
3. Design decisions and the alternatives rejected
Each decision below is recorded in our architecture decision records, cited by number and date.
ADR 0003, 13 July 2026. Transcribe on the device, decide meaning on the server. Rejected: a fully on-device intent model, which would need retraining for every scene's branch set and could not write a free-form clarification; fully cloud speech recognition, worse on privacy, latency and offline behaviour alike; and a streaming duplex voice agent, impossible without a consented voice and fighting the film's audio for the same speaker. The record also fixes the invariants both clients keep: token authentication, silence as a turn, roughly three nudges then the default, and a fallback on any failure so playback never stalls.
ADR 0001, 13 July 2026. No voice clone. The text-to-speech path was built and then deliberately removed. A stock voice was rejected because a wrong voice mid-call is more character-breaking than a caption, and the decision stands as a project rule.
ADR 0006, 13 July 2026. In-character prompts derive from the story bible. Any feature where a model speaks as a film character embeds a distilled canon block generated from the bible and regenerated whenever it changes. Rejected: the full bible in the prompt, on cost and latency, since the block ships on every turn, and a fine-tuned per-character model, premature for one character in one film. The record carries the lesson that cost the most debugging time: decision logic must be declared to override personality.
ADR 0007, 13 July 2026. One caption preference model. Film captions are generated with Whisper into WebVTT and muxed into the scene manifests as a subtitle track. Separate styling for call captions was rejected, so a viewer who set large amber captions for the film gets them from the character too.
ADR 0010, 26 August 2026, accepted 1 September 2026. The category is the call, not the format. This reads as positioning and functions as an engineering constraint, ruling out generated video and generated faces: the footage is filmed, the model listens and routes, and nothing it produces is rendered as image or voice. It also rules out game framing, which is why the branch prompt is a question from a person on a call and not a menu with a timer.
4. Measurements
These numbers do not exist yet. The engine shipped with telemetry and a unit-test harness, but no measurement run has been performed, and this section is published with placeholders so the method is reviewable before the results are.
- Median end-to-end decision latency: [P50 latency ms] (population: Turns that returned a decision)
- Tail decision latency: [P95 latency ms] (population: Turns that returned a decision)
- Fallback rate: [fallback rate %] (population: All conversational turns)
- Routing accuracy: [routing accuracy on N labelled utterances] (population: Offline labelled set)
- Clarify rate: [clarify rate %] (population: All conversational turns)
Latency comes from the client, because the number that matters is the one the viewer waits through. Both clients stamp the moment an utterance is submitted and emit a call_turn event carrying latency_ms when the decision arrives, with the decision, the nudge count, a silence flag and the utterance length. Export the events, filter call_turn to decisions of route and clarify, and take the median and 95th percentile of latency_ms. Note that the interval excludes on-device transcription and the 0.9 second endpointing delay, and add both to any figure quoted as mic-to-cut.
Fallback rate has two readings and both should be given. At turn level it is the share of call_turn events whose decision is fail. At call level it is the distribution of the outcome field on the call_ended event, which records routed_engine, routed_classifier, default or declined, so the share of calls that finished on the engine rather than the old classifier falls straight out of it. The call-level number answers the question the invariant poses.
Routing accuracy cannot come from production logs, by design: no transcript text is logged client-side, because the endpoint already sees every utterance and the corpus belongs server-side. It comes from an offline harness. Build a labelled set of N utterances per conversational scene from three pools of roughly equal size, namely clear selections in natural phrasing, ambiguities and questions back at the character, and empty strings representing silence. Label each with the branch a human reader would take, or with clarify where no branch is correct. Drive the exported prompt builder and the real service against the live model with the production canon block, and compare the returned choiceId against the label. Report overall accuracy, the route-versus-clarify confusion matrix, false-route and false-clarify rates separately because their costs differ, and the count of routes to an identifier that does not exist. The backend unit tests, which inject a fake fetch and a fake Firestore, supply the scaffolding; only the fake fetch becomes a live call.
Clarify rate is the share of call_turn events with decision clarify. Two companions make it interpretable: the distribution of nudge_count at the moment of routing, and the frequency of the max_nudges decision, which is the engine admitting it could not resolve the turn. Split all three by the silence flag: a clarification prompted by silence is a success, one prompted by speech the model could not place is closer to a miss.
5. Prior art
Menu-driven branching video. The reference points are Netflix's interactive specials, of which Bandersnatch in 2018 is the best known, and Eko, which built a platform and catalogue for branching video. The affordance in both is a menu: a layer appears over the video, enumerated options are shown, and a tap or remote press selects one within a timer. That absorbs exactly the problem described here by making the wait state explicit and the input space finite, at the cost of the viewer leaving the fiction to operate a control. Both have contracted commercially, Eko closing in 2024 and Netflix retiring most of its interactive titles in 2025, which is one reason we do not present this work as a continuation of that category. Our branch point keeps the fiction: the option set is never shown, and the input is whatever the viewer says to a person on the phone with them.
Language-model branching narrative. WHAT-IF, by Huang, Martin and Callison-Burch (arXiv:2412.10582, posted 13 December 2024, revised 20 October 2025), uses zero-shot meta-prompting to turn a prewritten story into a branching structure, storing generated alternatives in a graph for interactive fiction. The relationship is close in spirit and opposite in direction. WHAT-IF generates the branches; ours are filmed, finite and fixed months before the model sees them, and the model's only job is to decide which of them the viewer just chose. WHAT-IF also operates in text, where the reader's turn has no deadline, whereas ours is bounded by a running clip.
"Facilitating Video Story Interaction with Multi-Agent Collaborative System", by Zhang, Hao, Wang, Sheng and Zeng (arXiv:2505.03807, May 2025), applies a vision language model to video stories and combines retrieval-augmented generation with a multi-agent system so viewers can explore personalised stories with dynamic characters and customisable scenes. The direction again differs: that system's value is in composing new material, while ours renders no generated frame and no generated voice, by decision as well as by capability, under a real-time no-stall invariant on two consumer platforms.
Task-oriented dialogue. Route-or-clarify is not a new pattern: committing to an interpretation or asking a clarification question is the core loop of task-oriented dialogue, and forcing a structured decision through a tool schema is ordinary practice. The narrow claim is the setting, a wait state that cannot be paused, which forces a finite clarification budget, explicit escalation and a fallback that works with the model switched off. We are aware of no adoption of this design outside our own codebase, no third-party evaluation and no citations.
6. Limitations
The measurements in section 4 are placeholders. Until a run exists, every performance claim here is a design intention.
Android has no kill switch and relies on the failure path alone, its offline recognition preference is best effort, and the iOS latency stamp uses a wall clock where a monotonic one belongs. The fallback is the binary classifier that predates the engine: positive or negative is a poor floor for a scene with more than two filmed branches, so the safety net degrades exactly where the branching gets interesting.
The clarification ceiling of three, the 0.9 second endpointing window, the 7 second silence prompt and the 18 word reply limit were chosen by hand and none has been swept. The confidence score is recorded and unused. Latency has no perceptual target, because we have not established what mic-to-cut delay a viewer reads as thinking rather than hanging. The scope is one English-language character in one film with one canon block, so multi-character calls, other languages and code-switching are untested. The system has a single human author working with AI-assisted tooling throughout, so it has had no independent engineering review.
7. What practitioners can reuse
- Make the dumb mechanism the default and the smart one an addition. The engine can be switched off, or fail on every turn, and the film still reaches an ending. That is what made it safe to put a model in the critical path.
- Force the tool call and treat any deviation as failure. An out-of-range identifier should fail loudly, not snap to the nearest match.
- Declare that mechanical judgment overrides voice. If a character has personality in the prompt and also decides control flow, the personality wins unless the prompt says it does not.
- Treat silence as an utterance. Absence of input is information, and routing it through the same path as speech costs one timer and removes a class of dead ends.
- Budget the clarifications and escalate. Two chances to ask, then a requirement to commit, then take the decision away. An unbounded clarification loop is a stall with more talking.
- Let consent constraints shape the interface. The text-only caption path exists because the project ruled out cloning any performer's voice, and it proved cheaper, more accessible and more failure-tolerant.
- Send text, not audio, and log decisions, not transcripts. On-device recognition with local endpointing buys privacy, payload and latency at once, and a decision, a nudge count, a silence flag, an utterance length and a latency measure the system without keeping any words.
References
- R. Huang, L. J. Martin, C. Callison-Burch. "WHAT-IF: Exploring Branching Narratives by Meta-Prompting Large Language Models." arXiv:2412.10582, 13 December 2024, revised 20 October 2025.
- Y. Zhang, J. Hao, Z. Wang, H. Sheng, W. Zeng. "Facilitating Video Story Interaction with Multi-Agent Collaborative System." arXiv:2505.03807, May 2025.
- Netflix, Black Mirror: Bandersnatch, 2018. Menu-driven interactive feature.
- Eko. Interactive video platform and catalogue, closed 2024.
- Treezy Technologies Inc. Architecture decision records 0001, 0003, 0006 and 0007, 13 July 2026, and 0010, 26 August 2026, accepted 1 September 2026. Internal.
Colophon
The engine described here was designed and written by one human engineer working with AI-assisted development tooling throughout.
Publication path: this report here, with byline and date, then a system-paper version to arXiv under cs.HC. Two constraints apply there. arXiv has required an endorsement from an established author in the category for first-time authors since 21 January 2026, and since 31 October 2025 the computer science categories reject surveys and position papers that have not been peer reviewed. The submitted version must therefore be a system description with an evaluation, and section 4 must carry real numbers.