Averse Studio

Giving Façade's characters a live voice

Façade let you type anything to a couple whose marriage is coming apart, then answered with one of its recorded lines. We kept its story engine in charge and had a language model write every line live, performed by cloned voices of the original actors.

Video coming soon

A full evening with Trip and Grace, every line written and voiced live.

The game

Façade came out in 2005, made by Michael Mateas and Andrew Stern. You visit an old couple of friends, Trip and Grace, for drinks, and over twenty minutes their marriage falls apart in front of you. You take part by typing whatever you like.

Underneath, a drama manager picks the next story beat, and each character is run by behaviours written in ABL, a language the authors built for believable agents. What you type is read as one of a small set of discourse acts (agreeing, criticising, flirting, bringing up a topic), and the characters react with one of thousands of recorded lines. The drama could follow you. The words couldn't: whatever you said, the answer had been written years before you said it.

What changed, and what didn't

The story engine still decides everything it decided before. When it wants Trip to speak, it names the line the script had for that moment. We treat that line as the intent, what the moment has to do in the story, and write new words for it.

Story
The original drama manager and character behaviours, running unchanged, choose each beat and what the next line has to do.
Words
A language model writes the line from that intent, the evening so far and what the player just typed, with an acting cue that matches the face the story engine gave the character.
Voice
Cloned voices of the original actors perform the line, cue included, streamed so it starts playing about 0.75 seconds into voicing.
Timing
While one line plays, we read the story engine's behaviour tree to see what each character is about to say, and write and voice it ahead.
Trip at the door, with his greeting as a subtitle
Trip's greeting at the door. The script asks for a warm hello; the words were written when the player arrived, with the player's name in them.

A session, as it happened

Two minutes from a test run. The player's lines were typed by our test bot. Everything Trip and Grace say was written and voiced live in reply. Bracketed words are the acting cues the writer gave the voice; the short "uh…" lines are the original actors' recordings, played instantly while the real answer is written.

  1. Playergrace, you seem upset
  2. GraceUhh…
  3. Trip[nervous] No, no, hey, our friend just got here! Let's just… let the machine pick that up, okay? [forced cheer] We're having a night!
  4. Grace[surprised] Oh, no, I'm fine, really— [clear throat] excuse me, I should get the phone.
  5. Trip[confident] Ah, yeah, it really is something, isn't it, Alex? Twenty-three floors up, man [chuckling] I swear I never get tired of looking out over the city.
  6. Grace[disdainful] [clear throat] Right, well… I really need to call maintenance to clean these filthy windows.
  7. Playertrip, why are you so nervous?
  8. Graceuh…
  9. Grace[sarcastic, chuckling] Oh, nervous? [break] Trip isn't nervous, Alex, he's just maintaining his signature high-strung host energy. Aren't you, sweetie?
  10. Playerare you two okay?
  11. Grace[anxious] What? No, of course we're fine, Alex, everything is [clear throat] completely fine. Why wouldn't we be?

Making it fast enough

A drama about two people talking over each other dies if every line waits for a server. In the original, about 0.7 seconds pass between one speaker and the next. Our first working version took four.

Three things closed most of the gap. The voice streams, so a line plays as soon as its first words arrive instead of when the whole file is in. The writer's instructions, about 6,000 tokens of character notes and rules, are cached, so each line only sends what's new. And while one character talks, we look inside the story engine's running behaviour tree, find the lines it has already queued, and write and voice them before they're asked for.

Measured in a seven-minute test session, 7 October 2026
1.87 smedian from the story asking for a line to the line playing (3.9 s before this work)
21 of 46lines playing within 1.2 seconds of being asked for; 10 within half a second
41%of next lines named exactly by the behaviour-tree lookahead
0 of 51acting cues spoken aloud by mistake, checked with speech recognition
$0.0006writing cost per line

A lab instead of playtests

Two playthroughs of Façade are never the same evening, so comparing a change across runs measures luck as much as the change. Every session records what the story engine chose and when. A replay lab feeds those choices back without the game, through the real writer and voice, and puts out the evening as one audio file and a timeline. Every change is measured on the same story; a comparison takes minutes instead of a full playthrough.

What's next

The lookahead only sees lines the story has already queued. Most misses are branch points: what the beat does if the player stays quiet, or how it reacts if they speak. Writing the likely branches ahead is next, along with mouth shapes driven by the words rather than loudness, and letting a character finish a sentence when the player interrupts.