The serious page

How It Works

Every tractor is a retrieval-augmented language model built from hours of unscripted speech by the person it puns on. There is no prompt anywhere on this site that says "pretend to be a comedian". This page is how I built them, what I measure, and where the limits are.

AI tractors
21
Unscripted speech ingested
63 hours
Full psychological profiles
6
Direct prompts used
0
The principle

Tractors are experts, not prompts

The usual way to make an AI play a famous person is to tell a large language model to act like them. You get a generic voice with a famous name attached, saying whatever the model thinks a famous person would say. I didn't want that. It isn't interesting and it isn't fair on the person.

So each tractor is a SiteEngine expert, which is a retrieval system over a corpus of its namesake actually talking, with the model only allowed to answer from what it retrieves. Each vote and each private thought in a transcript is a query to that expert, in character, with its emotional state applied. When Hannah Furrow starts on about probabilities at the Round Bale, that's because the person behind her does it constantly and the system found those passages.

The same rule runs all the way down. Nobody writes the cartoon's dialogue. It's what the experts said when the simulated Round Bale ran, cut for length.

The corpus

Building a tractor from what somebody has said

Source plan

One plan per tractor before any download, naming the sources. Deception formats first, because a panel show where everyone is trying to catch you lying is the closest thing to a Round Bale that exists on tape.

Acquire

I pull the audio, split it into speakers, and find the subject by matching against a short clip of them speaking alone. Everyone else's speech is dropped.

Transcribe

Speech recognition with word timings, keeping the hesitations and false starts, because that's what unscripted speech sounds like, and the model should too.

Tag

Every source is tagged by mode: unscripted, performed (their own stand-up), scripted (someone else's words), or about (third parties writing about them). Retrieval weights by mode, so drama roles never leak into the voice.

Stage

The text is embedded and summarised, and an emotional profile is extracted from it. How much a tractor masks and how readily it lies comes out of its transcripts. I don't write any of that by hand.

Promote

Once a corpus passes its checks it is promoted to the serving system and bound to a tractor. Later additions go through staging again, so a tractor can always be improved without being rebuilt.

Playing the game

Memory, audiences and the Traitors' Barn

The Traitors is a game about who knows what, so the machinery has to be too. Every event in the game state carries its audience: everyone, the Traitors only, or a named list, and a tractor is shown only the events it was present for. What was said in the Traitors' Barn can be retrieved by the Traitors and nobody else.

Each tractor also carries notes on how it has played on the real broadcast so far, written from the aired episodes and stopping at the last one. A prediction for episode four is made by tractors that have seen episodes one to three and nothing after, and the same cut-off applies to retrieval, so an expert can't quietly read ahead.

Said aloud, thought privately

I ask for every answer in two parts: what the tractor says to the room and what it's actually thinking. The second part is why the transcripts are worth reading. When a Traitor votes for a Faithful "because he's already told us he's a Traitor", the private line underneath says what it's really doing.

Emotion and strain

Each tractor has an internal emotional state and the one it shows, and the gap between the two is tracked as masking strain. Lies it has to keep up add deception strain. The cartoon draws this straight from the numbers: a tractor under strain starts to leak exhaust, and a big enough lie belches smoke.

Measuring it

Does it actually sound like them?

I test this rather than assume it, and the honest answer is mostly.

I took twenty real interview answers from one subject, had the expert answer the same twenty questions, and asked a model judge to tell them apart. It got forty out of forty. Its own reasons showed it wasn't hearing the voice, though. The real answers were transcribed speech, broken up and full of recognition errors, and the generated ones were written prose. It was spotting a transcript.

So I rebuilt the test in audio, with both sides passed through the same synthesis so anything the machinery added would cancel out. On identity, a speaker-recognition model couldn't separate them: real and generated scored within a tenth of a standard deviation of each other. What was left was variety. The real person varied more from answer to answer than the model did, which means the next problem isn't accuracy, it's consistency, and too much of it.

Corpus size matters less than I expected and recording conditions matter more. Going from two hours of a subject's speech to six moved the similarity score by 0.155. Switching the reference recordings from a reverberant venue to a dry studio moved it by 0.214. The full write-ups are at traitorbot.com.

What it is not

The voices are cartoons on purpose

The research above measures how closely an AI can match a real person's way of speaking. That work stays in the lab. The voices in the cartoon are a different thing: pitched up, sped up, and pushed into South Park register on purpose. No recording of any real person is used in any episode, and none of the voices is meant to sound like the person it parodies. Rob Bucket sounds like a cartoon tractor because that's what he is.

The words are the same. The predictions and transcripts are fiction about how a game might go, generated by a model. They aren't claims about what any real person thinks or would do, and the private thoughts are the model's invention.

Read the full disclaimer