sotto
Hear, think, speak, and survive being interrupted. The conversational layer underneath everything else I run, and the four problems that only appear once a voice has to be listened to rather than read.
What it is
A conversation with a machine is usually built once, inside whatever product needs it, and thrown away with that product. sotto was pulled out of a hypnotherapy application because the loop turned out to be the reusable part and the therapy was not: something has to hear a person, decide what to say, say it out loud, and cope gracefully with being cut off halfway through. None of that is domain knowledge. All of it is hard.
Two services and a client library. The thinking runs on a self-hosted 30B over an OpenAI-compatible streaming endpoint, the hearing is faster-whisper held resident in process, and the speaking is an 82M parameter model on CPU. The lane that used to call a hosted model was deleted rather than left behind a flag, on the grounds that a disabled dependency is still a dependency.
Slowing a voice without slowing it
A guide that has to sound unhurried needs to speak more slowly, and every text-to-speech system offers a speed control that does exactly the wrong thing. Slowing the words themselves stretches the vowels and drags the consonants, and it sounds like a tape running slow, because that is what it is. The voice does not become calmer. It becomes damaged.
The idea that fixed it is that a model reading a sentence already puts silence in the right places. It pauses at the comma, longer at the ellipsis, longer still at the full stop, and those pauses are where unhurriedness actually lives. So the sentence is rendered once, at its natural rate, and then the silences it already contains are lengthened in proportion to how long each one already was. The words are never touched.
The proof that this is the right model of the problem is that every register in the ladder runs at a speed of 1.00 and still lands on its target words per second. The dial marked speed does nothing at all, in any of them. Silence does the entire job.
Punctuation is the lever, and the difference is measurable: a comma buys about 0.28 seconds of natural pause, an ellipsis 0.34, a full stop 0.38, and a spaced dash buys none at all.
The silence inside a word
The first version of that stretched every gap it could find, and it was subtly, awfully wrong. A stop consonant is a closure: to say the p in open, the mouth seals and the sound genuinely stops for forty to a hundred and twenty milliseconds before it releases. Measured by amplitude, that is indistinguishable from a short breath. On one line the detector found two pauses, one real at 340 milliseconds and one that was a consonant at 60, and stretched the consonant by 3.1 seconds. Three seconds of silence, in the middle of a word.
Three things came out of it. Nothing below 0.20 seconds is ever stretched, because below that a gap is articulation rather than punctuation. No mid-sentence pause may exceed 1.5 seconds, because past that it stops sounding like a pause and starts sounding like a fault. And whatever silence cannot be spent inside the sentence is not discarded, it is returned to the caller as a deficit and held as quiet after the line, so the pacing is still honoured without any single gap becoming a hole.
Speaking before the thinking is finished
A model that takes six or seven seconds to produce a full reply cannot be waited for. The stream is therefore cut into sentences as they complete, each sentence is synthesised as soon as it exists, and the audio starts on the first one while the rest is still being written. Renders overlap but each waits for the one before it to be handed over, so they are produced concurrently and still delivered in order.
Splitting a token stream into sentences is less trivial than it looks, because a full stop is not a sentence boundary in Dr. or e.g. or a numbered list, so the splitter only breaks on a terminator followed by whitespace and holds a list of abbreviations, digits and single letters that suppress it. The model also emits a control mark separating what should be spoken from machine-readable pacing instructions, which lets the consumer stop reading a stream early without losing them.
| Measured on the deployed box | Time |
|---|---|
| Brain, first token, warm | 0.9s |
| Voice, one sentence | 0.95 to 1.8s |
| A conversational turn, to first audio | 3.2 to 4.6s |
| The same reply, waited for in full | 6 to 7s |
When the first sentence still misses its budget, a pre-screened holding line is spoken while the real answer keeps arriving. It is the same mechanism a person uses when they say mm, let me think about that, and it is the only thing that makes a slow model tolerable rather than broken.
Being interrupted
The client holds the microphone open for the entire session rather than taking turns, with a small neural voice-activity detector deciding what is speech, on 32 millisecond frames, with separate thresholds for starting and stopping so a single noisy frame cannot trigger it. When the person starts talking, playback stops immediately.
The part that took a real failure to learn is what happens next. The obvious move is to abandon the response stream, since nobody is listening to it any more. That is what silently destroyed the first live session. The server needs to be told which sentences were actually heard so it can record the line as cut rather than as spoken, and it cannot know that if the client stops reading: the transcript loses the line, and the interrupt marks the wrong one. So the audio stops and the stream keeps being consumed to the end, quietly, and the conversation history ends up holding what the person really heard rather than what was generated.
That distinction matters more than it sounds. The model is shown its own previous turns, and if it is shown a paragraph the person never heard the end of, it will build on something that never happened.
Knowing nothing about the conversation
None of the above has any opinion about what is being discussed, and keeping it that way is the reason it is a library rather than a feature. Everything domain-specific sits behind an interface of seven properties and seven methods: what voice to use, how fast to pace, what to say next, whether the person is expected to reply, what the screening rules are, and what to do when a conversation ends. It is structural rather than inherited, so a consumer implements the shape and nothing else.
The therapy application keeps its phase machine and its contraindications behind that boundary. A plain assistant implements the same shape in a few dozen lines. Neither one reimplements streaming, sentence splitting, holding lines, interruption or screening, and there is a test whose only job is to assert that the interface is still exactly what the loop needs, so it cannot quietly grow.
One consequence is worth the words: when a conversation expects the person to stay quiet, the next line can be drafted during the pause and thrown away if they speak. Because that draft exists before anyone has heard it, it is the one place a repetition can be caught and deleted rather than apologised for. The similarity bar for throwing one away is set deliberately high, because the language of a guided session is formulaic on purpose and a stricter bar threw away lines that were good.
The bug that only a queue could fix
The voice service began refusing sentences under load, and the reason had nothing to do with audio. The render function was written as an ordinary synchronous def, which means the web framework helpfully ran it in a forty-thread worker pool, and nothing anywhere bounded how many renders could be in flight. Each one wants around a gigabyte while it works.
The fix was two render slots, and the important detail is that they queue rather than shed. A rate limiter that rejects the excess would have been easier and would have been wrong: a sentence that waits its turn is delivered late, and a sentence that is refused is never spoken at all. Underneath it, memory that appeared to leak turned out to be the allocator holding arenas rather than returning them, so it is now asked to release after each render.
The real lesson was the reporting. Twenty-two kills had happened before anyone noticed, because the process restarted cleanly each time and the only visible symptom was an occasional error mid-conversation. The health endpoint now counts renders refused, renders queued and restarts, on the principle that the reason this lasted was not that it was subtle but that nothing anywhere reported it.
A related one, same shape: the tensor library read the host's core count rather than the container's limit and oversubscribed its own threads. Matching them took one ten-word sentence from 6.7 seconds to 1.6.
What it cannot do
It is not deployed as a service for anyone but its first consumer, and the HTTP surface has barely been called from another machine. Conversations live in memory and die with the process, which also means it runs as a single worker, because a second one would confidently answer for sessions it has never heard of. Authentication is one shared key, and every authenticated caller is the same caller as far as the code is concerned.
The screening layer is a substring match over a short list of terms, and the repository is explicit that these are an engineering scaffold and not validated clinical criteria. The fifteen tests cover the loop with fakes, and they are good tests, named after the guarantee each one defends rather than the function it calls. They cover none of the voice service, which is the part that has actually broken in production.
And the open question is the one that cannot be measured. The pace ladder is hit, the latency is known, the failures are instrumented. Whether the voice is one anybody would want to be spoken to by is a listening decision, and it has not been made yet.
Timings are from the deployed box and from the farm over a tunnel. They are recorded in comments and commit messages at the moment they were taken, rather than produced by a benchmark that can be re-run, which is a gap worth naming.