It counts the voices
Nobody tells it how many people are in the room. The number of speakers is read out of the audio, from the structure of the voice fingerprints themselves.
early access
timbre turns a meeting recording into a transcript where every line carries a person's name. Not "Speaker 2". Introduce your team once and their voices are recognised in every recording afterwards.
the problem
Most meeting tools get the words roughly right and the people badly wrong. You end up with a wall of Speaker 1 and Speaker 2 that swaps identity halfway through, and deciding who committed to what means listening to the recording anyway. Which is the thing the transcript was supposed to save you from.
what you usually get
Speaker 204:12So the thing I keep coming back to is the timing.
Speaker 104:18Say more about that.
Speaker 204:19It only breaks when two people talk at once.
what timbre gives you
Anna04:12So the thing I keep coming back to is the timing.
Mehmet04:18Say more about that.
Anna04:19It only breaks when two people talk at once.
how it works
That is the fact every other tool ignores. Generic transcription treats each meeting as if it had never met anyone before, so it starts from zero every time and hands you anonymous labels.
timbre keeps a voice profile for each person you introduce. Name a speaker once, by correcting a single line, and that voice is recognised in every recording from then on: across meetings, across months, across whichever room and headset they happen to use.
Nobody tells it how many people are in the room. The number of speakers is read out of the audio, from the structure of the voice fingerprints themselves.
Voice detection can silently discard speech and still produce a clean-looking transcript. timbre measures its own coverage and warns when a transcript is probably partial.
A half-second reply buried in room noise carries almost no voice signal. But "Do a what?" is unmistakably an answer, so language arbitrates where acoustics cannot.
measured, not claimed
Diarization error rate is the standard measure of who-spoke-when: missed speech, false alarm and speaker confusion, as a fraction of reference speaking time. Lower is better. These are five meetings from the AMI corpus, a public set of real recorded meetings, scored against its published ground truth.
| meeting | speakers found | DER |
|---|---|---|
| IS1009c | 4 of 4 | 7.17% |
| IS1009b | 4 of 4 | 10.05% |
| ES2004b | 4 of 4 | 10.64% |
| ES2004a | 4 of 4 | 17.14% |
| IS1009a | 2 of 4 | 28.52% |
| overall | — | 12.14% |
Scored with a 0.25 s collar and overlapping speech included, the stricter of the two common conventions. Published numbers that exclude overlap are not directly comparable. IS1009a is the honest outlier: it finds two speakers where there are four, and that single mistake accounts for most of the remaining error. These are cold-start numbers, measured without any enrolled voice profiles.
where your audio goes
Meeting recordings are among the most sensitive material a company holds: salaries, performance, legal exposure, unreleased plans. Most transcription tools process them in the United States and reserve the right to improve their models with your data.
timbre runs on servers in Nuremberg, Germany. Recordings are deleted once the transcript is produced, and no customer audio is ever used to train models. Voice profiles are stored as mathematical fingerprints, from which the original speech cannot be reconstructed.
For recordings that are not allowed to leave your network at all, the same pipeline runs self-hosted, on your own hardware, with no outbound connection. It was built to run offline before it was ever offered as a service.
pricing
Relabelling speakers by hand in a one-hour meeting takes most people the better part of an hour. Early access pricing, fixed for the first year for everyone who joins now.
€19 / month
One person, their own meetings.
€59 / month
A team that meets the same people weekly.
Talk to us
Audio that cannot leave your network.
timbre is in early access and onboarding a small number of teams by hand, so that the first transcripts are checked by a person before anyone relies on them. Tell us what your meetings look like and we will tell you honestly whether it is ready for them yet.