early access

who said what.by name.

timbre turns a meeting recording into a transcript where every line carries a person's name. Not "Speaker 2". Introduce your team once and their voices are recognised in every recording afterwards.

Request early access See pricing
a real 39-minute meeting · AMI ES2004bresolving…
0:0019:3039:01
speaker 1 speaker 2 speaker 3 speaker 4
diarization error
12.14% on AMI
processed in
Germany · Nuremberg
audio kept after
0 · deleted
used for training
never

the problem

A transcript nobody can read is not a transcript.

Most meeting tools get the words roughly right and the people badly wrong. You end up with a wall of Speaker 1 and Speaker 2 that swaps identity halfway through, and deciding who committed to what means listening to the recording anyway. Which is the thing the transcript was supposed to save you from.

what you usually get

Speaker 204:12So the thing I keep coming back to is the timing.

Speaker 104:18Say more about that.

Speaker 204:19It only breaks when two people talk at once.

what timbre gives you

Anna04:12So the thing I keep coming back to is the timing.

Mehmet04:18Say more about that.

Anna04:19It only breaks when two people talk at once.

how it works

Your team is the same people every week.

That is the fact every other tool ignores. Generic transcription treats each meeting as if it had never met anyone before, so it starts from zero every time and hands you anonymous labels.

timbre keeps a voice profile for each person you introduce. Name a speaker once, by correcting a single line, and that voice is recognised in every recording from then on: across meetings, across months, across whichever room and headset they happen to use.

It counts the voices

Nobody tells it how many people are in the room. The number of speakers is read out of the audio, from the structure of the voice fingerprints themselves.

It admits what it missed

Voice detection can silently discard speech and still produce a clean-looking transcript. timbre measures its own coverage and warns when a transcript is probably partial.

It reads the words too

A half-second reply buried in room noise carries almost no voice signal. But "Do a what?" is unmistakably an answer, so language arbitrates where acoustics cannot.

measured, not claimed

Every number here comes from a benchmark you can rerun.

Diarization error rate is the standard measure of who-spoke-when: missed speech, false alarm and speaker confusion, as a fraction of reference speaking time. Lower is better. These are five meetings from the AMI corpus, a public set of real recorded meetings, scored against its published ground truth.

meetingspeakers foundDER
IS1009c4 of 47.17%
IS1009b4 of 410.05%
ES2004b4 of 410.64%
ES2004a4 of 417.14%
IS1009a2 of 428.52%
overall12.14%

Scored with a 0.25 s collar and overlapping speech included, the stricter of the two common conventions. Published numbers that exclude overlap are not directly comparable. IS1009a is the honest outlier: it finds two speakers where there are four, and that single mistake accounts for most of the remaining error. These are cold-start numbers, measured without any enrolled voice profiles.

where your audio goes

Germany, briefly, and then nowhere.

Meeting recordings are among the most sensitive material a company holds: salaries, performance, legal exposure, unreleased plans. Most transcription tools process them in the United States and reserve the right to improve their models with your data.

timbre runs on servers in Nuremberg, Germany. Recordings are deleted once the transcript is produced, and no customer audio is ever used to train models. Voice profiles are stored as mathematical fingerprints, from which the original speech cannot be reconstructed.

For recordings that are not allowed to leave your network at all, the same pipeline runs self-hosted, on your own hardware, with no outbound connection. It was built to run offline before it was ever offered as a service.

pricing

Priced against the hour you would spend fixing it.

Relabelling speakers by hand in a one-hour meeting takes most people the better part of an hour. Early access pricing, fixed for the first year for everyone who joins now.

Solo

€19 / month

One person, their own meetings.

  • 10 hours of audio a month
  • Unlimited voice profiles
  • Markdown, JSON and subtitle export

Team

€59 / month

A team that meets the same people weekly.

  • 40 hours of audio a month
  • Shared voice profiles across the team
  • Five seats included
  • Correction workspace

Self-hosted

Talk to us

Audio that cannot leave your network.

  • Runs on your own hardware
  • No outbound connection required
  • Annual licence
Request early access

timbre is in early access and onboarding a small number of teams by hand, so that the first transcripts are checked by a person before anyone relies on them. Tell us what your meetings look like and we will tell you honestly whether it is ready for them yet.