Transcribe voice messages to text (like Telegram) #70

Open
opened 2026-10-06 21:08:58 +00:00 by robocub · 0 comments
Member

Goal: a "Transcribe" button on every voice message (ours and other clients', m.audio with org.matrix.msc3245.voice). Tap it and the spoken text appears under the waveform, like Telegram's speech-to-text. Useful when you can't play audio, for searching what was said, and for accessibility.

Privacy first

Many rooms are end-to-end encrypted. A transcription feature must not quietly send decrypted audio to a server.

  • Default: on-device. Audio never leaves the device.
  • Optional server mode (faster or more accurate, e.g. a self-hosted Whisper service) only as an explicit opt-in. It must say plainly that audio leaves the device, and should default to off in encrypted rooms.

How (on-device)

  • whisper.cpp through our existing Rust library (rust_lib_commet, e.g. the whisper-rs crate). The same code runs on Linux, Windows and Android (arm64).
  • Models downloaded on first use, not bundled: tiny ≈ 75 MB / base ≈ 150 MB / small ≈ 470 MB. Default base, with a Settings choice and a size/speed/accuracy note. Language auto-detect, or the user's language as a hint.
  • Voice messages are Ogg/Opus (ours: record → Opus). Decode to 16 kHz mono PCM for Whisper (libopus via the Rust side, or reuse the app's player decoding).
  • Run off the UI thread, show progress, allow cancel. A 30-second message should take a few seconds on a phone with the tiny/base model, so measure on the operator's OnePlus 9 Pro and a weak laptop (#63).
  • Android alternative: the platform SpeechRecognizer can't take a file as input on most devices, so Whisper is the portable choice.

Storing and sharing transcripts

  • Cache transcripts locally per event id, so a second tap is instant and survives restarts.
  • Later, optional: send the transcript alongside your own voice messages (e.g. as the event's text body or extensible-events text) so every client can show it without transcribing. Only with the sender's consent, and never for other people's messages.

UX

  • "Transcribe" under each voice message. While running, a spinner; then the text, collapsible, with copy and a "low confidence" hint.
  • Settings › Voice messages: model size (download/delete), language, auto-transcribe (off by default; Wi-Fi only for the model download).

Phases

  1. Desktop prototype: whisper.cpp through rust_lib_commet, tiny/base model, a button on voice messages, local cache.
  2. Android build of the same, measured on a real phone.
  3. Settings (model management, language, auto-transcribe).
  4. Optional: transcript attached to your own outgoing voice messages; optional opt-in server mode.

Related: voice messages (#15). Upstream: no Commet issue or branch about transcription (searched 2026-10-06).

Goal: a **"Transcribe"** button on every voice message (ours and other clients', `m.audio` with `org.matrix.msc3245.voice`). Tap it and the spoken text appears under the waveform, like Telegram's speech-to-text. Useful when you can't play audio, for searching what was said, and for accessibility. ## Privacy first Many rooms are end-to-end encrypted. A transcription feature must not quietly send decrypted audio to a server. - **Default: on-device.** Audio never leaves the device. - **Optional server mode** (faster or more accurate, e.g. a self-hosted Whisper service) only as an explicit opt-in. It must say plainly that audio leaves the device, and should default to off in encrypted rooms. ## How (on-device) - **whisper.cpp** through our existing Rust library (`rust_lib_commet`, e.g. the `whisper-rs` crate). The same code runs on Linux, Windows and Android (arm64). - **Models downloaded on first use**, not bundled: tiny ≈ 75 MB / base ≈ 150 MB / small ≈ 470 MB. Default base, with a Settings choice and a size/speed/accuracy note. Language auto-detect, or the user's language as a hint. - Voice messages are Ogg/Opus (ours: `record` → Opus). Decode to 16 kHz mono PCM for Whisper (libopus via the Rust side, or reuse the app's player decoding). - Run off the UI thread, show progress, allow cancel. A 30-second message should take a few seconds on a phone with the tiny/base model, so measure on the operator's OnePlus 9 Pro and a weak laptop (#63). - **Android alternative:** the platform `SpeechRecognizer` can't take a file as input on most devices, so Whisper is the portable choice. ## Storing and sharing transcripts - Cache transcripts locally per event id, so a second tap is instant and survives restarts. - **Later, optional:** send the transcript alongside your *own* voice messages (e.g. as the event's text body or extensible-events text) so every client can show it without transcribing. Only with the sender's consent, and never for other people's messages. ## UX - "Transcribe" under each voice message. While running, a spinner; then the text, collapsible, with copy and a "low confidence" hint. - Settings › Voice messages: model size (download/delete), language, auto-transcribe (off by default; Wi-Fi only for the model download). ## Phases 1. Desktop prototype: whisper.cpp through `rust_lib_commet`, tiny/base model, a button on voice messages, local cache. 2. Android build of the same, measured on a real phone. 3. Settings (model management, language, auto-transcribe). 4. Optional: transcript attached to your own outgoing voice messages; optional opt-in server mode. Related: voice messages (#15). Upstream: no Commet issue or branch about transcription (searched 2026-10-06).
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
nether/vommet#70
No description provided.