Voice isolation on par with Discord #3

Open
opened 2026-10-04 20:18:48 +00:00 by robocub · 1 comment
Owner

Why

Background noise (keyboards, fans, kids, TV) is one of the first things Discord users notice when they try Matrix calls. Discord ships Krisp, which is proprietary; we need an open equivalent.

What exists

Only libwebrtc's built-in processing: commet/lib/client/components/voip/webrtc_default_devices.dart asks for echoCancellation: true, noiseSuppression: true, autoGainControl: false. The LiveKit path (matrix_livekit_backend.dart) passes only a device id. libwebrtc's suppressor handles steady hiss, not keyboards or voices in the room. Upstream has the same request open (Commet #603).

Options

Model License Notes
RNNoise BSD-3 Tiny, 48 kHz / 10 ms frames, proven (OBS, Mumble). Good first step.
DeepFilterNet 3 MIT / Apache-2.0 Written in Rust, clearly better on non-stationary noise, heavier CPU. The app already ships a Rust library via flutter_rust_bridge.

The hard part

Inserting processing into the capture path. flutter_webrtc has no audio-processing hook on desktop. Commet already pins its own forks of flutter-webrtc and the LiveKit Flutter SDK, so the realistic route is a capture post-processing hook in that fork (libwebrtc's APM supports a custom capture post-processor), calling into the Rust library. Needs a per-platform look at Linux and Windows first; Android has its own capture path.

Plan

  1. Spike: confirm the hook point in the forked flutter-webrtc on Linux and measure RNNoise CPU per call.
  2. RNNoise behind a setting ("Noise suppression: Off / Standard / Strong"), Standard = current behaviour.
  3. Evaluate DeepFilterNet as "Strong".
  4. Optional: a noise gate / voice-activity threshold, like Discord's input sensitivity slider.

Acceptance

A/B recordings with keyboard typing and a fan, the same mic, judged by ear by at least two people. CPU stays reasonable on a low-end laptop.


Filed with LLM assistance. This is a fork-only issue; never refile it on Commet's tracker.

## Why Background noise (keyboards, fans, kids, TV) is one of the first things Discord users notice when they try Matrix calls. Discord ships Krisp, which is proprietary; we need an open equivalent. ## What exists Only libwebrtc's built-in processing: `commet/lib/client/components/voip/webrtc_default_devices.dart` asks for `echoCancellation: true`, `noiseSuppression: true`, `autoGainControl: false`. The LiveKit path (`matrix_livekit_backend.dart`) passes only a device id. libwebrtc's suppressor handles steady hiss, not keyboards or voices in the room. Upstream has the same request open (Commet #603). ## Options | Model | License | Notes | |---|---|---| | RNNoise | BSD-3 | Tiny, 48 kHz / 10 ms frames, proven (OBS, Mumble). Good first step. | | DeepFilterNet 3 | MIT / Apache-2.0 | Written in Rust, clearly better on non-stationary noise, heavier CPU. The app already ships a Rust library via flutter_rust_bridge. | ## The hard part Inserting processing into the capture path. flutter_webrtc has no audio-processing hook on desktop. Commet already pins its own forks of flutter-webrtc and the LiveKit Flutter SDK, so the realistic route is a capture post-processing hook in that fork (libwebrtc's APM supports a custom capture post-processor), calling into the Rust library. Needs a per-platform look at Linux and Windows first; Android has its own capture path. ## Plan 1. Spike: confirm the hook point in the forked flutter-webrtc on Linux and measure RNNoise CPU per call. 2. RNNoise behind a setting ("Noise suppression: Off / Standard / Strong"), Standard = current behaviour. 3. Evaluate DeepFilterNet as "Strong". 4. Optional: a noise gate / voice-activity threshold, like Discord's input sensitivity slider. ## Acceptance A/B recordings with keyboard typing and a fan, the same mic, judged by ear by at least two people. CPU stays reasonable on a low-end laptop. --- _Filed with LLM assistance. This is a fork-only issue; never refile it on Commet's tracker._
Author
Owner

Research: where noise suppression can plug in (2026-10-05)

Short version: the hook already exists on desktop; nobody calls it.

What's there

  • Commet pins its own forks: commetchat/flutter-webrtc@hkdf (2d0dce5) and commetchat/livekit-client-sdk-flutter@hkdf.
  • On Linux and Windows, flutter-webrtc doesn't compile WebRTC. It downloads the prebuilt libwebrtc.zip from upstream flutter-webrtc v1.4.0 (March 2026), which wraps WebRTC in webrtc-sdk/libwebrtc.
  • That wrapper has had RTCAudioProcessing since July 2025 (webrtc-sdk/libwebrtc #108):
    class RTCAudioProcessing {
      class CustomProcessing { Initialize(rate, channels); Process(num_bands, num_frames, buffer_size, float* buffer); Reset(rate); Release(); };
      virtual void SetCapturePostProcessing(CustomProcessing*) = 0;
      virtual void SetRenderPreProcessing(CustomProcessing*) = 0;
    };
    
    and RTCPeerConnectionFactory::GetAudioProcessing().
  • The desktop plugin already fetches it (common/cpp/src/flutter_webrtc_base.cc:22: audio_processing_ = factory_->GetAudioProcessing();) but exposes no method to install a processor. Our Linux CI build links against it, so the v1.4.0 prebuilt does ship it.
  • The processor gets 10 ms of the first channel, full band, as floats on the int16 scale (audio->channels()[0]); at 48 kHz that's 480 samples. That is exactly RNNoise's frame size and sample scale, so no resampling or reframing is needed at 48 kHz.
  • Android / Apple have a separate AudioProcessingAdapter (ExternalAudioFrameProcessing.process(numBands, numFrames, ByteBuffer)). That's what LiveKit's noise-filter plugins use on mobile.

Proposed design (desktop first)

  1. Fork patch in a nether.codes fork of commetchat/flutter-webrtc@hkdf: a method-channel call setCapturePostProcessor(fn, userData) that takes a native function pointer (void process(void* user, int frames, float* buf)) and wraps it in a CustomProcessing; fn = 0 removes it. Roughly 100 lines of C++, no new dependencies, and the plugin stays model-agnostic.
  2. Model in Rust: Commet already ships rust_lib_commet on Linux and Windows. Add nnnoiseless (a pure-Rust RNNoise port, BSD-3) and export an extern "C" processor with one state per stream. A later "Strong" mode can swap in DeepFilterNet (MIT/Apache, Rust) behind the same function pointer.
  3. Dart: look up the Rust symbol with dart:ffi, pass its address to the plugin, and add a setting Off / Standard / Strong. Standard is today's WebRTC suppressor; Strong is RNNoise on top.
  4. Android later, via the existing AudioProcessingAdapter, through JNI or a Kotlin RNNoise binding.

Effort and risks

  • Medium: about a day for a working Linux build, then listening tests. The fork adds a third pinned fork to maintain. The C++ patch is generic enough to offer upstream to flutter-webrtc (hand-written; check their contribution policy first).
  • Things to verify in the spike: the processor runs at 48 kHz on common mics (otherwise use RNNoise only at 48 kHz and bypass at other rates); CPU per call; and that the WebRTC suppressor plus RNNoise stacked doesn't sound over-processed (if so, turn the WebRTC one off in Strong mode).

Filed with LLM assistance.

## Research: where noise suppression can plug in (2026-10-05) **Short version: the hook already exists on desktop; nobody calls it.** ### What's there - Commet pins its own forks: `commetchat/flutter-webrtc@hkdf` (2d0dce5) and `commetchat/livekit-client-sdk-flutter@hkdf`. - On **Linux and Windows**, flutter-webrtc doesn't compile WebRTC. It downloads the prebuilt `libwebrtc.zip` from upstream flutter-webrtc **v1.4.0** (March 2026), which wraps WebRTC in [webrtc-sdk/libwebrtc](https://github.com/webrtc-sdk/libwebrtc). - That wrapper has had `RTCAudioProcessing` since July 2025 (webrtc-sdk/libwebrtc #108): ```cpp class RTCAudioProcessing { class CustomProcessing { Initialize(rate, channels); Process(num_bands, num_frames, buffer_size, float* buffer); Reset(rate); Release(); }; virtual void SetCapturePostProcessing(CustomProcessing*) = 0; virtual void SetRenderPreProcessing(CustomProcessing*) = 0; }; ``` and `RTCPeerConnectionFactory::GetAudioProcessing()`. - The desktop plugin already fetches it (`common/cpp/src/flutter_webrtc_base.cc:22`: `audio_processing_ = factory_->GetAudioProcessing();`) but exposes **no** method to install a processor. Our Linux CI build links against it, so the v1.4.0 prebuilt does ship it. - The processor gets **10 ms of the first channel, full band, as floats on the int16 scale** (`audio->channels()[0]`); at 48 kHz that's 480 samples. That is exactly RNNoise's frame size and sample scale, so no resampling or reframing is needed at 48 kHz. - **Android / Apple** have a separate `AudioProcessingAdapter` (`ExternalAudioFrameProcessing.process(numBands, numFrames, ByteBuffer)`). That's what LiveKit's noise-filter plugins use on mobile. ### Proposed design (desktop first) 1. **Fork patch** in a nether.codes fork of `commetchat/flutter-webrtc@hkdf`: a method-channel call `setCapturePostProcessor(fn, userData)` that takes a native function pointer (`void process(void* user, int frames, float* buf)`) and wraps it in a `CustomProcessing`; `fn = 0` removes it. Roughly 100 lines of C++, no new dependencies, and the plugin stays model-agnostic. 2. **Model in Rust**: Commet already ships `rust_lib_commet` on Linux and Windows. Add [nnnoiseless](https://github.com/jneem/nnnoiseless) (a pure-Rust RNNoise port, BSD-3) and export an `extern "C"` processor with one state per stream. A later "Strong" mode can swap in DeepFilterNet (MIT/Apache, Rust) behind the same function pointer. 3. **Dart**: look up the Rust symbol with `dart:ffi`, pass its address to the plugin, and add a setting Off / Standard / Strong. Standard is today's WebRTC suppressor; Strong is RNNoise on top. 4. **Android later**, via the existing `AudioProcessingAdapter`, through JNI or a Kotlin RNNoise binding. ### Effort and risks - Medium: about a day for a working Linux build, then listening tests. The fork adds a third pinned fork to maintain. The C++ patch is generic enough to offer upstream to flutter-webrtc (hand-written; check their contribution policy first). - Things to verify in the spike: the processor runs at 48 kHz on common mics (otherwise use RNNoise only at 48 kHz and bypass at other rates); CPU per call; and that the WebRTC suppressor plus RNNoise stacked doesn't sound over-processed (if so, turn the WebRTC one off in Strong mode). --- _Filed with LLM assistance._
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
robocub/vommet#3
No description provided.