Overview

High
Severity
CVSS
No
Exploited ITW
Fixed
Fix Status
ImpactOut of bounds read in Web Speech
DescriptionOut of bounds read in Web Speech
ComponentWeb Speech
Bug ClassOOB
Tracker482828615
Fix commit88a0535d5bab (chromium/src) +30/-13
CISA KEVNot listed
CreditedGrischa Hauser
Disclosed2026-03-10

Changed Functions

FunctionChangeNotes
if
content/browser/speech/speech_recognizer_impl.cc
modified

Files Changed

  • content/browser/speech/speech_recognizer_impl.cc
  • content/browser/speech/speech_recognizer_impl.h
From 88a0535d5bab4c1e7e97af17882a3c00d021fba9 Mon Sep 17 00:00:00 2001
From: Thomas Guilbert <[email protected]>
Date: Wed, 18 Feb 2026 13:08:43 -0800
Subject: [PATCH] [CodeHealth] Spanify (de)interleaving in SpeechRecognizerImpl

This CL updates an instance of `AudioBus::ToInterleaved()` to use its
safer, spanified counterpart.

It also removes one round trip of deinterleaving/interleaving + copy
converting to/from float, which saves one memory copy. This extra round
trip might have had the intention of clipping/sanitizing incoming data,
which is only useful for `float`, not `int16_t`.

Finally, this CL also hardens by rejecting potentially bad messages
coming from the renderer, and using checked math when calculating
memory allocation sizes.

See linked bugs for additional details.

Bug: 373960632, 482828615
Change-Id: I03c0e3342651bc6b55aa3505206e54c1c56bac5e
Reviewed-on: https://chromium-review.googlesource.com/c/chromium/src/+/7573019
Reviewed-by: Frank Liberato <[email protected]>
Reviewed-by: Nasko Oskov <[email protected]>
Reviewed-by: Tom Sepez <[email protected]>
Commit-Queue: Thomas Guilbert <[email protected]>
Cr-Commit-Position: refs/heads/main@{#1586654}
---

diff --git a/content/browser/speech/speech_recognizer_impl.cc b/content/browser/speech/speech_recognizer_impl.cc
index b246f4fc..b94a4cec 100644
--- a/content/browser/speech/speech_recognizer_impl.cc
+++ b/content/browser/speech/speech_recognizer_impl.cc
@@ -156,12 +156,12 @@
   // See http://crbug.com/506051 for details.
   audio_converter_.Convert(output_bus_.get());
   // Create an audio chunk based on the converted result.
-  scoped_refptr<AudioChunk> chunk(new AudioChunk(
+  auto chunk = base::MakeRefCounted<AudioChunk>(
       output_parameters_.GetBytesPerBuffer(media::kSampleFormatS16),
-      kNumBitsPerAudioSample / 8));
+      kBytesPerAudioSample);
 
-  static_assert(SpeechRecognizerImpl::kNumBitsPerAudioSample == 16,
-                "kNumBitsPerAudioSample must match interleaving type.");
+  static_assert(SpeechRecognizerImpl::kBytesPerAudioSample == sizeof(int16_t),
+                "kBytesPerAudioSample must match interleaving type.");
   output_bus_->ToInterleaved<media::SignedInt16SampleTypeTraits>(
       chunk->SamplesData16AsWriteableSpan());
   return chunk;
@@ -334,17 +334,33 @@
     return;
   }
 
-  std::unique_ptr<media::AudioBus> data =
-      AudioBus::Create(buffer->channel_count, buffer->frame_count);
+  if (buffer->channel_count <= 0 || buffer->frame_count <= 0) {
+    mojo::ReportBadMessage("AudioDataS16: non-positive dimensions");
+    return;
+  }
 
-  data->FromInterleaved<media::SignedInt16SampleTypeTraits>(
-      buffer->data.data(), buffer->frame_count);
+  auto total_samples =
+      base::CheckMul(buffer->channel_count, buffer->frame_count);
 
-  scoped_refptr<AudioChunk> chunk(new AudioChunk(
-      buffer->channel_count * buffer->frame_count * kNumBitsPerAudioSample / 8,
-      kNumBitsPerAudioSample / 8));
-  data->ToInterleaved<media::SignedInt16SampleTypeTraits>(
-      chunk->SamplesData16AsWriteableSpan());
+  if (!total_samples.IsValid() ||
+      buffer->data.size() != total_samples.ValueOrDie<size_t>()) {
+    mojo::ReportBadMessage("AudioDataS16: size mismatch");
+    return;
+  }
+
+  // If ever the sample format changed to `float`, we would have to clip and
+  // sanitize data coming from the renderer. `int16_t` doesn't need it, since it
+  // cannot represent invalid values or values outside the [-1.0, 1.0] range.
+  static_assert(SpeechRecognizerImpl::kBytesPerAudioSample == sizeof(int16_t),
+                "`AddAudioFromRenderer()` expects `int16_t`.");
+
+  const auto chunk_size =
+      base::CheckMul(total_samples, kBytesPerAudioSample).ValueOrDie<size_t>();
+  auto chunk =
+      base::MakeRefCounted<AudioChunk>(chunk_size, kBytesPerAudioSample);
+
+  chunk->SamplesData16AsWriteableSpan().copy_from(buffer->data);
+
   FSMEventArgs event_args(EVENT_AUDIO_DATA);
   event_args.audio_chunk = std::move(chunk);
   GetIOThreadTaskRunner({})->PostTask(
diff --git a/content/browser/speech/speech_recognizer_impl.h b/content/browser/speech/speech_recognizer_impl.h
index 317e103..4760bcba 100644
--- a/content/browser/speech/speech_recognizer_impl.h
+++ b/content/browser/speech/speech_recognizer_impl.h
@@ -48,6 +48,7 @@
   static constexpr media::ChannelLayoutConfig kChannelLayoutConfig =
       media::ChannelLayoutConfig::Mono();
   static constexpr int kNumBitsPerAudioSample = 16;
+  static constexpr size_t kBytesPerAudioSample = kNumBitsPerAudioSample / 8;
   static constexpr int kNoSpeechTimeoutMs = 8000;
   static constexpr int kEndpointerEstimationTimeMs = 300;
 
Loading diff…

Original Bug Report

reported by [email protected]

Heap OOB read in SpeechRecognizerImpl::AddAudioFromRenderer

Summary

SpeechRecognizerImpl fails to validate that the size of the data array in the AudioDataS16 Mojo struct matches the declared channel_count * frame_count. This allows a compromised renderer to trigger a heap out-of-bounds read in the Browser Process.

VULNERABILITY DETAILS

Heap-buffer-overflow (OOB read) in the browser process via SpeechRecognizerImpl::AddAudioFromRenderer(). A compromised renderer can read attacker-controlled amounts of browser-process heap memory through the media.mojom.SpeechRecognitionAudioForwarder Mojo interface. No permissions are required, the audio forwarder path bypasses the microphone permission check.

The per-packet AudioDataS16 Mojo struct has independent channel_count, frame_count, and data fields with no validation that data.size() >= channel_count * frame_count. The deprecated raw-pointer FromInterleaved() overload at speech_recognizer_impl.cc:341 reads channel_count * frame_count int16 values from the undersized data buffer, causing an OOB heap read. The read size is (channel_count * frame_count - data.size()) * 2 bytes, fully attacker-controlled per packet.

There is also a secondary integer overflow at line 345 where channel_count * frame_count * 16 / 8 is computed in int32 this can undersize the AudioChunk allocation, but requires a 1GB+ AudioBus to trigger so it’s only DoS.

The OOB-read data is converted to float, written to an AudioChunk, then FLAC-encoded (level 0, lossless) and uploaded to Google’s speech API. The int16→float32→int16→FLAC pipeline is bit-for-bit lossless (float32 mantissa exceeds int16 precision), so the heap bytes survive encoding intact.

Introduced in commit 881ab7a9d5 (“Add MediaStreamTrack support to the Web Speech API”, Evan Liu, 2024-08-01, https://chromium-review.googlesource.com/c/chromium/src/+/5631655).

VERSION Chrome Version: 146.0.7673.0 (trunk) affected since Chrome 129 (Aug 2024) Operating System: Ubuntu 22.04 / WSL2

REPRODUCTION CASE

Attached: poc-speech-oob-read.html (primary), poc-variable-read-sizes.html (shows 100/1024/4096/65536-byte reads)

The PoCs use MojoJS to simulate a compromised renderer crafting raw Mojo messages. A real attacker with renderer RCE would construct identical messages directly.

cd chromium/src ASAN_OPTIONS=detect_odr_violation=0 ./out/asan-debug/chrome
–no-sandbox –enable-blink-features=MojoJS,MojoJSTest –disable-gpu
file:///path/to/poc-speech-oob-read.html 2>&1 | tee /tmp/asan.log

Build: is_debug=true is_asan=true is_lsan=true is_component_build=true --no-sandbox is a WSL2 requirement; the bug is in the browser-process Mojo handler and is hit regardless of sandbox state. detect_odr_violation=0 is needed for ASan component builds (expected cppgc ODR violation).

The PoC binds media.mojom.SpeechRecognizer, sends Start() with an audio_forwarder receiver (bypasses mic permission), waits for the session to initialize, then sends AddAudioFromRenderer with channel_count=2, frame_count=1024, data=[10 values] reads 4076 bytes past the 20-byte allocation.

In release builds the OOB read is silent with no crash.

FOR CRASHES Type of crash: browser process (Chrome_IOThread) Crash State (full ASan log attached)

==35997==ERROR: AddressSanitizer: heap-buffer-overflow on address 0x717bf9f18bf4 at pc 0x754d48186b8b bp 0x714be7753730 sp 0x714be7753728 READ of size 2 at 0x717bf9f18bf4 thread T8 (Chrome_IOThread) #0 0x754d48186b8a in void media::AudioBus::CopyConvertFromInterleavedSourceToAudioBus<media::FixedSampleTypeTraits<short>>(…) media/base/audio_bus.h:494:27 #1 0x754d481869ce in void media::AudioBus::FromInterleavedPartial<media::FixedSampleTypeTraits<short>>(…) media/base/audio_bus.h:422:3 #2 0x754d4817f758 in void media::AudioBus::FromInterleaved<media::FixedSampleTypeTraits<short>>(…) media/base/audio_bus.h:390:3 #3 0x754d48173ecc in content::SpeechRecognizerImpl::AddAudioFromRenderer(…) content/browser/speech/speech_recognizer_impl.cc:341:9 #4 0x754d3e126c20 in media::mojom::SpeechRecognitionAudioForwarderStubDispatch::Accept(…) gen/media/mojo/mojom/speech_recognition_audio_forwarder.mojom.cc:192:13

0x717bf9f18bf4 is located 0 bytes after 20-byte region [0x717bf9f18be0,0x717bf9f18bf4) allocated by thread T8 (Chrome_IOThread) here: #0 0x5db19b595a5d in operator new(unsigned long) … #8 0x754d3e03269f in mojo::ArrayTraits<std::vector<short>>::Resize(…) mojo/public/cpp/bindings/array_traits.h:149:17

SUMMARY: AddressSanitizer: heap-buffer-overflow media/base/audio_bus.h:494:27 in void media::AudioBus::CopyConvertFromInterleavedSourceToAudioBus<…>(…)

SUGGESTED FIX

Attached: suggested-fix.patch

Validate data.size() == channel_count * frame_count (with base::CheckMul for overflow) before the FromInterleaved call, and switch to the span-based FromInterleaved overload which has CHECK_LE. The speech service at chrome/services/speech/ already does this validation. The browser-process SpeechRecognizerImpl just lacks it.

CREDIT INFORMATION

Reporter credit: Grischa Hauser

View on issue tracker