Skip to content

audio: fix mel preprocessor in LFM2 audio - #29403

Merged
ngxson merged 1 commit into
ggml-org:masterfrom
ykhrustalev:ykhrustalev/lfm2a-mel-fix
Sep 25, 2026
Merged

ngxson merged 1 commit into
ggml-org:masterfrom
ykhrustalev:ykhrustalev/lfm2a-mel-fix

Conversation

@ykhrustalev

Copy link
Copy Markdown
Contributor

Overview

The previous implementation produced mel features that differed from the reference by ~3-4%,

which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words.

This change:

  • uses log(x + 2^-24) instead of clamping to the log floor
  • uses a symmetric Hann window, equivalent to torch.hann_window(periodic=False)
  • adds the normalization epsilon to the standard deviation instead of inside the square root

Only the lfm2a preprocessor opts into these behaviors. Other audio preprocessors are unchanged.

Tested on top of 84e76d8 using llama-server with CUDA and temperature=0, compared against http://github.com/Liquid4All/liquid-audio fp32.

Additional information

Test set:

  • 200 LibriSpeech test-clean utterances (EN)
  • 200 Common Voice ja test utterances (JP)
  • identical 16 kHz audio passed to both implementations
Greedy transcript identical to liquid-audio Without fix With fix
EN F16 191/200 200/200
JP F32 187/200 200/200
JP F16 187/200 199/200

The remaining JP F16 difference is a comma and matches the reference implementation's own bf16 output.

Mel relative L2 error versus liquid-audio:

  • EN: 3.2% -> ~2e-6 median
  • JP: 3.9% -> ~2e-6 median

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES, with help of Claude and reviewed by human .

…the reference by ~3-4%,

which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese
test utterances. In Japanese, some differences changed entire words.

This change:

* uses `log(x + 2^-24)` instead of clamping to the log floor
* uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)`
* adds the normalization epsilon to the standard deviation instead of inside the square root

Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged.

Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against
http://github.com/Liquid4All/liquid-audio fp32.

Test set:

* 200 LibriSpeech `test-clean` utterances (EN)
* 200 Common Voice `ja` test utterances (JP)
* identical 16 kHz audio passed to both implementations

| Greedy transcript identical to `liquid-audio` | Without fix |    With fix |
| --------------------------------------------- | ----------: | ----------: |
| EN F16                                        |     191/200 | 200/200 |
| JP F32                                        |     187/200 | 200/200 |
| JP F16                                        |     187/200 | 199/200 |

The remaining JP F16 difference is a comma and matches the reference implementation's own bf16
output.

Mel relative L2 error versus `liquid-audio`:

* EN: 3.2% -> ~2e-6 median
* JP: 3.9% -> ~2e-6 median
@ykhrustalev
ykhrustalev requested review from a team and ggerganov as code owners September 25, 2026 00:12
@github-actions github-actions Bot added testing Everything test related mtmd Related to multimodal functionality (video/image/audio) labels Sep 25, 2026
@CISC

CISC commented Sep 25, 2026

Copy link
Copy Markdown
Member

cc/ @tdakhran

@ngxson

ngxson commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator

/bot review

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Automated code review

Reviewed the diff against tools/mtmd/mtmd-audio.cpp/h, tests/test-mtmd-impl.cpp, and the surrounding preprocessor infrastructure. This is a well-scoped, single-purpose numerical fix, gated correctly so only the conformer (lfm2a) preprocessor opts into the new behaviors. I verified the key claims against the code:

Verification of correctness

  • Symmetric window: fill_hann_window(len, periodic=false) yields 0.5*(1-cos(2*pi*i/(len-1))), which is exactly torch.hann_window(periodic=False). Matches the NeMo reference. Other preprocessors keep periodic=true and are untouched.
  • Log guard: mel_floor_add changes log(max(x, floor)) to log(x + floor), matching NeMo's default log_zero_guard_type="add". The conformer does not set params.mel_floor, so it uses the existing default 5.960464477539063e-08f, which is exactly 2^-24, NeMo's default log_zero_guard. The floor is never applied to a negative value (filters and squared magnitudes are non-negative), so log cannot see <= 0.
  • Std epsilon: sqrt(var) + 1e-5 vs sqrt(var + 1e-5) is applied only behind the flag; unbiased variance (var /= n-1) is preserved. The PR's empirical data (~2e-6 median relative L2 error vs liquid-audio, 200/200 and 199/200 transcript matches) strongly supports this reading of the reference.
  • Test: correctly wired into the existing test-mtmd-impl.cpp registry via MAKE_TEST (no new test file), and n_len == 101 is consistent with center padding, window 400, hop 160, 16000 samples. The new clip_hparams constructor is a reasonable, minimal addition for testability; the reference member's lifetime is safe in the test.

No security concerns: no new attacker-controlled inputs, loops stay bounded, no allocation changes.

Nits (non-blocking)

  1. The conformer silently relies on filter_params::mel_floor's default being 2^-24 for the "add" guard. To make the NeMo contract explicit and protect against a future change to that default, consider setting params.mel_floor explicitly in mtmd_audio_preprocessor_conformer::preprocess (it already sets mel_floor_add, so adding the value it guards with would be consistent).
  2. The test asserts 4 sampled mel values at 1e-3 tolerance. That is a reasonable smoke check, but a single scalar drift could pass it; if you still have the reference mel handy, an aggregate assertion (e.g. max abs difference over the full frame) would make the test much harder to pass accidentally.

Nothing blocking here. The changes are minimal, correctly scoped to the conformer path, and the PR description provides the exact kind of reference-comparison data this numerical change needs.

This review was generated automatically by pi coding agent using zai-org/GLM-5.3. It may contain mistakes. Maintainers make the final call.

@ngxson
ngxson merged commit fcc8915 into ggml-org:master Sep 25, 2026
12 checks passed
feal87 added a commit to feal87/myllama.cpp that referenced this pull request Sep 25, 2026
Merge upstream commits:
  - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364)
  - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294)
  - metal: split fa kernels into per-dtype libraries (ggml-org#29329)
  - metal: FWHT kernels for block widths above 512 (ggml-org#29095)
  - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393)
  - common: extract shared unicode path/string helpers (ggml-org#29415)
  - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432)
  - rpc: include nb in the get_alloc_size cache key (ggml-org#29283)
  - [SYCL] support sparse FA (ggml-org#28796)
  - musa: fix PH1 operator failures and build issues (ggml-org#29193)
  - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231)
  - opencl: add q5_k bin kernel (ggml-org#29401)
  - hexagon: add q5_k quant type support (ggml-org#29123)
  - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404)
  - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403)
  - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409)
  - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422)
  - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417)

Assisted-by: Pi
sky-mighty pushed a commit to sky-mighty/llama.cpp that referenced this pull request Sep 26, 2026
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese
test utterances. In Japanese, some differences changed entire words.

This change:

* uses `log(x + 2^-24)` instead of clamping to the log floor
* uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)`
* adds the normalization epsilon to the standard deviation instead of inside the square root

Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged.

Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against
http://github.com/Liquid4All/liquid-audio fp32.

Test set:

* 200 LibriSpeech `test-clean` utterances (EN)
* 200 Common Voice `ja` test utterances (JP)
* identical 16 kHz audio passed to both implementations

| Greedy transcript identical to `liquid-audio` | Without fix |    With fix |
| --------------------------------------------- | ----------: | ----------: |
| EN F16                                        |     191/200 | 200/200 |
| JP F32                                        |     187/200 | 200/200 |
| JP F16                                        |     187/200 | 199/200 |

The remaining JP F16 difference is a comma and matches the reference implementation's own bf16
output.

Mel relative L2 error versus `liquid-audio`:

* EN: 3.2% -> ~2e-6 median
* JP: 3.9% -> ~2e-6 median
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese
test utterances. In Japanese, some differences changed entire words.

This change:

* uses `log(x + 2^-24)` instead of clamping to the log floor
* uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)`
* adds the normalization epsilon to the standard deviation instead of inside the square root

Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged.

Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against
http://github.com/Liquid4All/liquid-audio fp32.

Test set:

* 200 LibriSpeech `test-clean` utterances (EN)
* 200 Common Voice `ja` test utterances (JP)
* identical 16 kHz audio passed to both implementations

| Greedy transcript identical to `liquid-audio` | Without fix |    With fix |
| --------------------------------------------- | ----------: | ----------: |
| EN F16                                        |     191/200 | 200/200 |
| JP F32                                        |     187/200 | 200/200 |
| JP F16                                        |     187/200 | 199/200 |

The remaining JP F16 difference is a comma and matches the reference implementation's own bf16
output.

Mel relative L2 error versus `liquid-audio`:

* EN: 3.2% -> ~2e-6 median
* JP: 3.9% -> ~2e-6 median
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese
test utterances. In Japanese, some differences changed entire words.

This change:

* uses `log(x + 2^-24)` instead of clamping to the log floor
* uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)`
* adds the normalization epsilon to the standard deviation instead of inside the square root

Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged.

Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against
http://github.com/Liquid4All/liquid-audio fp32.

Test set:

* 200 LibriSpeech `test-clean` utterances (EN)
* 200 Common Voice `ja` test utterances (JP)
* identical 16 kHz audio passed to both implementations

| Greedy transcript identical to `liquid-audio` | Without fix |    With fix |
| --------------------------------------------- | ----------: | ----------: |
| EN F16                                        |     191/200 | 200/200 |
| JP F32                                        |     187/200 | 200/200 |
| JP F16                                        |     187/200 | 199/200 |

The remaining JP F16 difference is a comma and matches the reference implementation's own bf16
output.

Mel relative L2 error versus `liquid-audio`:

* EN: 3.2% -> ~2e-6 median
* JP: 3.9% -> ~2e-6 median
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese
test utterances. In Japanese, some differences changed entire words.

This change:

* uses `log(x + 2^-24)` instead of clamping to the log floor
* uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)`
* adds the normalization epsilon to the standard deviation instead of inside the square root

Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged.

Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against
http://github.com/Liquid4All/liquid-audio fp32.

Test set:

* 200 LibriSpeech `test-clean` utterances (EN)
* 200 Common Voice `ja` test utterances (JP)
* identical 16 kHz audio passed to both implementations

| Greedy transcript identical to `liquid-audio` | Without fix |    With fix |
| --------------------------------------------- | ----------: | ----------: |
| EN F16                                        |     191/200 | 200/200 |
| JP F32                                        |     187/200 | 200/200 |
| JP F16                                        |     187/200 | 199/200 |

The remaining JP F16 difference is a comma and matches the reference implementation's own bf16
output.

Mel relative L2 error versus `liquid-audio`:

* EN: 3.2% -> ~2e-6 median
* JP: 3.9% -> ~2e-6 median

(cherry picked from commit fcc8915)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

mtmd Related to multimodal functionality (video/image/audio) testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants