Repository navigation
audio: fix mel preprocessor in LFM2 audio - #29403
Conversation
…the reference by ~3-4%, which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words. This change: * uses `log(x + 2^-24)` instead of clamping to the log floor * uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)` * adds the normalization epsilon to the standard deviation instead of inside the square root Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged. Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against http://github.com/Liquid4All/liquid-audio fp32. Test set: * 200 LibriSpeech `test-clean` utterances (EN) * 200 Common Voice `ja` test utterances (JP) * identical 16 kHz audio passed to both implementations | Greedy transcript identical to `liquid-audio` | Without fix | With fix | | --------------------------------------------- | ----------: | ----------: | | EN F16 | 191/200 | 200/200 | | JP F32 | 187/200 | 200/200 | | JP F16 | 187/200 | 199/200 | The remaining JP F16 difference is a comma and matches the reference implementation's own bf16 output. Mel relative L2 error versus `liquid-audio`: * EN: 3.2% -> ~2e-6 median * JP: 3.9% -> ~2e-6 median
|
cc/ @tdakhran |
|
/bot review |
Automated code reviewReviewed the diff against Verification of correctness
No security concerns: no new attacker-controlled inputs, loops stay bounded, no allocation changes. Nits (non-blocking)
Nothing blocking here. The changes are minimal, correctly scoped to the conformer path, and the PR description provides the exact kind of reference-comparison data this numerical change needs. This review was generated automatically by pi coding agent using |
Merge upstream commits: - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364) - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294) - metal: split fa kernels into per-dtype libraries (ggml-org#29329) - metal: FWHT kernels for block widths above 512 (ggml-org#29095) - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393) - common: extract shared unicode path/string helpers (ggml-org#29415) - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432) - rpc: include nb in the get_alloc_size cache key (ggml-org#29283) - [SYCL] support sparse FA (ggml-org#28796) - musa: fix PH1 operator failures and build issues (ggml-org#29193) - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231) - opencl: add q5_k bin kernel (ggml-org#29401) - hexagon: add q5_k quant type support (ggml-org#29123) - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404) - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403) - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409) - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422) - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417) Assisted-by: Pi
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words. This change: * uses `log(x + 2^-24)` instead of clamping to the log floor * uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)` * adds the normalization epsilon to the standard deviation instead of inside the square root Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged. Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against http://github.com/Liquid4All/liquid-audio fp32. Test set: * 200 LibriSpeech `test-clean` utterances (EN) * 200 Common Voice `ja` test utterances (JP) * identical 16 kHz audio passed to both implementations | Greedy transcript identical to `liquid-audio` | Without fix | With fix | | --------------------------------------------- | ----------: | ----------: | | EN F16 | 191/200 | 200/200 | | JP F32 | 187/200 | 200/200 | | JP F16 | 187/200 | 199/200 | The remaining JP F16 difference is a comma and matches the reference implementation's own bf16 output. Mel relative L2 error versus `liquid-audio`: * EN: 3.2% -> ~2e-6 median * JP: 3.9% -> ~2e-6 median
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words. This change: * uses `log(x + 2^-24)` instead of clamping to the log floor * uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)` * adds the normalization epsilon to the standard deviation instead of inside the square root Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged. Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against http://github.com/Liquid4All/liquid-audio fp32. Test set: * 200 LibriSpeech `test-clean` utterances (EN) * 200 Common Voice `ja` test utterances (JP) * identical 16 kHz audio passed to both implementations | Greedy transcript identical to `liquid-audio` | Without fix | With fix | | --------------------------------------------- | ----------: | ----------: | | EN F16 | 191/200 | 200/200 | | JP F32 | 187/200 | 200/200 | | JP F16 | 187/200 | 199/200 | The remaining JP F16 difference is a comma and matches the reference implementation's own bf16 output. Mel relative L2 error versus `liquid-audio`: * EN: 3.2% -> ~2e-6 median * JP: 3.9% -> ~2e-6 median
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words. This change: * uses `log(x + 2^-24)` instead of clamping to the log floor * uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)` * adds the normalization epsilon to the standard deviation instead of inside the square root Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged. Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against http://github.com/Liquid4All/liquid-audio fp32. Test set: * 200 LibriSpeech `test-clean` utterances (EN) * 200 Common Voice `ja` test utterances (JP) * identical 16 kHz audio passed to both implementations | Greedy transcript identical to `liquid-audio` | Without fix | With fix | | --------------------------------------------- | ----------: | ----------: | | EN F16 | 191/200 | 200/200 | | JP F32 | 187/200 | 200/200 | | JP F16 | 187/200 | 199/200 | The remaining JP F16 difference is a comma and matches the reference implementation's own bf16 output. Mel relative L2 error versus `liquid-audio`: * EN: 3.2% -> ~2e-6 median * JP: 3.9% -> ~2e-6 median
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words. This change: * uses `log(x + 2^-24)` instead of clamping to the log floor * uses a symmetric Hann window, equivalent to `torch.hann_window(periodic=False)` * adds the normalization epsilon to the standard deviation instead of inside the square root Only the `lfm2a` preprocessor opts into these behaviors. Other audio preprocessors are unchanged. Tested on top of 84e76d8 using `llama-server` with CUDA and `temperature=0`, compared against http://github.com/Liquid4All/liquid-audio fp32. Test set: * 200 LibriSpeech `test-clean` utterances (EN) * 200 Common Voice `ja` test utterances (JP) * identical 16 kHz audio passed to both implementations | Greedy transcript identical to `liquid-audio` | Without fix | With fix | | --------------------------------------------- | ----------: | ----------: | | EN F16 | 191/200 | 200/200 | | JP F32 | 187/200 | 200/200 | | JP F16 | 187/200 | 199/200 | The remaining JP F16 difference is a comma and matches the reference implementation's own bf16 output. Mel relative L2 error versus `liquid-audio`: * EN: 3.2% -> ~2e-6 median * JP: 3.9% -> ~2e-6 median (cherry picked from commit fcc8915)
Overview
The previous implementation produced mel features that differed from the reference by ~3-4%,
which resulted in different greedy transcripts for 4.5% of English and 6.5% of Japanese test utterances. In Japanese, some differences changed entire words.
This change:
log(x + 2^-24)instead of clamping to the log floortorch.hann_window(periodic=False)Only the
lfm2apreprocessor opts into these behaviors. Other audio preprocessors are unchanged.Tested on top of 84e76d8 using
llama-serverwith CUDA andtemperature=0, compared against http://github.com/Liquid4All/liquid-audio fp32.Additional information
Test set:
test-cleanutterances (EN)jatest utterances (JP)liquid-audioThe remaining JP F16 difference is a comma and matches the reference implementation's own bf16 output.
Mel relative L2 error versus
liquid-audio:Requirements