Skip to content

very bad quality of output inference #2841

Description

@TheAnkulai

when i tried to inference I got this result with default settings
not_normal.wav
But with old rvc and same settings i got better result
good_result.wav
I dont know what's wrong, but that's sad
I made it on windows, rvc was used DirectML and inference started on cpu

Activity

  1. Hayri-maker commented on Aug 10, 2026

    @Hayri-maker
  2. TheAnkulai commented on Aug 12, 2026

    @TheAnkulai
    Author

    I don't know how, but it fixed by redownload and replace modules.py.
    In first attempts on all methods I got this error (first lines was different):

    [INFO]: device is not None, use cpu
      [INFO]    > call by:torchfcpe.tools.spawn_infer_cf_naive_mel_pe_from_pt
      [WARN] args.model.use_harmonic_emb is None; use default False
      [WARN]    > call by:torchfcpe.tools.spawn_cf_naive_mel_pe
    Traceback (most recent call last):
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\vc\modules.py", line 263, in vc_single
        audio_opt = self.pipeline.pipeline(
                    ^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\vc\pipeline.py", line 328, in pipeline
        self.vc(
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\vc\pipeline.py", line 222, in vc
        synthesized = run_cuda_graph(
                      ^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\tools\cuda_graph.py", line 190, in run_cuda_graph
        return function(*inputs)
               ^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\vc\pipeline.py", line 225, in <lambda>
        lambda phone, lengths, coarse, continuous, speaker: net_g.infer(
                                                            ^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\module\models.py", line 677, in infer
        o = self.dec(z * x_mask, nsff0, g=g, n_res=return_length2)
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\.venv\Lib\site-packages\torch\nn\modules\module.py", line 1553, in _wrapped_call_impl
        return self._call_impl(*args, **kwargs)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\.venv\Lib\site-packages\torch\nn\modules\module.py", line 1562, in _call_impl
        return forward_call(*args, **kwargs)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\module\models.py", line 484, in forward
        har_source, noi_source, uv = self.m_source(f0, self.upp)
                                     ^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\.venv\Lib\site-packages\torch\nn\modules\module.py", line 1553, in _wrapped_call_impl
        return self._call_impl(*args, **kwargs)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\.venv\Lib\site-packages\torch\nn\modules\module.py", line 1562, in _call_impl
        return forward_call(*args, **kwargs)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\module\models.py", line 391, in forward
        sine_wavs, uv, _ = self.l_sin_gen(x, upp)
                           ^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\.venv\Lib\site-packages\torch\nn\modules\module.py", line 1553, in _wrapped_call_impl
        return self._call_impl(*args, **kwargs)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\.venv\Lib\site-packages\torch\nn\modules\module.py", line 1562, in _call_impl
        return forward_call(*args, **kwargs)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\module\models.py", line 335, in forward
        sine_waves = self._f02sine(f0, upp) * self.sine_amp
                     ^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\infer\module\models.py", line 315, in _f02sine
        rad += F.pad(rad_acc, (0, 0, 1, -1))
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
      File "F:\Retrieval-based-Voice-Conversion-WebUI\.venv\Lib\site-packages\torch\nn\functional.py", line 4552, in pad
        return torch._C._nn.pad(input, pad, mode, value)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
    RuntimeError: torch-dml currently doesn't support a mix of positive and negative padding values, must be all positive or all negative.
    

    But now after replace modules.py it starts work properly. So, at this time this issue can be considered irrelevant.
    If you want to get .pth and .index that i used for this, there it is: https://dropmefiles.com/zm3vS (i can't upload zip in this message, idk why. And you able to download it only for 1 week). This model was trained not by me, so I don't know exactly dataset that was used for train. This is all I can give at this momentzz

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions