Skip to content

feat: Implement fallback for audio files - #420

Merged
lukasdotcom merged 1 commit into
mainfrom
audio-fallback
Sep 22, 2026
Merged

lukasdotcom merged 1 commit into
mainfrom
audio-fallback

Conversation

@lukasdotcom

Copy link
Copy Markdown
Member

Fallback for audio files that they get transcribed to text if the model doesn't support audio natively.

🤖 AI (if applicable)

  • The content of this PR was partly or fully generated using AI

@edward-ly

Copy link
Copy Markdown
Contributor

Were there any models you encountered that fall under this category? And did you notice any difference in output quality between transcript vs. direct audio file?

@lukasdotcom

Copy link
Copy Markdown
Member Author

Were there any models you encountered that fall under this category? And did you notice any difference in output quality between transcript vs. direct audio file?

Not all models support audio files. For example qwen does not support audio files. I think gemma understands audio files in a similar way anyway as it sometimes acts confused if you ask it to summarize an audio file and says it only sees a transcript.

For the topic of quality I don't know if for sure I would guess this is lower quality than an actual understanding, but this is really meant for models that don't support audio files anyway.

@edward-ly

Copy link
Copy Markdown
Contributor

I think gemma understands audio files in a similar way anyway as it sometimes acts confused if you ask it to summarize an audio file and says it only sees a transcript.

Hm, this doesn't sound very user-friendly. Would it be worthwhile to just prevent this case from happening by e.g. disabling audio file input for unsupported models, or would it be too much work?

@lukasdotcom

Copy link
Copy Markdown
Member Author

I think gemma understands audio files in a similar way anyway as it sometimes acts confused if you ask it to summarize an audio file and says it only sees a transcript.

Hm, this doesn't sound very user-friendly. Would it be worthwhile to just prevent this case from happening by e.g. disabling audio file input for unsupported models, or would it be too much work?

Sorry I think I wasn't clear enough gemma supports audio files natively, and had that problem. This method works fine the model thinks it was actually given an audio file and would respond to understanding what the audio is about (other than for cases where the transcript doesn't give enough information. Eg: what bird is that)

Signed-off-by: Lukas Schaefer <lukas@lschaefer.xyz>
}
return [[
'type' => 'text',
'text' => 'Filename:' . $file->getName() . "\nTranscription:\n" . $resultTask->getOutput()['output'],

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The force push fixed a typo here the \n was missing the n for one of them

@edward-ly

Copy link
Copy Markdown
Contributor

Hm, I still haven't been able to trigger the transcribe task yet. Any particular models or configuration settings I need to set to reach this situation?

@lukasdotcom

Copy link
Copy Markdown
Member Author

Hm, I still haven't been able to trigger the transcribe task yet. Any particular models or configuration settings I need to set to reach this situation?

It should work with any provider including the exapp. I personally used openai with whisper. Make sure to disable the audio input in this selection otherwise it will just send the file instead of the transcript.
image

@edward-ly edward-ly left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah I understand now. The transcript was already generated in a previous task, so we just retrieve it for the current task. I was also able to generate a new transcript if it didn't exist yet.

Tested and working for me with transcription model set to whisper-1 and provider model set to gpt-3.5-turbo.

@lukasdotcom
lukasdotcom merged commit c779716 into main Sep 22, 2026
25 checks passed
@lukasdotcom
lukasdotcom deleted the audio-fallback branch September 22, 2026 12:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants