Repository navigation
Add voice resolution and defaulting to the Text to Speech experiment - #5
Merged
dkotter merged 7 commits intoAug 26, 2026
Merged
Conversation
Providers like ElevenLabs require a voice ID for text to speech and fail when none is configured. Introduce a Voice_Resolver that reads the resolved model's metadata for voices declared as supported values on the outputSpeechVoice option: - The Voice setting becomes a dynamic select when the provider declares its voices, defaulting to the first available voice at generation time when the setting is empty. Falls back to the existing free-text field when no voices are declared. - A new wpai_tts_default_voice filter lets site or provider plugins supply a default voice even when the model metadata declares none. - The raw provider exception for a missing voice is mapped to an actionable error pointing at the Voice setting. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message. To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook. |
…ice no longer exists
…o it's used for each chunk
dkotter
approved these changes
Aug 26, 2026
dkotter
merged commit Aug 26, 2026
ed2d9b5
into
dkotter:feature/text-to-speech
25 of 28 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What?
Follows up on the ElevenLabs discussion in WordPress#888. Adds a
Voice_Resolverto the Text to Speech experiment that dynamically resolves the voices a provider supports and defaults to the first one when no voice is configured.Why?
Some providers (e.g. ElevenLabs) require a voice ID for text to speech. The experiment currently only sends a voice when the free-text Voice setting is filled in, so those providers fail out of the box with a raw SDK exception.
How?
Voice_Resolverreads the resolved TTS model's metadata for supported values declared on theoutputSpeechVoiceoption (mirroringSpeech_Generator's provider/model resolution order, cached in a transient for an hour). This establishes a generic contract: providers can declare their voice IDs assupportedValuesonSupportedOption( OptionEnum::outputSpeechVoice() ).elements/DataForm pattern, no JS changes) with a "Provider default (first available voice)" option, andSpeech_Generator::generate_chunk()fills an empty voice with the first declared one — covering both the cron path and theai/speech-generationability.wpai_tts_default_voicefilter lets site or provider plugins supply a default voice, andNote: the ElevenLabs provider 0.3.0 has been released with the complementary fixes — it now declares inline
outputFileTypesupport and falls back to a premade default voice when none is configured. The two fixes work together: an explicit voice from this plugin always overrides the provider default.Use of AI Tools
AI assistance: Yes
Tool(s): Claude Code
Model(s): Fable 5
Used for: Planning and implementation. Review and testing done by me
Testing Instructions
add_filter( 'wpai_tts_default_voice', fn () => '21m00Tcm4TlvDq8ikWAM' );— generation succeeds via the cron path and viaPOST /wp-json/wp-abilities/v1/abilities/ai/speech-generation/runSupportedOption( OptionEnum::outputSpeechVoice(), array( 'v1', 'v2' ) )— the Voice setting renders as a select and leaving it on "Provider default" usesv1Changelog Entry
🤖 PR message Partly generated with Claude Code