Skip to content

Add voice resolution and defaulting to the Text to Speech experiment - #5

Merged
dkotter merged 7 commits into
dkotter:feature/text-to-speechfrom
saarnilauri:tts-voice-resolver
Aug 26, 2026
Merged

dkotter merged 7 commits into
dkotter:feature/text-to-speechfrom
saarnilauri:tts-voice-resolver

Conversation

@saarnilauri

@saarnilauri saarnilauri commented Aug 24, 2026 •

Copy link
Copy Markdown

What?

Follows up on the ElevenLabs discussion in WordPress#888. Adds a Voice_Resolver to the Text to Speech experiment that dynamically resolves the voices a provider supports and defaults to the first one when no voice is configured.

Why?

Some providers (e.g. ElevenLabs) require a voice ID for text to speech. The experiment currently only sends a voice when the free-text Voice setting is filled in, so those providers fail out of the box with a raw SDK exception.

How?

  • New Voice_Resolver reads the resolved TTS model's metadata for supported values declared on the outputSpeechVoice option (mirroring Speech_Generator's provider/model resolution order, cached in a transient for an hour). This establishes a generic contract: providers can declare their voice IDs as supportedValues on SupportedOption( OptionEnum::outputSpeechVoice() ).
  • When voices are declared, the Voice setting renders as a select (existing elements/DataForm pattern, no JS changes) with a "Provider default (first available voice)" option, and Speech_Generator::generate_chunk() fills an empty voice with the first declared one — covering both the cron path and the ai/speech-generation ability.
  • When no voices are declared (all providers today), the field stays free-text and behavior is unchanged, except:
    • a new wpai_tts_default_voice filter lets site or provider plugins supply a default voice, and
    • the raw "outputSpeechVoice is required" exception is mapped to an actionable error pointing at the Voice setting.

Note: the ElevenLabs provider 0.3.0 has been released with the complementary fixes — it now declares inline outputFileType support and falls back to a premade default voice when none is configured. The two fixes work together: an explicit voice from this plugin always overrides the provider default.

Use of AI Tools

AI assistance: Yes
Tool(s): Claude Code
Model(s): Fable 5
Used for: Planning and implementation. Review and testing done by me

Testing Instructions

  1. Connect a provider that requires a voice (e.g. the ElevenLabs provider, version 0.3.0 or later) and enable the Text to Speech experiment
  2. With the Voice setting empty, generate audio for a post — with ElevenLabs 0.3.0's default-voice fallback it succeeds; with a voice-requiring provider that has no default, the editor now shows the actionable "requires a voice" error instead of a raw exception
  3. Add add_filter( 'wpai_tts_default_voice', fn () => '21m00Tcm4TlvDq8ikWAM' ); — generation succeeds via the cron path and via POST /wp-json/wp-abilities/v1/abilities/ai/speech-generation/run
  4. Enter a voice ID in the Voice setting — it takes precedence
  5. Register a test provider whose TTS model declares SupportedOption( OptionEnum::outputSpeechVoice(), array( 'v1', 'v2' ) ) — the Voice setting renders as a select and leaving it on "Provider default" uses v1
  6. With OpenAI connected, verify no regression (empty voice → provider default path)

Changelog Entry

Added - Voice selection for the Text to Speech experiment: voices declared by the provider render as a dropdown, the first declared voice is used by default, and a wpai_tts_default_voice filter allows supplying a default for providers that require one.

🤖 PR message Partly generated with Claude Code

Providers like ElevenLabs require a voice ID for text to speech and fail
when none is configured. Introduce a Voice_Resolver that reads the
resolved model's metadata for voices declared as supported values on the
outputSpeechVoice option:

- The Voice setting becomes a dynamic select when the provider declares
  its voices, defaulting to the first available voice at generation time
  when the setting is empty. Falls back to the existing free-text field
  when no voices are declared.
- A new wpai_tts_default_voice filter lets site or provider plugins
  supply a default voice even when the model metadata declares none.
- The raw provider exception for a missing voice is mapped to an
  actionable error pointing at the Voice setting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

The following accounts have interacted with this PR and/or linked issues. I will continue to update these lists as activity occurs. You can also manually ask me to refresh this list by adding the props-bot label.

If you're merging code through a pull request on GitHub, copy and paste the following into the bottom of the merge commit message.

Co-authored-by: saarnilauri <laurisaarni@git.wordpress.org>
Co-authored-by: dkotter <dkotter@git.wordpress.org>

To understand the WordPress project's expectations around crediting contributors, please review the Contributor Attribution page in the Core Handbook.

@dkotter
dkotter merged commit ed2d9b5 into dkotter:feature/text-to-speech Aug 26, 2026
25 of 28 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants