SpeakTrue

Eleven v4 compatibility

Model contract

SpeakTrue discovers eleven_v4 and eleven_v4_turbo through the existing provider model catalog. These IDs use the following explicit capability gates; unknown future models do not inherit them through a prefix or turbo match.

Capability Eleven v4 and Eleven v4 Turbo
Request text limit 10,000 Unicode code points
Provider voice controls Stability and Similarity
Provider Speed, Style, Speaker Boost, legacy latency optimization Omitted
SSML wrapper and break tags Unsupported; do not insert
Live WebSocket wss://api.elevenlabs.io/v1/text-to-dialogue/stream-input
Single-use token endpoint /v1/single-use-token/ttd_websocket
Connection token type ttd_websocket
Additive session protocol field websocket_protocol: "text_to_dialogue"

Playback controls remain separate from provider generation controls. Hidden provider settings remain saved so switching back to an older model restores the user’s preferences. Existing model defaults remain unchanged.

Input editors preserve the draft when switching to a model with a lower limit; generation validates against the selected model rather than trimming the draft. Web controls are scoped to each TTS, STS, and Settings panel’s own selection.

The native and web live clients register one voice in the initial voices array, send speech in inputs entries with the registered voice_id, use flush for short text, and finalize with close_socket. Dialogue responses use is_final; legacy isFinal remains accepted. Provider credentials remain on the server; clients receive only a single-use token.

Android queues text and the final request while the socket is connecting. Its initial frame, queued text, and final frame are serialized; connecting is not treated as closed during drainage. Socket callbacks and keep-alive timers are bound to their originating connection so a stopped session cannot alter a new one.

The web speech studio also keeps its existing server HTTP generation path for ordinary TTS and batch workflows.

Rollout and recovery

Backend session minting and client frame handling must be delivered together for v4 live voice. Changing the socket URL alone is insufficient: the old tts_websocket token is rejected by the dialogue endpoint. Existing models retain their prior provider settings and socket routes. Selecting an existing model provides the compatibility fallback.

Device playback, intelligibility, speaker similarity, and production workflow acceptance require separate verification after delivery. Current delivery evidence is recorded below.

Provider evidence — 2026-09-29

The configured provider account returned both IDs with TTS capability and 10,000-character limits in a live catalog read. Small synthetic HTTP requests for both models returned PCM audio. Dialogue requests for both returned audio and a final frame when authenticated with ttd_websocket tokens. A diagnostic request using tts_websocket reproduced invalid_token_type. These checks validate the provider contract, not a deployed SpeakTrue application flow. No generated audio, credentials, or raw provider payloads are stored here.

Sources:

Some reference wording still mentions only v3. The current streaming guide includes v4; the token type and both 10,000-character caps above were checked against the live provider rather than inferred from the older wording.

Local verification — 2026-09-29

Deployment — 2026-09-29

Recovery: both web hosts retain speaktrue-web:rollback-15ca9d5503dc-v4 and /tmp/speaktrue-v4-web-rollback.json with the previous revision and asset version. Prior backend bundles and their deployment inventory are retained locally under /private/tmp/speaktrue-v4-backend-rollback and /private/tmp/speaktrue-v4-deploy-function-inventory.json. Restore only the affected functions with their original JWT modes, and verify health after restoring a web image/revision. These temporary recovery files need retention for any later rollback.

Apple processing rejection and packaging repair — 2026-09-29

Apple subsequently rejected build 69 with ITMS-90208 for the embedded onnxruntime.framework. The archive’s framework plist claims minimum iOS 15.1, but its executable’s LC_BUILD_VERSION requires 26.0. The app minimum is already 26.0. The upstream ONNX artifact is a static library: Xcode removes its static executable during embedding and injects a codeless-framework stub at the app deployment target, retaining the original upstream plist.

The app now normalizes the generated ONNX framework plist to the embedded executable’s actual minimum OS after copying, refreshing the framework signature before app signing. It does not modify the upstream package or lower the executable requirement. The packaging guard rejects frameworks requiring a newer OS than the app and can verify a completed archive before upload:

python3 scripts/verify_ios_framework_minimum_os.py \
  --app /path/to/SpeakTrue.xcarchive/Products/Applications/SpeakTrue.app
python3 -m unittest discover -s scripts/tests \
  -p test_ios_framework_minimum_os.py

The replacement build is iOS 1.4.1 (70); Android remains 1.4 (35). The new archive passes the read-only guard and strict nested/app code-signature verification. Its log confirms normalization runs after ONNX stub generation and before final app signing. All 10 guard regression tests and the complete local Android CI-equivalent gate passed. The exported IPA also passes the guard and strict app signature check. The old build-69 archive is rejected by the guard for the reported 15.1/26.0 mismatch.

Xcode reports upload/export succeeded for 1.4.1 (70). The missing ONNX dSYM warning remains for the generated stub and did not block upload. Apple processing and tester availability have not been read back; an upload acknowledgment alone does not prove Apple processing succeeded.