SpeakTrue

Local TTS UI simplification plan

Status: Implemented in the uncommitted iOS and Android working tree on 2026-09-22, based on bf9fc0310d3d1a2a39c1b4bce1f493bc83e7dc39. Simulator/unit verification and physical-device acceptance are tracked separately; no release or deployment is claimed.

Outcome and boundaries

Make the ordinary path choose a voice → enter text → generate → play or export. When a prerequisite is missing, show one direct next action in the persistent action area. Keep model comparison, voice administration, sampler controls, chunking details, history, and performance diagnostics available without making them part of first use.

Preserve all existing model, reference, download, generation, history, and export contracts. A model download remains an explicit action with its approximate size and device suitability visible before it starts. Save to Soundboard remains an explicit cloud action and must say that it uploads the generated WAV and metadata while reference audio stays local. The uncommitted clip-metadata/Soundboard changes already in this working tree must be integrated or isolated before UI implementation; this plan does not take ownership of that work.

Baseline evidence before this UI slice

Implemented screen behavior

The default surface contains one compact Model · status row, a Voice picker, text input, language only when the selected model offers more than one language, a pinned primary action, and the latest result when one exists. The model row opens Manage models; an empty voice state offers Add voice directly. Settings, section preview, history, and diagnostics remain discoverable secondary sections. Keep About/privacy/attribution available from the screen.

The primary action is derived from the first actionable state, in this order:

State Primary action or message
Generating or preparing Show current stage/progress and the existing safe cancel action.
Recording a reference Stop recording; retain the recording flow and do not offer Generate.
Engine unavailable or device unsupported Explain why; offer Choose another model only when an eligible alternative exists.
Selected model not loaded Download & load or Load model, with size and support status before download.
Reference voice missing Add voice; open the existing voice form directly.
Required reference transcript missing Add transcript; edit and explicitly save the selected reference’s transcript.
Text empty Focus the text editor; keep the reason beside the action.
Ready Generate locally.

Do not auto-download, auto-upload, silently alter a saved voice, or change the model during an active generation. On first use, recommend an already downloaded, supported model first; otherwise recommend a supported catalog model with its size. Preserve a user’s explicit selection and show a compatible alternative if it is unsupported. The recommendation means passes catalog preflight, not physical-device quality or reliability proof.

After success, show playback, captured voice/model, Regenerate, and Export. Export presents distinct Share, Save WAV, and Save to Soundboard choices; the cloud option names the upload. If regeneration needs the original model after unload or a model switch, offer to prepare that model before retrying instead of returning only a disabled action. Keep the generated result associated with its captured settings even if the editor changes.

Implementation slices

1. Readiness and direct recovery

Add a small presentation state that maps existing readiness facts to one label, explanation, target, and enabled action. Use the same state table on iOS and Android while keeping each platform’s engine and device rules authoritative. Wire the pinned action area to open the existing model/voice controls or focus transcript/text, and to stop an active reference recording. Remove repeated general status/error copy from Android’s header and editor once the actionable message is shown near the primary action.

Likely paths: ios/SpeakTrue/LocalQwenViewModel.swift, ios/SpeakTrue/LocalQwenPrototypeView.swift, android/app/src/main/java/com/speaktrue/features/tts/local/presentation/LocalTtsUiState.kt, LocalTtsLabScreen.kt, LocalTtsLabComponents.kt, and their focused tests.

Acceptance: every blocker has exactly one understandable next step; unsupported or unavailable models never offer download/load; preparation/generation retain cancellation; recording shows Stop recording even if another voice is selected; draft text and the selected voice survive setup navigation.

2. Model and voice setup

Make the selected model row open a single model-management area with one context-sensitive preparation action. Keep refresh, unload, delete, detailed preflight, and the full catalog as secondary controls. Add a direct empty-voice action and progressively reveal recording/import, name, transcript, and save fields. Keep transcript optional when the selected engine does not require it and retain existing authorization/save confirmations.

Android’s current transcript field is a new-voice draft, while readiness reads the saved selected reference. Add an explicit Edit transcript operation for an existing selected reference: populate it from that reference, save only its transcript after user action, preserve profile/reference IDs and audio, and never upload it. Update the root reference or additionalReferences entry as appropriate. iOS should likewise direct Add transcript to the reference actually used for generation. A model switch from transcript-optional PocketTTS to transcript-required Qwen is the key recovery case; merely focusing Android’s current creation field would not resolve it.

Likely additional paths: android/app/src/main/java/com/speaktrue/features/tts/local/data/LocalVoiceProfileStore.kt, LocalTtsViewModel.kt, their store/ViewModel tests, and the iOS voice profile store if its current reference edit path cannot persist the selected transcript directly.

Use catalog support and downloaded state for a first-use recommendation. If a saved or loaded selection exists, retain it; when unsupported, suggest rather than silently switch. Do not promise that the lower-RAM model has passed physical-device listening or thermal checks.

Acceptance: first use reaches preparation and voice creation without searching disclosures; selected model size and support are clear; no background download begins; model/voice deletion and cloud profile operations remain reachable; saving a transcript on an existing reference makes a transcript-required model ready without duplicating or replacing its audio.

3. Composer and result

Move section-count explanation and family sampling controls out of the primary reading path while keeping their current behavior and per-family defaults. Avoid changing chunking, audio settings, or generation semantics in this UI slice. Consolidate result actions under Export, label the Soundboard upload, add captured voice/model on Android, and recover the original model before regeneration when necessary.

Acceptance: ready state shows voice, text, optional language, and Generate without an expanded technical panel; all existing settings and export paths remain available; a result remains correctly identified after editor changes; regenerate works or gives the direct prepare action after model unload/switch.

4. Verification and closeout

Delivery boundary

The UI changes were integrated as reviewable slices on top of the existing uncommitted clip-metadata/Soundboard work. They do not change backend contracts or providers. The implementation stays uncommitted; a signed device walkthrough and listening/thermal acceptance remain separate from source and Simulator verification.