Current UI correction (build 62): one Temperature (variation) slider replaces the duplicate Variation and Temperature controls. Both previously modified the same value. Sampling ranges/defaults are unchanged; references to two separate controls below are historical planning only.
Status: historical implementation sequence from 2026-09-05. Settings contracts and per-voice defaults have since been implemented; build 59 makes the current editor values effective immediately for the next run. The current guide and multi-model plan describe shipped source; later exploratory items below remain proposals unless marked implemented there.
Give the iOS Local Qwen Lab understandable generation controls, persistent settings per voice, and advanced tuning without suggesting unsupported model capabilities. Applies to the existing 0.6B Base 4/6/8-bit variants. Keep generation local and use one selected reference with its matching transcript. No backend or Android feature changes are planned.
Initial build-40 baseline: model selection, language, default voice, selectable reference recordings, section previews and a fixed 5 ms crossfade. Generation used temperature 0.9, top-p 1.0, repetition penalty 1.1 and a 4096 generated-audio-token ceiling; top-k defaulted to 0. Sections were bounded at 420 characters. Build 40 requested increased memory after a confirmed build 39 per-process-limit termination. This baseline is retained as planning history, not current behavior.
| Control | Planned behavior | Initial default / proposed bounds |
|---|---|---|
| Voice and recording | Keep existing selectors; show which reference is active | Saved default voice and preferred recording |
| Language | Supported Qwen languages plus Auto after verification | Preserve existing English default |
| Model | Keep downloaded/loaded status and memory guidance | Preserve user’s selected model; never switch silently |
| Variation | Main slider mapped directly to temperature; numerical value in Advanced | 0.9; provisional 0.3–1.2 |
| Speaking speed | Pitch-preserving processing of the generated file | 1.0×; provisional 0.75–1.25× |
| Paragraph pause | Explicit silence between paragraphs only | 0 ms additional; provisional 0–1000 ms |
| Top-p / top-k | Advanced sampling filters; explain 0 disables top-k | 1.0 / 0; provisional 0.5–1.0 / 0–100 |
| Repetition penalty | Advanced audio-token repetition control | 1.1; provisional 1.0–1.3 |
| Audio token ceiling | Advanced maximum per section; may truncate speech | 4096; provisional 512–4096 |
| Section length | Advanced character bound with live preview | 420; provisional 120–420 |
| Join crossfade | Advanced adjustment at continuous section boundaries | 5 ms; provisional 0–10 ms |
Bounds are conservative product proposals for testing, not proven quality ranges. All advanced settings have help text and Reset defaults. Variation and temperature are one value, not competing controls. The SDK also forwards min-p; leave it at 0 in this scope to avoid overlapping sampling controls. Seed/KV-cache fields in the shared parameter type are not sufficient proof of support: the current Qwen settings adapter does not forward them. Do not expose those controls without separate implementation and verification.
Do not add emotion, similarity, style-strength, multi-reference fusion, or an exact generated-duration promise. Consistent/Balanced/Varied presets are a later step after listening comparisons, with no quality guarantees.
Record the installed build/model and reproduce short and multi-section generation using an authorized reference. Keep audio/transcripts local, record only aggregate timing/memory/outcomes in docs. Confirm no termination, intelligible complete speech, cancellation and a usable second run. If this fails, resolve it before changing generation defaults. Keep the current build as the behavioral comparison.
Add a Codable, validated LocalQwenGenerationSettings value in the existing iOS feature structure. Put an immutable settings snapshot in each generation request so changes cannot affect an in-flight run. Separate sampling, segmentation and output-processing fields.
Build 59 behavior: existing profiles decode with current defaults. Generation captures the current editor values, including unsaved edits, for both saved and unsaved voices. Save per-voice overrides explicitly using “Save as voice defaults”; reference selection stays independently remembered. Switching voices loads their saved overrides; editing does not silently persist them. Reset restores built-in defaults in the editor and affects the next run, but is persisted only by Save. Keep model choice separate to avoid implicit memory-heavy loads.
Gate: legacy profile decoding, round-trip persistence, invalid/NaN/out-of-range values, voice switching, reset/save and immutable request tests. Preserve reference file protection and backup exclusions.
Expose language and Variation, followed by collapsible Advanced controls for temperature, top-p, top-k, repetition penalty and audio-token ceiling. Wire validated values through the actual Qwen adapter. Keep existing defaults unchanged. Label token ceilings as generated audio tokens, not input tokens or exact duration. Detect or clearly warn about possible cutoff when the ceiling is reached; inspect SDK termination information before promising detection.
Gate: request-to-runtime mapping tests and device short-generation checks at defaults and selected bounds. Test VoiceOver labels, Dynamic Type, slider values and narrow-device layout. Prevent concurrent generations; clarify whether edits apply to the next run.
Replace plain string chunks with section metadata distinguishing continuous, sentence and paragraph boundaries. Preserve paragraph structure before whitespace normalization. Keep the 420-character maximum until evidence supports raising it. Update preview to show section sizes and explicit pauses.
Continuous sections use the crossfade and no inserted silence. At paragraph boundaries, honor the requested added pause and avoid crossfading across it; use short fades at audio-to-silence edges if needed. A zero pause adds no silence but does not remove pauses the model itself generates.
Gate: punctuation/Unicode/long-word/newline splitting tests, no dropped or duplicated text, exact added-silence sample counts and crossfade tests. Listen to a long sentence, multiple sentences and multiple paragraphs for clicks, static and unnatural gaps.
Use offline pitch-preserving time stretch with AVAudioEngine/AVAudioUnitTimePitch or an equivalent verified AVFoundation path. Apply it to generated sections before final paragraph spacing, so a 500 ms pause remains 500 ms at every speed. Keep 1.0× as a bypass. Preview, playback and export must use the same final file.
Run processing off the main actor, check cancellation, bound buffers, clean partial files on failure and publish only a complete result. Avoid keeping several full-length audio copies in memory. Do not add pitch shifting in this scope.
Gate: duration ratio, sample format, finite/nonempty output, pitch preservation on synthetic tones, cancellation and temporary-file cleanup. Listen at 0.75×, 1× and 1.25×; profile peak memory with model weights still loaded.
Compare current defaults against small, documented candidate settings on fixed authorized examples. Include a short reference, the previously used longer reference, one long sentence and paragraphs. Test 4-bit first, then 6-bit and 8-bit only where device memory allows; record unsupported combinations explicitly. Do not claim universal support based on the 8 GB phone.
Only publish named presets if they provide repeatable useful differences; retain Custom and Reset defaults. A/B comparisons evaluate intelligibility, completeness, voice likeness, naturalness, joins and runtime rather than assuming higher precision or temperature is better.
Gate: focused iPhone tests, actual listening, repeated generation/cancel/retry, export playback, memory/thermal observations, installed build and launch evidence. Increment iOS build for app changes; update the Android build assertion where required and run the full Android CI-equivalent gate before committing/pushing it. Update prototype guide and Obsidian work log/status/map, then verify git diff –check and account for changed files. Commit/push only within authorized release scope.
Primary existing files: LocalQwenContracts.swift, LocalQwenViewModel.swift, LocalQwenRuntime.swift, LocalQwenPrototypeView.swift, LocalVoiceProfileStore.swift, LocalQwenPrototypeTests.swift. A dedicated settings value and output processor may become separate Swift files if needed for clarity. Do not patch cached third-party package checkouts.
Focused command: xcodebuild test for SpeakTrueTests/LocalQwenPrototypeTests using the connected physical iPhone destination and existing package cache. Contract tests do not establish speech quality. Final acceptance also requires the actual generation and listening cases above.
Start with phase 0, then implement phase 1 as the first bounded code change. Complete and verify each phase before expanding into the next; this document is the review point before implementation begins.