SpeakTrue

iOS Local Qwen Prototype Guide

Local TTS workspace simplification (working tree, 2026-09-22)

SpeakTrue 1.4 build 67 makes the primary Local TTS path choose a saved voice, enter text, generate locally, then play or export the captured result. The pinned action offers one direct recovery step for model preparation, missing voice, missing reference transcript, empty text, or an active recording; unavailable engines and unsupported models explain the block. A compatible downloaded model is preferred for first use, while explicit model selection remains under Manage models. Model download and its approximate size remain an explicit choice.

The selected reference transcript can be edited and saved without replacing its audio or identity. Generation settings, model and voice administration, history, and performance details remain available in disclosures. The result retains its captured voice/model and offers Regenerate and Export: Share local WAV, Save local WAV, or Save to Soundboard. Soundboard save explicitly uploads the generated audio and metadata; reference audio and its transcript remain local. Regenerate after an unload or model switch directs preparation of the original model before retrying. The UI and persistence tests do not establish physical-device audio quality, thermals, or memory behavior.

Local TTS lab generate and gated download (2026-09-15)

Generation settings show only the knobs the selected engine reads (PocketTTS temperature; NeuTTS temperature/top-k/ceiling; MOSS none; Qwen and Chatterbox keep their fuller sets plus segmentation). NeuTTS Air is hidden from iOS model cards until generate is unblocked. Qwen/Chatterbox generate uses the SDK streaming path and copies each audio chunk to CPU floats immediately. The non-streaming generate() path concatenated reference plus all new codec frames and was jetsammed on device at about 5.5 GiB (vm-pageshortage). PocketTTS keeps Mimi decoder state across sentences. Loading a second model unloads the resident MLX weights and every ONNX/llama engine first, so Pocket and MOSS sessions cannot stay in memory together. MOSS opens encode, TTS, and codec graphs sequentially, keeps KV cache as native ORT values, and caps clone prompts at the most recent 10 seconds. Gated NeuTTS downloads still require a Hugging Face token in the lab notice after you accept the model terms.

Release preparation and optional account profiles (build 63)

Qwen new/reset settings now use Top-K 50 and repetition penalty 1.05; temperature 0.9 and Top-P 1 remain. Explicit saved overrides stay unchanged. Other model families retain their previous defaults. The app includes its own required-reason privacy manifest. Hosted consent gates outbound requests, including direct ElevenLabs and account-bound STS continuation, while local navigation remains available without hosted permission.

New reference profiles are account-scoped on device. Legacy installation profiles stay unclaimed and visible as such. Sign-out hides account profiles and clears active reference/playback state; it does not upload or erase independent legacy profiles. Generated history remains installation-scoped. Account deletion purges the requesting device’s account-profile directory after successful server deletion.

The optional account library supports explicit upload confirmation, saved revisions, separate copies, checksummed downloads and distinct cloud deletion. It is disabled by default until backend rollout and real two-device validation. See release and cloud profile plan and the backend function README for rollout dependencies and limits. Cloud export normalizes mono references to PCM16 WAV. No model weights or generated history are uploaded by that feature.

OmniVoice is omitted unless the Boolean Info.plist key LocalTTSOmniVoiceBetaEnabled is true. SPEAKTRUE_CLOUD_VOICE_PROFILES_ENABLED also defaults false. Set these intentionally before building a beta artifact; Release alone does not distinguish TestFlight from App Store distribution. Configure SPEAKTRUE_CLOUD_VOICE_STORAGE_ORIGIN to the approved HTTPS Storage origin when it differs from the API host. Do not edit signed artifacts after build.

Qwen stopping investigation (2026-09-09; no runtime change)

The device tokenizer was copied read-only and the latest script encoded locally: 59 text tokens. The pinned SDK imposes min(maxTokens, max(75, textTokens * 6)), so this run permits 354 codec frames despite the app’s 4096 ceiling. At 12.5 Hz, 354 frames produce exactly 28.32 s (679680 samples at 24 kHz), matching the WAV. Because the SDK checks EOS before appending its frame, this is strong evidence of limit exhaustion without EOS, rather than a natural stop followed by padding. The token trace was not retained, so this is reconstructed evidence.

Inspected EOS behavior: the configured codec EOS is excluded from special-token suppression, preserved through sampling filters, temperature-scaled with other logits, and checked before decoding. The ICL prompt includes text EOS and matches the upstream non-streaming prompt layout. No missing EOS check or duplicate final audio append was found. The app currently discards .token events and retains neither EOS status nor termination reason in history.

The run used top-k 0 and repetition penalty 1.1. Upstream Qwen defaults are 50 and 1.05, respectively. This is a controlled-test candidate, not an established cause. Next: retain EOS/limit diagnostics per section, then compare the same script/reference with one sampling parameter changed at a time and repeat runs. Do not shorten the token cap, force EOS, add amplitude trimming, or add transcription-based stopping as a substitute for diagnosing natural termination. No settings, model files, audio, or app code changed in this investigation.

Reference: Qwen generation and ICL implementation.

Qwen 6-bit tail and duplicate-control review (build 62)

Latest retrieved Qwen 6-bit run: 23.2931 s generation, 28.32 s WAV, first internal audio chunk at 1.8229 s; temperature 0.9, top-p 1, top-k 0, repetition penalty 1.1, token ceiling 4096. Local Whisper-base transcription places the final requested words around 17.7 s. The remaining roughly 10.6 s has low-level nonzero audio (one-second RMS measurements mostly -38 to -51 dBFS), not digital silence. ASR does not prove that every trailing sound is non-speech. The pinned SDK emits incremental chunks and only the remainder at completion; no duplicate whole-waveform append was found in the inspected path. The exact cause of late termination remains unresolved. The WAV-based RTF is 0.823 but includes the tail; it must not be presented as a useful-speech performance benchmark.

Variation and Temperature were duplicate bindings to the same sampling value. Build 62 removes the duplicate, retaining one Temperature (variation) slider with an explanation. Sampling values and generation behavior are unchanged. A separate 18.30 s trimmed preview was prepared outside the repository for user comparison; the original device clip is unchanged. No automatic duration or loudness trim is introduced because that could remove quiet intended speech.

Manual memory unload (build 61)

The expanded model controls now offer Unload from memory whenever a model is loaded. The button identifies the model actually in memory, even if another model is selected. It is disabled during generation or model preparation/unload. Unloading releases model weights, reference waveform/tokens and the MLX cache, then refreshes storage indicators. Downloaded files, saved voice profiles, current output and history are retained; generation requires loading a model again. Switching selection alone still does not unload a model; loading another model still releases the old one automatically.

Build 60 device quality follow-up (2026-09-09)

The latest saved OmniVoice run completed with 18.56 s output in 50.2075 s (RTF 2.705): reference preparation 0.0386 s, reference encoding 2.2327 s, diffusion 46.2285 s, decoding 1.6797 s, assembly 0.0257 s. Settings were 16 steps, guidance 2, speed 1; reference tokens were not reused. This differs from the earlier text/baseline, so it is not a controlled speedup measurement.

The user reports inaccurate speech; local cached Whisper-base transcription of the retrieved WAV also shows omissions and rearranged words. ASR is imperfect and is supporting evidence, not a listening verdict. This is a completed but unacceptable generation, not a validated quality fix. The selected saved reference is 34.3846 s and its transcript broadly agrees with local ASR.

OmniVoice upstream guidance recommends 3–10 s references; longer references can slow inference and degrade quality. The generic in-app 10–30 s copy needs family-specific correction. Long reference conditioning is a plausible contributor, not a proven cause; real-codec window equivalence and reduced-step fidelity remain unverified.

Next comparison: a separate 7.10 s excerpt and corresponding first-sentence transcript were prepared locally for import as an alternative reference. The original recording/profile is preserved. Retry the same target text at the same 16 steps, guidance and speed; check fidelity before further latency tuning. No reference audio, transcript or target text is stored in repository/vault docs.

Status

At the build 63 checkpoint, the Local TTS Lab was an experimental, installable iOS prototype with Qwen3-TTS 4/6/8-bit, Chatterbox multilingual, and OmniVoice. The app compiles as a universal iPhone/iPad target; per-model speech quality, thermal behavior, and memory suitability still require device validation.

Build 60 (2026-09-09) addresses a reported late OmniVoice crash with bounded codec decoding: 100 token frames per window plus 32 real context frames on each side, trimmed to the centre without extra fades, silence or crossfades. The pinned convolutional codec geometry is validated before loading. Conditional and unconditional diffusion passes are evaluated separately to avoid overlapping activation graphs. Decoder graphs are evaluated and released per window; unexpected lengths and non-finite samples become errors. Release logs record phase/memory counters without text or audio. Phone reports now confirm memory-pressure termination: SpeakTrue was the killed process with reason vm-pageshortage at 21:27:03 and 21:29:31 on 2026-09-09. The latter recorded 358,358 resident pages at 16,384 bytes/page (5.47 GiB). Reports do not identify the exact inference stage. Initial CoreDevice 12040/12010 mounting failures later cleared; Release 1.3 (60) was built, installed and launched, with version verified on the phone. The mitigations still require a repeat generation and listening test; no successful post-fix inference or memory reduction is claimed.

The agreement service now distinguishes a failed status lookup from an actual negative consent result. Same-account verified acceptance survives a transient refresh failure. Unknown status shows a retry screen without reopening the agreement notice; account changes/reset invalidate pending requests. Local OmniVoice still makes no ElevenLabs generation request. The existing app-wide agreement requirement is otherwise preserved.

Verification: 52 focused Simulator tests passed, including consent failure/account isolation and CPU-only window boundary/length checks. Direct MLX evaluation in the simulator aborted during Metal initialization, so these tests do not claim codec or physical-device inference validation. Android/iOS parity passed.

Build 59 puts text entry first, beneath a compact model/voice header. Expand the model row for coloured model buttons, downloads and device details; Manage voices contains reference recording/import, transcripts and profile editing. The voice picker and reference preview stay in the compact header. Generate/Cancel remains visible above the tab bar, with the blocking reason, elapsed time and actual generation stage. OmniVoice exposes reference encoding and audio decoding; other families retain the coarser stages their SDK exposes. Remaining-time estimates appear only after three runs with matching saved profile/reference, model, language and settings, and similar input lengths; estimates can be wrong when thermal state or cache warmth changes.

Generation captures the visible settings immediately. Save as voice defaults persists them for future sessions; editing and reset alone do not persist. At the build 59 checkpoint, results offered playback, regenerate, share and Save WAV, with timings under Performance details. Section colours remain editable; their explanation is collapsed. The latest 20 successful clips and their text/settings are retained in Application Support/LocalTTSHistory, protected and excluded from backup. They survive sign-out and can be played, reused or individually deleted without regeneration. Save WAV exports a separate copy; deleting history does not delete exports or reference profiles. Reusing an entry whose reference was deleted requires choosing another voice. Only the current unsaved reference can be regenerated directly; history does not retain reference audio. That checkpoint did not yet include cloud Soundboard save; the current working-tree behavior is described above.

Build 58 replaces the local model dropdown and duplicate inventory list with coloured selection buttons. Each shows its model name, approximate download size and current storage/preparation state; a loaded model shows both Downloaded and Loaded in memory. A checkmark and border identify selection independently of colour. Selecting a button does not download or load automatically; the existing action below the grid does that. The grid adapts to width and larger accessibility text, and selection remains disabled during preparation/generation.

Build 57 moves the lab into the TTS tab under a Standard / Local TTS selector. STT’s Copy to TTS action replaces both drafts, then opens TTS in its selected mode. Subsequent edits remain independent. The tab container owns the local view model, so switching modes preserves its draft, selected profile, loaded model and active download/generation. Leaving the local view stops playback and reference recording; leaving the main tab container cancels operations and cleans temporary reference audio. STS is unchanged.

Current controls and review corrections

Build 56 implements in-memory OmniVoice reference-token reuse. The reference is encoded once before the section loop, then reused for all sections and later runs with the same audio while the model remains loaded. The single-entry cache uses SHA-256 of the reference file plus sample rate, so fresh temporary copies reuse it and changed audio misses it. Model replacement/unload and profile deletion clear it; failed or cancelled encoding never enters the cache. Tokens are not persisted to disk.

The timing panel separates reference preparation, reference encoding (including a reuse label), speech generation, audio decoding and WAV assembly, and labels Debug versus Release builds. These use monotonic clocks; MLX evaluation is completed at encoding/diffusion/decoding boundaries before the phase is timed. Total/first-audio timing now starts before reference-file preparation. Phase totals omit small orchestration/cache-cleanup overhead and need not equal the total exactly. Other model families retain their existing opaque runtime path.

Both default and custom OmniVoice settings now show diffusion-step progress. Playback still waits for the complete joined WAV. Speed improvement and voice quality with the new adapter require physical-device validation; the earlier 141.93-second / 37.48-second two-section result at 16 steps is the user baseline.

Maintenance: LocalOmniVoiceModel.swift and LocalOmniVoiceModelConfig.swift adapt the orchestration/configuration from mlx-audio-swift commit d9e6e7c59cc11dd58e6ee314f3d46144469f8868, retaining its MIT notice. They reuse the SDK’s Qwen3 backbone, OmniVoice audio tokenizer and diffusion-parameter type. The model loads from the app-verified directory; the codec still uses the repaired default cache. The local adapter adds prepared-token input, timings and strict missing-model-weight rejection without changing sampling math or defaults. Review this adapter against upstream whenever updating the pinned SDK; do not patch Xcode’s external package checkout.

Build 55 shows generation sections as alternating text colours directly in the editable generation box. The separate section-boundary disclosure is removed. Colour ranges follow the same chunker as generation while preserving original line breaks, spacing and emoji. The section count and character limit remain below the editor. This changes presentation only, not generation segmentation.

After build 54 the user reported good OmniVoice quality on iPhone 16 Pro Max, but slow generation even for roughly 90 characters. That report prompted the build-56 reference reuse and timing work above; device benchmarking remains open.

Latency investigation (2026-09-08): the user reports excessive latency even at 16 diffusion steps. Previous device deliveries used Debug (-Onone). Compare an optimized Release build using the same model, saved reference, text and settings before attributing latency to hardware alone. Record total generation time and output duration; a build succeeding is not measured speed evidence. The pinned OmniVoice API re-encodes reference audio for each section and exposes no pre-encoded-reference parameter. Its waveform arrives only after a complete section. This motivated the local adapter described above. The earlier Release 1.3 (55) build succeeded with -O verified for SpeakTrue, MLX and MLXAudioTTS. Installation is pending: the paired iPhone was unavailable when delivery was attempted. No Release-versus-Debug latency result is available yet.

Settings are saved per voice and model family. Switching families reloads that family’s saved editor values. Runs capture the current editor values, including unsaved edits. Use Save as voice defaults to remember them across sessions.

Build 53 corrects two regressions in the owned download transport: cancellation now stops the actual URLSession task and resumes the waiting preparation exactly once, and an independent monotonic watchdog detects 25 seconds without progress even if no byte callback arrives. Stalled transfers retry up to three attempts; user cancellation stops retries and backoff. Progress uses the pinned manifest’s size if the server omits Content-Length. The redundant debug HEAD request is removed, so preparation does not wait for an extra network probe.

Build 54 repairs OmniVoice’s cache handoff. The pinned SDK ignores the supplied cache in its nested resolvers, so the runtime now replaces entries in HubCache.default from the verified model folder before loading. Existing entries are replaced even when present; deleting OmniVoice also removes that loader cache. On the affected phone the loader cache contained a 281,103,571-byte model while the app-managed manifest expects 2,450,344,102 bytes. The latest output was a valid 6.4-second WAV with nearly constant energy, consistent with the reported noise. The cache defect is confirmed; speech quality after repair still needs listening verification. Reload OmniVoice after installing the update; a complete verified download can be reused.

The sections below retain the earlier Qwen-specific troubleshooting history.

The lab is intentionally separate from the hosted ElevenLabs voice-cloning path. It links the MLX runtime directly into the iOS app; Ollama, LM Studio, and a Mac-hosted model server are not required after installation.

Open and use the prototype

  1. Open TTS and choose Local TTS.
  2. Select the Qwen3-TTS 0.6B · 4-bit button, then choose Download & load while connected to Wi-Fi and power.
  3. Expand Manage voices to record a clean voice reference or import an audio file.
  4. Enter the exact words spoken in the reference. The transcript is part of the voice conditioning and should match the audio closely.
  5. Name and save the profile. The reference audio and transcript persist in app storage for later sessions.
  6. Enter generation text, choose a language code, and generate. Longer text is split into bounded chunks and reassembled into one WAV.
  7. Play the result in the app or use Share WAV to export it.

The timing panel reports time to first audio, total generation time, output duration, and real-time factor. A real-time factor below 1.0 means generation completed faster than the resulting audio duration.

Download troubleshooting

Build 31 corrects the Hugging Face repository IDs to end in 4bit, 6bit, and 8bit. Earlier builds requested IDs with an extra hyphen (such as 4-bit), which caused all three downloads to fail immediately with a generic HuggingFace.HTTPClientError message. Install build 31 or later and retry Download & load. The visible model labels still use “4-bit”, “6-bit”, and “8-bit”. Successful device download and inference still require device validation.

Reference-generation troubleshooting

Build 33 also repairs incomplete cached downloads before loading: both the root model checkpoint and speech_tokenizer/model.safetensors must be present and nonempty, along with the tokenizer/configuration files. An older cache may have root weights but no speech-codec weights; the SDK can still report it loaded and produce near-silent noise. Choose Download & load to fetch missing files. Saved voice profiles remain intact.

The runtime limits MLX’s unused allocation cache to 2 MB and clears it after model loading and before reference processing. Build 31 was observed exiting with signal 9 during reference conditioning; build 32 completed the same 34.4-second reference attempt but produced near-silent output, and device inspection found missing speech-codec weights. Memory pressure is suspected for the original termination; a matching jetsam report was not available. The missing codec was verified repaired on the iPhone in build 33. A second consecutive load then exited with signal 9 before generation. Build 34 prevents loading an already loaded model and releases the previous model before switching. The user confirmed usable 4-bit speech after repair in build 34, with quality still below their target.

Build 34 separates byte-based download progress from indeterminate tensor loading. Cancel immediately enters a stopping state, cancels the preparation task, and discards a canceled load rather than marking it ready. A synchronous MLX operation already executing must finish before its resources can be released. Other profile/text controls remain available while preparing; model switching, deletion, and generation are blocked until preparation has stopped.

A subsequent 6-bit download stalled after small files arrived, and cancellation remained pending in the SDK snapshot transport. Build 35 replaces snapshot/Xet transfer with sequential HTTP downloads for required files. Progress names the current file and reports its bytes. Cancel invalidates the active URLSession; requests time out after 30 seconds without activity. Completed cached files are kept, while an interrupted file may need to restart. The user confirmed cancellation works in build 35. Six-bit completion remains unverified. Build 36 keeps the model card full width in every preparation state and keeps a linear progress bar visible while connecting. It shows percentage and bytes once the server reports a file size, or connecting/received-byte text when the size is still unknown.

Build 37 pins each model to a published Hugging Face revision and validates required file sizes plus SHA-256 for both model checkpoints. This replaces the nonempty-file readiness check after device inspection found a six-bit root checkpoint of 995,086,264 bytes instead of the published 1,164,476,202 bytes. A canceled/truncated or same-size corrupt checkpoint cannot be marked ready. Invalid files are redownloaded with the SDK cache disabled into staging, verified, and only then moved into place. Valid existing files are preserved. Hashing uses bounded 1 MB reads and checks cancellation. The UI displays “Verifying model files…” during this step. Repaired six-bit speech remains under device validation.

Debug builds report reference duration and aggregate MLX memory use without logging reference audio, transcript text, or generation text.

Model inventory, default voice, and section boundaries

Build 38 lists all three model variants as Loaded, Downloaded, Needs download/repair, or Not downloaded. Inventory uses expected file sizes; checkpoint hashes are still verified before loading. Select a saved profile and choose Use as default voice to select it automatically when reopening the lab. Clearing or deleting the default removes that preference without changing other profiles. The preference is local, protected, and excluded from backup.

Generation shows the section count, a 420-character per-section limit, and an expandable preview of the exact text sections. Splits preserve punctuation and prefer sentence endings, then word boundaries; exceptionally long words are split to preserve the bound. The limit is characters, not text tokens. The SDK is asked for at most 4,096 generated audio tokens per section and also caps output based on input length, so duration varies.

The old fixed 140 ms inserted gap is removed. Independent sections overlap with a 5 ms linear crossfade to reduce boundary clicks. This does not promise removal of artifacts or pauses generated inside a section by the model. Streaming frames within a section remain in their original order.

There is no enforced reference-duration maximum. The user-confirmed 4-bit run used a 34.4-second reference; longer references require more memory/processing. A 10–30-second clean reference is a practical comparison range, not a hard limit or quality guarantee. Build 39 supports multiple selectable recordings per profile. Select a saved voice, tap Add recording to this profile, record or import its audio, enter its exact transcript, and save. The recording picker restores that recording’s transcript; the preferred recording persists across launches. Existing single-recording profiles remain compatible. Deleting a profile removes all its recordings. The current Swift generation path uses only the selected reference/transcript; adding alternatives does not automatically improve quality. Upstream lists represent batched prompts, not automatic multi-clip fusion.

Model options

Option Approximate download Prototype guidance
Qwen 0.6B Base 4-bit 1.71 GB Recommended starting point for both target devices. Lowest memory pressure and best chance of responsive generation.
Qwen 0.6B Base 6-bit 1.85 GB Quality/latency comparison option. Expect higher memory and thermal pressure.
Qwen 0.6B Base 8-bit 1.99 GB Highest-fidelity quantized pilot. Treat as experimental, particularly on iPhone.

The app evaluates reported physical memory, low-power mode, and thermal state before generation. Those messages are operational guidance, not certification that a model will fit every workload.

Persistence and storage

Install later from the Mac Studio

  1. Pull or open this repository in Xcode on the Mac Studio.
  2. Open ios/SpeakTrue.xcodeproj and allow Xcode to resolve the pinned Swift packages. Xcode may ask you to validate the MLX package build plugin.
  3. Select the SpeakTrue target and your Apple development team if signing is not already configured.
  4. Connect and trust the iPhone or iPad, select it as the run destination, and choose Run.
  5. On first use, approve microphone access if recording a reference. Imported reference audio does not require microphone access.

For command-line verification, package plugin validation can be skipped explicitly:

xcodebuild \
  -skipPackagePluginValidation \
  -project ios/SpeakTrue.xcodeproj \
  -scheme SpeakTrue \
  -destination 'generic/platform=iOS' \
  CODE_SIGNING_ALLOWED=NO \
  build

Current verification boundary

The prototype has been package-resolved, compiled for a generic physical iOS destination, compiled for iPhone and iPad simulator destinations, and covered by focused unit tests for model selection, device assessment, long-text chunking, timing metrics, and profile persistence. Simulator inference is intentionally disabled because it is not representative of physical-device MLX memory behavior.

Before treating the pilot as device-ready, validate on both target devices:

The MLX Audio Swift package is pinned to revision d9e6e7c59cc11dd58e6ee314f3d46144469f8868 so later dependency changes cannot silently alter this pilot.

Build 40 memory-limit correction

The build 39 Generate failure on the physical iPhone at 11:16 on 2026-09-05 was an iOS jetsam termination with reason per-process-limit. SpeakTrue was frontmost at 216,125 resident 16 KiB pages (approximately 3.30 GiB). Build 40 requests com.apple.developer.kernel.increased-memory-limit, which was absent from earlier builds. Apple permits a higher per-app memory limit on supported devices; this is not a guarantee that every model/reference fits. The same generation must be retried before claiming this resolves inference. Source: https://developer.apple.com/documentation/bundleresources/entitlements/com.apple.developer.kernel.increased-memory-limit