SpeakTrue

2026-09-16 — Auto-unload when loading a second Local TTS model

Loading a second Local TTS model unloads the first on iOS and Android. iOS drops MLX weights and every ONNX/llama engine before the replacement download/load. Android unloads every registered engine and clears the previous Ready state. Reloading the same model is unchanged. Test CPU spikes during xcodebuild are the Swift build plus File Provider/Spotlight on the iCloud-synced Documents tree; lsd is Launch Services, idle when not installing Simulator apps.

2026-09-09 — Successful Qwen 6-bit sampling runs, user-confirmed quality

Latest device history confirms two runs with the same 223-character script and reference, temperature 0.9, top-k 50, repetition penalty 1.05452 and ceiling 4096. Top-p 1 produced 16 s audio in 14.7589 s (RTF 0.922, first internal audio 2.2071 s). Top-p 0.90248 produced 14 s audio in 12.6342 s (RTF 0.902, first internal audio 1.2889 s). Device tokenizer encodes this script to 47 tokens, giving a hidden SDK cap of 22.56 s; both outputs ended below that cap, consistent with natural EOS. Explicit token traces are still unavailable. User reports the unwanted tail is fixed and quality/speed are decent; no independent listening claim. Different script from the earlier failed 269-character run prevents isolating causality to sampling. Both top-p values worked; do not credit top-p reduction alone.

Sampling settings are now user-validated candidates; no app defaults or runtime code changed. Broader long-text/repeated-run validation remains pending. Device history retrieved read-only and token count checked locally. No audio/transcript content copied into documentation. Repository dirty state preserved atop 13dfb226, build remains 62, no new repository files, no commit/push/build. Map refreshed and checked at closeout.

2026-09-09 — Qwen stopping investigation; build remains 62

Read-only device tokenizer retrieval and local encoding verified 59 text tokens for the latest run. Pinned SDK hidden limit is min(user ceiling, max(75, text tokens times 6)): 354 codec frames / 12.5 Hz = 28.32 seconds, exactly matching the WAV. EOS is checked before appending audio frames, so evidence strongly indicates limit exhaustion without natural EOS; no token trace was retained. The reason EOS was not selected remains unresolved. Inspected suppression, sampling, ICL text-EOS prompt and stream finalization; no missing EOS check or duplicate final append found. Runtime currently ignores token events. Current sampling differs from upstream defaults (top-k 0 versus 50, repetition penalty 1.1 versus 1.05); causality unproven.

Natural stopping investigation is active. Next proposed slice: retain per-section EOS/limit diagnostics, then controlled repeated comparisons with the same script/reference and one sampling setting changed at a time. User excludes blunt duration trimming and transcription-based stopping. No runtime settings or generation behavior changed, and no new build/device generation was performed. Additional device config retrieval timed out in CoreDeviceService after tokenizer retrieval succeeded; device config was not reverified. No private transcript/audio included here.

Changed this slice: docs/guides/ios-local-qwen-prototype.md only; prior 12-file dirty state preserved atop 13dfb226. No new uncommitted files. Verified tokenizer count, WAV sample duration, pinned source and upstream source; git diff –check passed. No commit/push. External map regenerated and checked at closeout.

2026-09-09 — Qwen tail diagnosis and single temperature control, build 62

Retrieved latest Qwen 6-bit history/WAV from the phone. Total generation 23.2931 s; output 28.32 s; first internal chunk 1.8229 s. Settings: temperature 0.9, top-p 1, top-k 0, repetition penalty 1.1, ceiling 4096. Local cached Whisper-base places final requested words around 17.7 s; trailing approximately 10.6 s contains low-level nonzero sound, mostly -38 to -51 dBFS in one-second RMS windows. No external ASR received audio. ASR is supporting evidence, not proof of silence/non-speech. Inspected pinned Qwen streaming path emits each chunk once plus final remainder; no full-output duplication found. Late end-of-sequence behavior remains unresolved. File RTF 0.823 includes unwanted tail and is not a useful-speech speed claim.

Confirmed Variation and Temperature were duplicate bindings. Removed the redundant Variation control; one Temperature (variation) slider remains with clear help. Values/ranges and generation behavior unchanged. Prepared separate 18.30 s trimmed preview under /private/tmp/speaktrue-qwen6-trimmed-preview.wav; original phone clip unchanged. No automatic trimming introduced; avoid cutting quiet intended speech. No audio or transcript content stored in repo/vault.

Release 1.3 (62) built, installed over the existing iPhone app, launched and device version verified. Android/iOS parity passed 12 surfaces/24 gates; git diff –check passed. No new tests needed for the duplicate-control removal; prior build-61 suite had 52 passing focused tests. No physical tap-through or new generation result claimed.

This slice changes ios/SpeakTrue/LocalQwenPrototypeView.swift, ios/SpeakTrue.xcodeproj/project.pbxproj, android/app/src/test/java/com/speaktrue/contracts/AndroidReleaseReadinessContractTest.kt (iOS build assertion), docs/guides/ios-local-qwen-prototype.md, docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md and docs/product/LOCAL_QWEN_GENERATION_SETTINGS_PLAN.md. Prior changes preserved; 12 modified tracked files atop 13dfb226, no new uncommitted files. No commit/push, GitHub Actions or TestFlight upload. Full local Android gate remains required before commit/push. External map refreshed/checked at closeout.

2026-09-09 — Manual model unload, build 61

Added Unload from memory in the expanded Local TTS model controls. It names and unloads the model actually in memory, even when a different model is selected. Busy state prevents generation, load/delete, selection and repeated unload during the operation. Runtime unload releases weights, cached reference waveform/tokens and MLX cache; storage indicators refresh afterward. Downloaded model files, voice profiles, generated output and history remain. Generation requires reloading a model. Loading a replacement still unloads automatically; selection alone does not.

Verification: 52 focused LocalQwenPrototypeTests passed; signed Release 1.3 (61) built and installed over the existing app on Yoseif’s iPhone. Android/iOS parity passed 12 surfaces/24 gates; git diff –check passed. Runtime/button wiring reviewed, but no instrumented on-device RAM measurement or tap-through result claimed. OmniVoice accuracy remains unresolved from the earlier report; this task does not change generation math.

Changed this slice: ios/SpeakTrue/LocalQwenViewModel.swift, LocalQwenPrototypeView.swift, ios/SpeakTrue.xcodeproj/project.pbxproj; android/app/src/test/java/com/speaktrue/contracts/AndroidReleaseReadinessContractTest.kt (iOS build assertion only); docs/guides/ios-local-qwen-prototype.md and docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md. Prior build-60 changes preserved. Total working tree: 11 modified tracked files atop 13dfb226; no new uncommitted files. No commit, push, GitHub Actions or TestFlight upload. Full local Android gate is required before committing/pushing the accumulated parity assertion. External map refreshed/checked at closeout.

2026-09-09 — Build 60 completes but speech fidelity fails

Build 60 completed the latest on-device OmniVoice run, but the user reports inaccurate speech; local cached Whisper-base ASR supports omissions/reordered words. Output 18.56 s, total 50.2075 s, RTF 2.705; diffusion 46.2285 s (92.1%), decoding 1.6797 s, reference encoding 2.2327 s (cache miss), preparation 0.0386 s and assembly 0.0257 s. Settings: 16 steps, guidance 2, speed 1. This is not the same text as the earlier baseline, so no controlled speedup is claimed.

The selected reference is 34.3846 s, with broadly matching transcript according to local ASR. Upstream OmniVoice recommends 3–10 s and warns that longer references can slow inference and degrade quality (https://github.com/k2-fsa/OmniVoice#voice-cloning). Our generic 10–30 s app guidance is unsuitable for this family and remains to be corrected. Long conditioning is a hypothesis, not confirmed causality; real-codec window equivalence also remains unverified. Quality status remains failing pending repeat comparison.

Prepared a separate 7.10 s first-sentence reference and matching transcript under /private/tmp/speaktrue-omnivoice-reference-test for user import as an alternative. Original phone profile/recording unchanged. Excerpt duration and local ASR checked. User should compare the same target text with unchanged 16-step/guidance/speed settings before further tuning. No external transcription service or external model received the audio; no audio/transcript content copied to the vault or repository.

This turn changed only docs/guides/ios-local-qwen-prototype.md and docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md; prior nine-file build-60 dirty state remains uncommitted atop 13dfb226. No new uncommitted repository files. No new app code, build, install, commit/push, CI run or TestFlight action. git diff –check passed; external map refreshed/checked at closeout.

Build 60 is implemented, signed Release-built, installed and launched on Yoseif’s iPhone; device app inventory confirms 1.3 (60). Device jetsam reports at 21:27:03 and 21:29:31 on 2026-09-09 explicitly identify SpeakTrue as killed for vm-pageshortage. The latest has 358358 resident pages at 16384 bytes/page, approximately 5.47 GiB. The exact generation stage is unproven. Initial CoreDevice 12040/12010 developer-image errors later cleared. Reports remain outside the repository/vault; no raw logs or user audio/text copied here.

Memory mitigations: evaluate conditional and unconditional diffusion passes separately, and decode 100-frame windows with 32 context frames on each side using the unchanged pinned codec. Trim shared context without additional fades, silence or crossfades; materialize CPU samples and release each decoder graph before the next window. Validate codec geometry, decoded length and finite samples; emit content-free Release phase/memory logs. The default sampling calculations and final normalization remain unchanged. Physical repeat-generation, listening continuity and measured memory/latency improvement remain pending.

Consent bug: failed remote lookups no longer mean missing consent. Preserve same-account verified acceptance on transient refresh failure, show a retry screen for unknown status, and reject stale responses after account reset/change. Only an actual negative result opens the notice. This fixes a code path that could explain the reported prompt, without claiming that the exact network failure was captured. The app-wide consent requirement remains; OmniVoice does not invoke ElevenLabs generation.

Verification: 52 focused Simulator tests passed (consent failures/account isolation, CPU-only decoder-window boundaries/length/errors and existing local TTS tests). Initial MLX-in-Simulator tests aborted in Metal initialization; the final tests exercise window orchestration without claiming real codec inference. Final Release build passed, Android/iOS parity passed 12 surfaces/24 gates, and git diff –check passed. No GitHub Actions triggered; no commit or push. Full Android local gate remains required before a future commit/push of the build assertion.

Working tree atop 13dfb226: nine modified tracked files: ios/SpeakTrue/AIDataSharingAgreementService.swift, AIDataSharingConsentGateView.swift, LocalOmniVoiceModel.swift, LocalQwenRuntime.swift; ios/SpeakTrueTests/LocalQwenPrototypeTests.swift; ios/SpeakTrue.xcodeproj/project.pbxproj; android/app/src/test/java/com/speaktrue/contracts/AndroidReleaseReadinessContractTest.kt (iOS build assertion only); docs/guides/ios-local-qwen-prototype.md; docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md. No new uncommitted files. No archive/TestFlight upload.

2026-09-09 — Build 59 changes committed and pushed

2026-09-09: All 18 pending files committed and pushed to origin/main as 13dfb226ea6631619acb44f144c0818d67fbed41 (iOS 1.3 build 59). Includes Local TTS reliability, OmniVoice orchestration/configuration, TTS/STT integration, model controls, private history, tests and product documentation. Android source change is solely the release test assertion for iOS build 59. Commit includes [skip ci] as requested; no workflow configuration was changed. The full local Android CI-equivalent gate passed (unit suite, lint, release bundle, parity/readiness and backend contracts); external live checks remain unverified. Previously verified iOS Release build/install and focused tests remain the iOS evidence. Physical audio quality and latency testing remain pending. Working tree is clean; all three formerly untracked Swift files are now committed.

Earlier entries below are historical snapshots. The pending local commit gate is now satisfied; no archive or TestFlight upload was performed.

Current device delivery

2026-09-08: Optimized Release 1.3 (59) rebuilt successfully, installed over the existing app on Yoseif’s iPhone 16 Pro Max, and launched successfully. Device app inventory confirms version 1.3 / build 59. No uninstall was performed. Physical generation, listening quality, latency and full UI smoke testing remain pending. No archive or TestFlight upload. Repository source unchanged by this delivery; existing changes remain uncommitted and unpushed atop d182a818.

Current generation-first Local TTS workspace

2026-09-08, build 59: all seven approved UI recommendations are implemented locally. A compact model/voice header leads into editable generation text; coloured models and device details expand from the model row, and Manage voices holds reference creation/editing. The voice menu and preview remain accessible in the header. A persistent Generate/Cancel bar shows blocking reasons, actual stage, elapsed time and optional timing estimates. Estimates require three comparable runs with the same saved profile/reference, model, language and settings and similar text lengths; thermal/cache effects remain unmeasured. OmniVoice emits a decoding event; other SDK families retain coarser progress. Inline section colours remain, with collapsed explanatory help.

Generation snapshots current editor settings immediately; Save as voice defaults controls persistence only. Results offer playback, regenerate, Share and Save WAV, with detailed metrics in Performance details. A device-only actor store retains the latest 20 generated WAVs and text/settings metadata in Application Support/LocalTTSHistory with file protection and backup exclusion. History survives sign-out and supports replay, reuse and individual deletion. It stores no reference recordings; deleted or unavailable historical references require selecting a voice. WAV export makes an independent copy. No backend or cloud Soundboard changes.

Verification: 46 focused Simulator tests passed, zero failed; a subsequent targeted rendering test passed after the final voice-menu layout correction. Normal and accessibility-size standalone Local TTS renderings were inspected; an overlapping native voice picker and an oversized sticky explanation were corrected. Optimized generic-iOS Release 1.3 (59) built successfully. Android/iOS parity passed 12 surfaces/24 gates; git diff –check passed. This is compile, contract and simulator-render evidence, not full-device UI/inference or measured latency validation. Installed and launched on the phone after verification; generation testing remains pending. No archive or TestFlight upload. Uncommitted atop d182a818; full Android CI-equivalent gate remains required before commit/push.

Current model selection buttons

2026-09-08, build 58: implemented coloured buttons for all five Local TTS models, replacing the dropdown and duplicate inventory list. Buttons show model name, approximate download size, storage/preparation status and an explicit Loaded in memory indicator; loaded models also show Downloaded. A selection checkmark and border provide non-colour cues and VoiceOver receives the selected trait. The grid adapts to available width and becomes a flexible single column for accessibility text sizes. Selection remains disabled during preparation/generation; download/load/cancel remain separate actions. Model inventory now refreshes after preparation failures as well as success/cancellation/deletion.

Verification: optimized generic-iOS Release 1.3 (58) built successfully; Android/iOS parity passed 12 surfaces/24 gates and git diff –check passed. Source/layout review only; no rendered-device or inference claim. No new tests were added for this small UI change; the preceding build-57 focused suite had 41 passing tests. Device layout/status smoke testing remains pending. Uncommitted atop d182a818; no install, archive or TestFlight upload. Full Android gate remains required before commit/push.

Current TTS tab integration

2026-09-08, build 57: implemented locally. The TTS screen now offers Standard / Local TTS modes, replacing the Voice Cloning lab entry. STT Copy to TTS fills both independent drafts through the main tab container even when a mode is hidden; later edits remain independent. STS is unchanged. The main tab owns the local model state, preserving the draft, loaded model, profiles and active operations across mode switches. Leaving the local view pauses playback and cancels active reference recording; leaving the main tab container performs full temporary-reference cleanup.

Verification: 41 focused Simulator tests passed with zero failures, including shared transcript replacement, repeat copies, Unicode/newlines and independent later edits. Optimized Release 1.3 (57) built successfully. Android/iOS parity passed 12 surfaces/24 gates; git diff –check passed. Physical-device navigation, generation quality and latency remain pending. No installation, archive or TestFlight upload. Work remains uncommitted atop d182a818; the full Android gate is still required before commit/push. Earlier build-56 cache/timing work remains included.

Current OmniVoice reference reuse

2026-09-08, build 56: code-only OmniVoice latency work is implemented locally. A single in-memory reference-token cache encodes once before the section loop and reuses unchanged audio across later runs. Keys use reference-file SHA-256 plus sample rate, not temporary paths; model replacement/unload and profile deletion clear it, and failed/cancelled encoding is not cached. The app-owned LocalOmniVoiceModel/Config adapter reuses the SDK backbone, codec and parameter type while exposing prepared-token generation and phase timings. Default and custom settings now show diffusion-step progress. The timing panel reports reference preparation/encoding, diffusion, decoding and WAV assembly, cache reuse and build mode. Physical-device speed and quality remain unverified while the user is away.

Validation: 40 focused Simulator tests passed, zero failed/skipped; parity verifier passed 12 surfaces/24 gates. The diffusion loop and 11 core helper bodies match pinned upstream d9e6e7c after excluding timing instrumentation and local type naming. This is source/regression evidence, not a measured latency improvement. Code is uncommitted atop d182a818. Earlier proposed-cache status below is superseded; presets, lower-precision evaluation and early section playback remain pending.

Active OmniVoice latency investigation

Measured user baseline (2026-09-08): OmniVoice at 16 diffusion steps, two generation sections, first-audio metric 52.52 s, total generation 141.93 s, output 37.48 s, real-time factor 3.79. Local ffprobe independently confirms the attached WAV is mono float32 PCM at 24 kHz and 37.475 s. No listening-quality assessment is claimed from that metadata check. The first-audio metric marks an internally completed section; playback currently waits for the assembled output. Stage-specific timing and the benefit of reference-token reuse remain unmeasured. Optimized Release 1.3 (55) remains ready but uninstalled; the paired iPhone is still unavailable. No reference audio or transcript contents copied into these notes.

Delivery status: optimized Release iOS 1.3 (55) built successfully; -O verified for SpeakTrue, MLX and MLXAudioTTS. Installation could not proceed because Yoseif’s paired iPhone 16 Pro Max is unavailable. Reconnect the phone to install and measure the same generation. No speed gain is yet measured.

2026-09-08 latency investigation: user still finds OmniVoice too slow at 16 steps. Prior device installs were Debug builds with Swift -Onone. Preparing an optimized Release build of the same iOS 1.3 (55) source as the next performance baseline. Source inspection confirms reference encoding repeats per section and the public OmniVoice API does not accept encoded reference tokens; waveform delivery waits for a whole section. No latency improvement or new reference cache is claimed until measured on the phone.

Profiling is active; encoded-reference caching remains proposed. The earlier inline editor and cache repair remain implemented and uncommitted. Earlier snapshots follow.

Current inline section colours

Build 55 (2026-09-07): generation sections now use alternating text colours inside the editable generation box. The separate section-boundary preview is removed; the section count and character limit remain. Original whitespace and emoji are preserved, and the colours use the actual generation chunker. All 36 focused Simulator tests passed, including Unicode/whitespace range coverage; Android/iOS parity checks passed. User reports very good OmniVoice quality after build 54, with slow generation on iPhone 16 Pro Max. Profiling and encoded-reference reuse remain proposed, not implemented.

Implementation remains uncommitted atop d182a818 with prior fixes preserved. This summary supersedes the earlier status snapshots below. Full Android gate remains required before commit/push.

Current OmniVoice repair

Build 54 OmniVoice repair (2026-09-07): device inspection confirmed a separate SDK-default model cache containing only 281,103,571 bytes versus the pinned 2,450,344,102-byte model present in the app-managed folder. Nested OmniVoice resolvers ignore the supplied cache; the old seed helper targeted its own source directory. The runtime now atomically publishes verified files into HubCache.default before loading, replaces existing entries and removes that cache on deletion. A regression test covers truncated and same-size corrupt cache files, repeated hard links and source preservation. All 35 focused Simulator tests passed. This confirms the cache repair, not audible speech quality; reload and generation/listening validation remain required.

State: implemented locally atop d182a818, uncommitted. Earlier build-53 review fixes remain included. The full Android gate is still required before committing/pushing its build assertion. This summary supersedes earlier current-state snapshots below.

Current local synchronization

Review 2026-09-07: origin/main remains d182a818a720dae19c5a311bdb78f216751ce63e. Local iOS 1.3 (53) corrections are implemented and uncommitted: cancellation reaches the owned transfer, an independent 25-second watchdog detects silent stalls, initial progress and known-size fallback remain visible, family switches reload saved editors, Chatterbox exposes only its supported temperature/top-p controls, and custom OmniVoice first-audio timing records the first section. All 34 focused Simulator tests and the signed generic iOS device build passed; the Android/iOS parity verifier passed. Physical tests were blocked by the locked iPhone; no new installation or listening validation is claimed. The full Android gate remains required before committing/pushing the updated build assertion. Chatterbox/OmniVoice listening, paragraph/speed work and presets remain pending; the prototype guide was refreshed.

Older status statements below are retained as historical context; this summary and the latest dated work-log entries supersede them.

Status: locally synchronized as of 2026-09-07.

Repo: /Users/haddadios/Documents/Programming/SpeakTrue Base commit: d182a818a720dae19c5a311bdb78f216751ce63e Snapshot: local source is iOS 1.3 (53), with nine modified tracked files atop origin/main. Review fixes passed 34 Simulator tests and a signed device build; no new phone installation or generation listening check.

SpeakTrue is a brownfield, multi-surface speech platform for text-to-speech, speech-to-text, speech-to-speech, voice cloning, user settings, billing/entitlements, and Soundboard media workflows.

Source Of Truth

Repository source, tests, tracked docs, and verified live state are authoritative. This vault is the synchronized navigation, status, and work-history layer.

Current Product State

Runtime And Architecture

Active Surfaces

STS Backend Contract

Active Work

  1. Automate same-commit LA/Canada web deployment through a secure dual-site promotion path. Cloudflare failover itself is implemented and drill-verified.
  2. Finish Android Play release evidence: signed upload, Play Console, screenshots, and live smoke proof.
  3. Retire legacy runtime compatibility namespaces and replace monkeypatch-heavy tests with focused service seams.
  4. Capture deployed native generate-to-Soundboard smoke evidence.
  5. Capture one new live combined-row database/UI proof that the deployed combined_from.clips snapshot exactly matches its selected source clips.
  6. Install build 28 on the iPhone 16 Pro Max and iPad Pro M4, then capture cold download, warm load, first-audio, real-time factor, peak memory, thermal, battery, long-output, and same-reference ElevenLabs comparison evidence.
See [[SpeakTrue Planning Status SpeakTrue Planning Status]] for proposed, decision-gated, implemented, and retired work.

Documentation Closeout Gate

Every future agent task that changes tracked files or verified product/deployment/planning state must:

  1. add a dated entry to [[SpeakTrue Work Log SpeakTrue Work Log]];
  2. refresh the codebase map with python3 scripts/refresh_obsidian_docs.py;
  3. update this note and/or [[SpeakTrue Planning Status SpeakTrue Planning Status]] when semantics changed;
  4. pass python3 scripts/refresh_obsidian_docs.py --check.

The repo-level requirement lives in AGENTS.md; the executable runbook is docs/ops/OBSIDIAN_DOCUMENTATION_SYNC.md.

Verification Anchors

python3 scripts/refresh_obsidian_docs.py
python3 scripts/refresh_obsidian_docs.py --check
bash scripts/verify_android_ci_local.sh
bash scripts/regression_migration_paths.sh
python3 scripts/verify_android_ios_parity_release.py
python3 scripts/verify_runtime_import_boundaries.py
python3 scripts/verify_legacy_runtime_retirement.py
git diff --check
git status --short

Notes For Future Agents

Proposed local generation controls

The Local Qwen generation settings plan (docs/product/LOCAL_QWEN_GENERATION_SETTINGS_PLAN.md) is in progress. Phase 0 is closed on user confirmation that generation works at defaults. Phase 1 (validated contract, per-voice persistence, immutable per-request snapshots) is committed at cce80bf6. Phase 2 (committed at 2a335f1c) wires the sampling settings through the Qwen adapter — defaults resolve to the prior fixed behavior — and adds always-visible sampling controls with a Variation slider, save/reset per voice, help text, a low-ceiling truncation warning, and a determinate generation progress bar with section and tokens/s readouts. Build 1.3 (41) is installed on Yoseif’s iPhone; focused device tests pass 21/21 and the Android CI-equivalent gate passed. Device listening checks at defaults and selected bounds are pending with the user; phases 3-5 remain open.