# SpeakTrue Planning Status

## 2026-09-09 — Successful Qwen 6-bit sampling runs, user-confirmed quality

Latest device history confirms two runs with the same 223-character script and reference, temperature 0.9, top-k 50, repetition penalty 1.05452 and ceiling 4096. Top-p 1 produced 16 s audio in 14.7589 s (RTF 0.922, first internal audio 2.2071 s). Top-p 0.90248 produced 14 s audio in 12.6342 s (RTF 0.902, first internal audio 1.2889 s). Device tokenizer encodes this script to 47 tokens, giving a hidden SDK cap of 22.56 s; both outputs ended below that cap, consistent with natural EOS. Explicit token traces are still unavailable. User reports the unwanted tail is fixed and quality/speed are decent; no independent listening claim. Different script from the earlier failed 269-character run prevents isolating causality to sampling. Both top-p values worked; do not credit top-p reduction alone.

Sampling settings are now user-validated candidates; no app defaults or runtime code changed. Broader long-text/repeated-run validation remains pending. Device history retrieved read-only and token count checked locally. No audio/transcript content copied into documentation. Repository dirty state preserved atop 13dfb226, build remains 62, no new repository files, no commit/push/build. Map refreshed and checked at closeout.

## 2026-09-09 — Qwen stopping investigation; build remains 62

Read-only device tokenizer retrieval and local encoding verified 59 text tokens for the latest run. Pinned SDK hidden limit is min(user ceiling, max(75, text tokens times 6)): 354 codec frames / 12.5 Hz = 28.32 seconds, exactly matching the WAV. EOS is checked before appending audio frames, so evidence strongly indicates limit exhaustion without natural EOS; no token trace was retained. The reason EOS was not selected remains unresolved. Inspected suppression, sampling, ICL text-EOS prompt and stream finalization; no missing EOS check or duplicate final append found. Runtime currently ignores token events. Current sampling differs from upstream defaults (top-k 0 versus 50, repetition penalty 1.1 versus 1.05); causality unproven.

Natural stopping investigation is active. Next proposed slice: retain per-section EOS/limit diagnostics, then controlled repeated comparisons with the same script/reference and one sampling setting changed at a time. User excludes blunt duration trimming and transcription-based stopping. No runtime settings or generation behavior changed, and no new build/device generation was performed. Additional device config retrieval timed out in CoreDeviceService after tokenizer retrieval succeeded; device config was not reverified. No private transcript/audio included here.

Changed this slice: docs/guides/ios-local-qwen-prototype.md only; prior 12-file dirty state preserved atop 13dfb226. No new uncommitted files. Verified tokenizer count, WAV sample duration, pinned source and upstream source; git diff --check passed. No commit/push. External map regenerated and checked at closeout.

## 2026-09-09 — Qwen tail diagnosis and single temperature control, build 62

Retrieved latest Qwen 6-bit history/WAV from the phone. Total generation 23.2931 s; output 28.32 s; first internal chunk 1.8229 s. Settings: temperature 0.9, top-p 1, top-k 0, repetition penalty 1.1, ceiling 4096. Local cached Whisper-base places final requested words around 17.7 s; trailing approximately 10.6 s contains low-level nonzero sound, mostly -38 to -51 dBFS in one-second RMS windows. No external ASR received audio. ASR is supporting evidence, not proof of silence/non-speech. Inspected pinned Qwen streaming path emits each chunk once plus final remainder; no full-output duplication found. Late end-of-sequence behavior remains unresolved. File RTF 0.823 includes unwanted tail and is not a useful-speech speed claim.

Confirmed Variation and Temperature were duplicate bindings. Removed the redundant Variation control; one Temperature (variation) slider remains with clear help. Values/ranges and generation behavior unchanged. Prepared separate 18.30 s trimmed preview under /private/tmp/speaktrue-qwen6-trimmed-preview.wav; original phone clip unchanged. No automatic trimming introduced; avoid cutting quiet intended speech. No audio or transcript content stored in repo/vault.

Release 1.3 (62) built, installed over the existing iPhone app, launched and device version verified. Android/iOS parity passed 12 surfaces/24 gates; git diff --check passed. No new tests needed for the duplicate-control removal; prior build-61 suite had 52 passing focused tests. No physical tap-through or new generation result claimed.

This slice changes ios/SpeakTrue/LocalQwenPrototypeView.swift, ios/SpeakTrue.xcodeproj/project.pbxproj, android/app/src/test/java/com/speaktrue/contracts/AndroidReleaseReadinessContractTest.kt (iOS build assertion), docs/guides/ios-local-qwen-prototype.md, docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md and docs/product/LOCAL_QWEN_GENERATION_SETTINGS_PLAN.md. Prior changes preserved; 12 modified tracked files atop 13dfb226, no new uncommitted files. No commit/push, GitHub Actions or TestFlight upload. Full local Android gate remains required before commit/push. External map refreshed/checked at closeout.


## 2026-09-09 — Manual model unload, build 61

Added Unload from memory in the expanded Local TTS model controls. It names and unloads the model actually in memory, even when a different model is selected. Busy state prevents generation, load/delete, selection and repeated unload during the operation. Runtime unload releases weights, cached reference waveform/tokens and MLX cache; storage indicators refresh afterward. Downloaded model files, voice profiles, generated output and history remain. Generation requires reloading a model. Loading a replacement still unloads automatically; selection alone does not.

Verification: 52 focused LocalQwenPrototypeTests passed; signed Release 1.3 (61) built and installed over the existing app on Yoseif’s iPhone. Android/iOS parity passed 12 surfaces/24 gates; git diff --check passed. Runtime/button wiring reviewed, but no instrumented on-device RAM measurement or tap-through result claimed. OmniVoice accuracy remains unresolved from the earlier report; this task does not change generation math.

Changed this slice: ios/SpeakTrue/LocalQwenViewModel.swift, LocalQwenPrototypeView.swift, ios/SpeakTrue.xcodeproj/project.pbxproj; android/app/src/test/java/com/speaktrue/contracts/AndroidReleaseReadinessContractTest.kt (iOS build assertion only); docs/guides/ios-local-qwen-prototype.md and docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md. Prior build-60 changes preserved. Total working tree: 11 modified tracked files atop 13dfb226; no new uncommitted files. No commit, push, GitHub Actions or TestFlight upload. Full local Android gate is required before committing/pushing the accumulated parity assertion. External map refreshed/checked at closeout.


## 2026-09-09 — Build 60 completes but speech fidelity fails

Build 60 completed the latest on-device OmniVoice run, but the user reports inaccurate speech; local cached Whisper-base ASR supports omissions/reordered words. Output 18.56 s, total 50.2075 s, RTF 2.705; diffusion 46.2285 s (92.1%), decoding 1.6797 s, reference encoding 2.2327 s (cache miss), preparation 0.0386 s and assembly 0.0257 s. Settings: 16 steps, guidance 2, speed 1. This is not the same text as the earlier baseline, so no controlled speedup is claimed.

The selected reference is 34.3846 s, with broadly matching transcript according to local ASR. Upstream OmniVoice recommends 3–10 s and warns that longer references can slow inference and degrade quality (https://github.com/k2-fsa/OmniVoice#voice-cloning). Our generic 10–30 s app guidance is unsuitable for this family and remains to be corrected. Long conditioning is a hypothesis, not confirmed causality; real-codec window equivalence also remains unverified. Quality status remains failing pending repeat comparison.

Prepared a separate 7.10 s first-sentence reference and matching transcript under /private/tmp/speaktrue-omnivoice-reference-test for user import as an alternative. Original phone profile/recording unchanged. Excerpt duration and local ASR checked. User should compare the same target text with unchanged 16-step/guidance/speed settings before further tuning. No external transcription service or external model received the audio; no audio/transcript content copied to the vault or repository.

This turn changed only docs/guides/ios-local-qwen-prototype.md and docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md; prior nine-file build-60 dirty state remains uncommitted atop 13dfb226. No new uncommitted repository files. No new app code, build, install, commit/push, CI run or TestFlight action. git diff --check passed; external map refreshed/checked at closeout.


## 2026-09-09 — OmniVoice memory termination and consent retry, build 60

Build 60 is implemented, signed Release-built, installed and launched on Yoseif’s iPhone; device app inventory confirms 1.3 (60). Device jetsam reports at 21:27:03 and 21:29:31 on 2026-09-09 explicitly identify SpeakTrue as killed for vm-pageshortage. The latest has 358358 resident pages at 16384 bytes/page, approximately 5.47 GiB. The exact generation stage is unproven. Initial CoreDevice 12040/12010 developer-image errors later cleared. Reports remain outside the repository/vault; no raw logs or user audio/text copied here.

Memory mitigations: evaluate conditional and unconditional diffusion passes separately, and decode 100-frame windows with 32 context frames on each side using the unchanged pinned codec. Trim shared context without additional fades, silence or crossfades; materialize CPU samples and release each decoder graph before the next window. Validate codec geometry, decoded length and finite samples; emit content-free Release phase/memory logs. The default sampling calculations and final normalization remain unchanged. Physical repeat-generation, listening continuity and measured memory/latency improvement remain pending.

Consent bug: failed remote lookups no longer mean missing consent. Preserve same-account verified acceptance on transient refresh failure, show a retry screen for unknown status, and reject stale responses after account reset/change. Only an actual negative result opens the notice. This fixes a code path that could explain the reported prompt, without claiming that the exact network failure was captured. The app-wide consent requirement remains; OmniVoice does not invoke ElevenLabs generation.

Verification: 52 focused Simulator tests passed (consent failures/account isolation, CPU-only decoder-window boundaries/length/errors and existing local TTS tests). Initial MLX-in-Simulator tests aborted in Metal initialization; the final tests exercise window orchestration without claiming real codec inference. Final Release build passed, Android/iOS parity passed 12 surfaces/24 gates, and git diff --check passed. No GitHub Actions triggered; no commit or push. Full Android local gate remains required before a future commit/push of the build assertion.

Working tree atop 13dfb226: nine modified tracked files: ios/SpeakTrue/AIDataSharingAgreementService.swift, AIDataSharingConsentGateView.swift, LocalOmniVoiceModel.swift, LocalQwenRuntime.swift; ios/SpeakTrueTests/LocalQwenPrototypeTests.swift; ios/SpeakTrue.xcodeproj/project.pbxproj; android/app/src/test/java/com/speaktrue/contracts/AndroidReleaseReadinessContractTest.kt (iOS build assertion only); docs/guides/ios-local-qwen-prototype.md; docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md. No new uncommitted files. No archive/TestFlight upload.


## 2026-09-09 — Build 59 changes committed and pushed

2026-09-09: All 18 pending files committed and pushed to origin/main as 13dfb226ea6631619acb44f144c0818d67fbed41 (iOS 1.3 build 59). Includes Local TTS reliability, OmniVoice orchestration/configuration, TTS/STT integration, model controls, private history, tests and product documentation. Android source change is solely the release test assertion for iOS build 59. Commit includes [skip ci] as requested; no workflow configuration was changed. The full local Android CI-equivalent gate passed (unit suite, lint, release bundle, parity/readiness and backend contracts); external live checks remain unverified. Previously verified iOS Release build/install and focused tests remain the iOS evidence. Physical audio quality and latency testing remain pending. Working tree is clean; all three formerly untracked Swift files are now committed.

Earlier entries below are historical snapshots. The pending local commit gate is now satisfied; no archive or TestFlight upload was performed.

## Current device delivery

2026-09-08: Optimized Release 1.3 (59) rebuilt successfully, installed over the existing app on Yoseif's iPhone 16 Pro Max, and launched successfully. Device app inventory confirms version 1.3 / build 59. No uninstall was performed. Physical generation, listening quality, latency and full UI smoke testing remain pending. No archive or TestFlight upload. Repository source unchanged by this delivery; existing changes remain uncommitted and unpushed atop d182a818.

## Current generation-first Local TTS workspace

2026-09-08, build 59: all seven approved UI recommendations are implemented locally. A compact model/voice header leads into editable generation text; coloured models and device details expand from the model row, and Manage voices holds reference creation/editing. The voice menu and preview remain accessible in the header. A persistent Generate/Cancel bar shows blocking reasons, actual stage, elapsed time and optional timing estimates. Estimates require three comparable runs with the same saved profile/reference, model, language and settings and similar text lengths; thermal/cache effects remain unmeasured. OmniVoice emits a decoding event; other SDK families retain coarser progress. Inline section colours remain, with collapsed explanatory help.

Generation snapshots current editor settings immediately; Save as voice defaults controls persistence only. Results offer playback, regenerate, Share and Save WAV, with detailed metrics in Performance details. A device-only actor store retains the latest 20 generated WAVs and text/settings metadata in Application Support/LocalTTSHistory with file protection and backup exclusion. History survives sign-out and supports replay, reuse and individual deletion. It stores no reference recordings; deleted or unavailable historical references require selecting a voice. WAV export makes an independent copy. No backend or cloud Soundboard changes.

Verification: 46 focused Simulator tests passed, zero failed; a subsequent targeted rendering test passed after the final voice-menu layout correction. Normal and accessibility-size standalone Local TTS renderings were inspected; an overlapping native voice picker and an oversized sticky explanation were corrected. Optimized generic-iOS Release 1.3 (59) built successfully. Android/iOS parity passed 12 surfaces/24 gates; git diff --check passed. This is compile, contract and simulator-render evidence, not full-device UI/inference or measured latency validation. Installed and launched on the phone after verification; generation testing remains pending. No archive or TestFlight upload. Uncommitted atop d182a818; full Android CI-equivalent gate remains required before commit/push.

## Current model selection buttons

2026-09-08, build 58: implemented coloured buttons for all five Local TTS models, replacing the dropdown and duplicate inventory list. Buttons show model name, approximate download size, storage/preparation status and an explicit Loaded in memory indicator; loaded models also show Downloaded. A selection checkmark and border provide non-colour cues and VoiceOver receives the selected trait. The grid adapts to available width and becomes a flexible single column for accessibility text sizes. Selection remains disabled during preparation/generation; download/load/cancel remain separate actions. Model inventory now refreshes after preparation failures as well as success/cancellation/deletion.

Verification: optimized generic-iOS Release 1.3 (58) built successfully; Android/iOS parity passed 12 surfaces/24 gates and git diff --check passed. Source/layout review only; no rendered-device or inference claim. No new tests were added for this small UI change; the preceding build-57 focused suite had 41 passing tests. Device layout/status smoke testing remains pending. Uncommitted atop d182a818; no install, archive or TestFlight upload. Full Android gate remains required before commit/push.

## Current TTS tab integration

2026-09-08, build 57: implemented locally. The TTS screen now offers Standard / Local TTS modes, replacing the Voice Cloning lab entry. STT Copy to TTS fills both independent drafts through the main tab container even when a mode is hidden; later edits remain independent. STS is unchanged. The main tab owns the local model state, preserving the draft, loaded model, profiles and active operations across mode switches. Leaving the local view pauses playback and cancels active reference recording; leaving the main tab container performs full temporary-reference cleanup.

Verification: 41 focused Simulator tests passed with zero failures, including shared transcript replacement, repeat copies, Unicode/newlines and independent later edits. Optimized Release 1.3 (57) built successfully. Android/iOS parity passed 12 surfaces/24 gates; git diff --check passed. Physical-device navigation, generation quality and latency remain pending. No installation, archive or TestFlight upload. Work remains uncommitted atop d182a818; the full Android gate is still required before commit/push. Earlier build-56 cache/timing work remains included.

## Current OmniVoice reference reuse

2026-09-08, build 56: code-only OmniVoice latency work is implemented locally. A single in-memory reference-token cache encodes once before the section loop and reuses unchanged audio across later runs. Keys use reference-file SHA-256 plus sample rate, not temporary paths; model replacement/unload and profile deletion clear it, and failed/cancelled encoding is not cached. The app-owned LocalOmniVoiceModel/Config adapter reuses the SDK backbone, codec and parameter type while exposing prepared-token generation and phase timings. Default and custom settings now show diffusion-step progress. The timing panel reports reference preparation/encoding, diffusion, decoding and WAV assembly, cache reuse and build mode. Physical-device speed and quality remain unverified while the user is away.

Validation: 40 focused Simulator tests passed, zero failed/skipped; parity verifier passed 12 surfaces/24 gates. The diffusion loop and 11 core helper bodies match pinned upstream d9e6e7c after excluding timing instrumentation and local type naming. This is source/regression evidence, not a measured latency improvement. Code is uncommitted atop d182a818. Earlier proposed-cache status below is superseded; presets, lower-precision evaluation and early section playback remain pending.

## Active OmniVoice latency investigation

Measured user baseline (2026-09-08): OmniVoice at 16 diffusion steps, two generation sections, first-audio metric 52.52 s, total generation 141.93 s, output 37.48 s, real-time factor 3.79. Local ffprobe independently confirms the attached WAV is mono float32 PCM at 24 kHz and 37.475 s. No listening-quality assessment is claimed from that metadata check. The first-audio metric marks an internally completed section; playback currently waits for the assembled output. Stage-specific timing and the benefit of reference-token reuse remain unmeasured. Optimized Release 1.3 (55) remains ready but uninstalled; the paired iPhone is still unavailable. No reference audio or transcript contents copied into these notes.

Delivery status: optimized Release iOS 1.3 (55) built successfully; -O verified for SpeakTrue, MLX and MLXAudioTTS. Installation could not proceed because Yoseif's paired iPhone 16 Pro Max is unavailable. Reconnect the phone to install and measure the same generation. No speed gain is yet measured.

2026-09-08 latency investigation: user still finds OmniVoice too slow at 16 steps. Prior device installs were Debug builds with Swift -Onone. Preparing an optimized Release build of the same iOS 1.3 (55) source as the next performance baseline. Source inspection confirms reference encoding repeats per section and the public OmniVoice API does not accept encoded reference tokens; waveform delivery waits for a whole section. No latency improvement or new reference cache is claimed until measured on the phone.

Profiling is active; encoded-reference caching remains proposed. The earlier inline editor and cache repair remain implemented and uncommitted. Earlier snapshots follow.

## Current inline section colours

Build 55 (2026-09-07): generation sections now use alternating text colours inside the editable generation box. The separate section-boundary preview is removed; the section count and character limit remain. Original whitespace and emoji are preserved, and the colours use the actual generation chunker. All 36 focused Simulator tests passed, including Unicode/whitespace range coverage; Android/iOS parity checks passed. User reports very good OmniVoice quality after build 54, with slow generation on iPhone 16 Pro Max. Profiling and encoded-reference reuse remain proposed, not implemented.

Implementation remains uncommitted atop d182a818 with prior fixes preserved. This summary supersedes the earlier status snapshots below. Full Android gate remains required before commit/push.

## Current OmniVoice repair

Build 54 OmniVoice repair (2026-09-07): device inspection confirmed a separate SDK-default model cache containing only 281,103,571 bytes versus the pinned 2,450,344,102-byte model present in the app-managed folder. Nested OmniVoice resolvers ignore the supplied cache; the old seed helper targeted its own source directory. The runtime now atomically publishes verified files into HubCache.default before loading, replaces existing entries and removes that cache on deletion. A regression test covers truncated and same-size corrupt cache files, repeated hard links and source preservation. All 35 focused Simulator tests passed. This confirms the cache repair, not audible speech quality; reload and generation/listening validation remain required.

State: implemented locally atop d182a818, uncommitted. Earlier build-53 review fixes remain included. The full Android gate is still required before committing/pushing its build assertion. This summary supersedes earlier current-state snapshots below.

## Current local synchronization

Review 2026-09-07: origin/main remains d182a818a720dae19c5a311bdb78f216751ce63e. Local iOS 1.3 (53) corrections are implemented and uncommitted: cancellation reaches the owned transfer, an independent 25-second watchdog detects silent stalls, initial progress and known-size fallback remain visible, family switches reload saved editors, Chatterbox exposes only its supported temperature/top-p controls, and custom OmniVoice first-audio timing records the first section. All 34 focused Simulator tests and the signed generic iOS device build passed; the Android/iOS parity verifier passed. Physical tests were blocked by the locked iPhone; no new installation or listening validation is claimed. The full Android gate remains required before committing/pushing the updated build assertion. Chatterbox/OmniVoice listening, paragraph/speed work and presets remain pending; the prototype guide was refreshed.

Older status statements below are retained as historical context; this summary and the latest dated work-log entries supersede them.

Status: locally synchronized as of 2026-09-07.

This note mirrors the planning taxonomy in `docs/index.md`. Repository source, tests, tracked docs, and verified live state remain authoritative.

## Active Execution And Release Work

- Local TTS validation remains active. Settings phases 1–2 and multi-model phases through 5b are committed; build 53 review corrections are implemented and locally verified, uncommitted. Chatterbox/OmniVoice listening and remaining settings-plan work are pending. The guide refresh is complete. Full Android gate and physical-device validation remain outstanding.

### Web And iOS Security Remediation

- Database/index decision resolved by user approval: existing per-user Supabase tables are authoritative in every legacy storage-adapter mode. Full category/clip action migration is implemented and locally verified, not deployed. New objects are uniquely namespaced/indexed; existing references remain stable across lifecycle operations. No automatic ownership assignment/import of historical shared data. DB connectivity and migrated schema are required, with no shared-index fallback; Supabase-only startup policy remains unchanged.

- Scoped implementation and local verification complete, uncommitted; deployment is not authorized or performed. Owner-bound generated media/regeneration, guarded static delivery, provider-specific permission, text-safe messages, headers, iOS Keychain migration and local-reference protection/lifecycle are implemented.
- Verification: 1045 primary final web tests passed; four relevant JavaScript runtime checks passed. Actual isolated PostgreSQL checks cover ownership, moves, restore, collision rejection, atomic conversion and regeneration CAS. Synthetic desktop/mobile account-category and move-destination interaction passed. Earlier iOS evidence remains 163 unit/15 primary focused passes (unchanged this continuation).
- Rollout remains outstanding: additive restore SQL correction `20260904120000_fix_restore_clip_sort_order.sql`, deployment approval, complete Android local CI-equivalent gate before commit/push of the shared backend change, and deployed account/storage verification. No current release or all-clear certification is claimed.
- Open decisions/evidence: recording provenance and any history removal; server-generated-media retention cleanup; legacy inline CSP and browser-local token hardening; approved staging rollout and physical-device lock-state checks. Old ownerless artifacts require regeneration; shared ownerless cache and legacy combined-download URLs are disabled.
- Repo anchor: `docs/ops/WEB_IOS_SECURITY_REMEDIATION.md`. No production data removal, commit, push, or deployment.

### LA And Canada Deployment Automation

- Canada warm standby is implemented on VM `2201` with a healthy Docker app,
  dedicated `speaktrue-web-ca` tunnel, and verified daily PBS backup.
- Cloudflare LA-primary/Canada-fallback routing is implemented and
  failure-drill verified on release `15ca9d55`.
- Remaining work: replace the manual Canada checkout promotion with a secure
  workflow that deploys and verifies the same immutable commit at both sites.
- Repo anchor: `docs/ops/SPEAKTRUE_LA_CA_LOAD_BALANCER_PLAN.md`.

### Android Play Release

- Source readiness and local release gates exist.
- Current native speech code includes structured failure preservation, restored realtime session compatibility, pronunciation-dictionary application, and simplified playback controls.
- Remaining evidence: signed upload, Play Console configuration, screenshots, and live smoke proof.
- Repo anchors: `docs/ops/ANDROID_PLAY_RELEASE_CHECKLIST.md`, `docs/ops/ANDROID_PLAY_RELEASE_EVIDENCE_TEMPLATE.md`.

### Legacy Runtime Retirement

- Production routes are off `legacy_runtime_adapter.py`.
- Remaining work is compatibility namespace retirement for `src/legacy_runtime.py`, `src/services/legacy_runtime_adapter.py`, and monkeypatch-heavy tests.
- Repo anchors: `docs/ops/legacy-runtime-inventory.md`, `docs/ops/legacy-runtime-compatibility-removal.md`, `docs/ops/legacy-runtime-seam-deletion.md`.

### Native Speech Parity Proof

- Canonical artifact metadata and server-side generated-artifact save contracts are implemented locally.
- Deployed native generate-to-Soundboard smoke evidence remains.
- Repo anchor: `docs/contracts/native-speech-parity-anchor.md`.

## Proposed Product Roadmap

### 90-Day Sequence

- Direction: move SpeakTrue from separate speech tools toward a communication workspace that preserves context, makes speech reusable, and stays useful offline.
- Sequence: contract and ephemeral UX proof, durable conversation/library contracts, native parity and fallback, integration/shortcuts, then staged hardening.
- Repo anchor: `docs/product/NEXT_PRODUCT_DIRECTION_90_DAY_ROADMAP.md`.

### Conversation Workspace

- Goal: keep completed Turn-Based STS turns in a durable replayable timeline across web, iOS, and Android.
- Important boundary: current STS revoices committed transcript text; it does not yet implement a separate source-to-target translation step.
- Planned slices: accurate semantics, ephemeral timelines per platform, RLS-backed conversation persistence, native resume, immutable edit/re-speak, recovery, and staged rollout.
- Repo anchor: `docs/product/CONVERSATION_WORKSPACE_PLAN.md`.

### Unified Speech Library

- Goal: add global recent, search, favorites, tags, and source filters over existing Soundboard clips without replacing categories.
- Planned contract: paginated owner-scoped query, lazy signed-URL resolution, first-class favorite/tag state, and web/iOS/Android parity.
- Repo anchor: `docs/product/UNIFIED_SPEECH_LIBRARY_PLAN.md`.

### Quick Speak Accessibility Mode

- Goal: reduce the path from typed or pinned phrase to audible speech, with clearly labeled platform voice fallback when cloud TTS is unavailable.
- Privacy boundary: phrase text stays out of telemetry, logs, shortcut URLs, and crash breadcrumbs.
- Planned slices: local iOS and Android MVPs, device-voice fallback, synced pins, Library integration, opaque-ID shortcuts, and accessibility hardening.
- Repo anchor: `docs/product/QUICK_SPEAK_ACCESSIBILITY_PLAN.md`.

## Future Or Decision-Gated Work

### Supabase To GCS Backup

- Design-only future direction.
- No worker, scheduler, or restore drill currently exists.
- Repo anchor: `docs/ops/SUPABASE_GCS_BACKUP_FUTURE_DIRECTION.md`.

### iOS Subscriptions

- Decision-gated reintroduction plan; current product remains free/open-access.
- Repo anchor: `docs/ops/IOS_SUBSCRIPTION_REINTRODUCTION_PLAN.md`.

### Server-Side Strict-Mode Drafts

- Intentionally deferred.
- Local drafts remain authoritative until an explicit Supabase persistence migration is designed and implemented.

## Implemented, Retired, Or Historical References

### Implemented

- iOS Local Qwen voice prototype: core lab committed in 4cc7b69d; build 31-40
  download/generation stabilization committed in 311cf548 on 2026-09-06 (pinned HF
  revisions with per-file expected sizes and SHA-256, staged transfers with
  verifyHashes:true, progress/cancellation UX, profile lifecycle hardening,
  increased-memory-limit entitlement, build 40, regression tests, parity assertion,
  guide update). Build 40 is installed/launched on the physical iPhone. The Voice Cloning surface can download/load pinned MLX Qwen3-TTS 0.6B
  Base quantizations, persist on-device reference-audio/transcript profiles,
  perform chunked local WAV generation, share output, and expose latency and
  real-time-factor metrics. Generic physical-iOS plus iPhone/iPad simulator
  builds pass and focused tests pass. Physical iPhone 16 Pro Max and iPad Pro
  M4 inference, memory, thermal, battery, latency, and quality proof remains an
  active validation step; this is not yet a production-ready mobile runtime.
  Repo anchor: `docs/guides/ios-local-qwen-prototype.md`.

- Current-speaker ASR feasibility benchmark: the full VibeVoice ASR HF model
  and Canary-Qwen 2.5B were downloaded and executed on the M1 Ultra against
  seven approved current-voice samples. On three clips with known references,
  Canary produced 33.8%-71.4% word error and VibeVoice's best passes produced
  64.7%-100%; VibeVoice also failed on the long technical sample. By contrast,
  MAI-Transcribe 1.5 produced 11.8%-21.4% unprompted word error, while Scribe
  V2 with a narrowly tailored per-speaker vocabulary reached 0%-4.4%. Current
  product direction is Scribe plus a consented per-user glossary, with MAI as
  the comparison or fallback; fine-tuning is deferred until more corrected
  reference transcripts exist. Neither local architecture can be served
  faithfully by stock LM Studio, Ollama, or Docker Model Runner on macOS;
  local testing uses their official Transformers and NeMo runtimes.
- OpenRouter staging transcription trials: implemented and live on isolated
  dev release `dev-d88346208a85` from exact commit `d8834620`. The web STT and
  Standard batch-STS surfaces expose MAI-Transcribe 1.5, Whisper Large v3, and
  Voxtral Mini Transcribe only to the exact authenticated staging account.
  These three replaced the initial choices after a live thirteen-model
  accessibility-speech comparison. Server-side feature/key/allowlist gates
  remain active; realtime STS and native surfaces were not changed.
- OpenRouter free-model assessment: no zero-cost model is currently available
  through the dedicated transcription catalog. Newly listed Fish Audio
  Transcribe 1 returned provider HTTP 502 across raw and 16 kHz test inputs, so
  it was not exposed as a staging choice.
- LA-primary/Canada failover for `web.speaktrue.cc`: implemented with two
  tunnel-backed pools, HTTPS `/health` monitoring, ordered failover, Canada
  fallback, session affinity off, and a successful controlled failover/failback
  drill.
- Generate Speech to Soundboard target: implemented acceptance and regression reference.
- Canonical generation artifact metadata: implemented across backend and supported clients.
- Soundboard category ZIP export: implemented on the legacy web surface with progress feedback.
- Mobile pronunciation-dictionary application: implemented across supported native TTS/STS paths.

### Retired

- Web long-form TTS/Studio pilot: retired after provider entitlement validation. The product uses standard direct TTS with a 10,000-character ceiling; `tts-longform` is a `410 Gone` tombstone.
- Repo anchor: `docs/product/WEB_LONG_FORM_TTS_PLAN.md`.

### Historical

- Realtime STS workflow plan: implemented realtime checkpoint, superseded as the default by Turn-Based STS.
- Monorepo migration inventory: completed structural migration record.
- Android visual refresh: completed historical design evidence, not an active queue.

## Current STS Direction

- Primary: Turn-Based STS (`live_interpreter`).
- Compatibility: `batch` and `realtime`.
- `sts-live-interpreter-session` and `sts-live-interpreter-tts-session` are implemented, not scaffolds.
- Automatic and manual turn completion are both supported.
- Current backend/native compatibility work preserves structured failure metadata and accepts the legacy Android realtime request shape.

## Maintenance Rule

Update this note whenever `docs/index.md` or verified implementation state changes status classification. Do not promote proposed or future work to active/implemented without code, workflow, or verified external-state evidence. Every agent task must also add an entry to [[SpeakTrue Work Log|SpeakTrue Work Log]] and pass the repo’s Obsidian sync gate.

## Related

- [[SpeakTrue|SpeakTrue]]
- [[SpeakTrue Work Log|SpeakTrue Work Log]]
- [[Codebase Map/SpeakTrue Codebase Map|SpeakTrue Codebase Map]]

## In progress — Local Qwen generation settings (2026-09-05)

Plan prepared at docs/product/LOCAL_QWEN_GENERATION_SETTINGS_PLAN.md. Sequence: capture inference baseline; validated settings contract and explicit per-voice save/reset; sampling controls; structured section boundaries and paragraph pauses; pitch-preserving offline speed; listening-based presets and device closeout. Phase 0 is closed on user confirmation that on-device generation works at defaults (aggregate timing/memory metrics remain uncaptured in docs). Phase 1 (settings contract and per-voice persistence) is committed at cce80bf6. Phase 2 is committed at 2a335f1c: the runtime resolves the request's settings snapshot through a revalidating sampling resolver and forwards temperature, top-p, top-k, repetition penalty and the audio-token ceiling to the Qwen adapter (defaults match the prior fixed behavior); the prototype card shows a Variation slider, always-visible sampling controls with help text and low-ceiling warning, save/reset-for-voice actions, and a structured determinate progress bar with section and tokens/s readouts. Focused device tests pass (21/21) and the full Android CI-equivalent gate passed after the build 41 parity bump. Device listening checks at defaults and selected bounds are pending with the user; phases 3-5 (section boundaries/pauses, speaking speed, presets/closeout) remain open.

## In progress — Local TTS Lab multi-model (2026-09-07)

Plan prepared at docs/product/LOCAL_TTS_LAB_MULTI_MODEL_PLAN.md. Generalizes the Local Qwen Lab into a Local TTS Lab hosting multiple on-device model families with the same download-verify-load-generate flow. Phases 0-4 are implemented (2026-09-07, pending on-device verification): phase 0 discovery pinned `mlx-community/Chatterbox-TTS-fp16` @ 77c7c8f9 (3-file manifest, 2.70 GB) and `mlx-community/OmniVoice` @ defe4bdc (9-file manifest, 3.27 GB fp32 required for cloning) with SHA-256 manifests, plus the documented package-managed S3TokenizerV2 auxiliary download; phase 1 renamed user-facing copy to "Local TTS Lab"; phase 2 introduced the family-aware catalog (`LocalTTSModelFamily`/`LocalTTSModelProfile`/`LocalTTSModelCatalog`) with loader model-type strings from the family, family-grouped model picker, per-family language lists, and memory thresholds; phases 3-4 added the Chatterbox and OmniVoice profiles with quality/cloning copy (OmniVoice sampling sliders do not apply — diffusion defaults). OmniVoice's local load is routed through fromPretrained after hard-linking the verified files into the flat package-cache location (vendored dispatcher lacks a local closure — upstream PR recorded). Downloads are an owned, fully-traced transfer: runtime-private URLSessionDownloadTask with a session-level delegate (the async per-task delegate path did not deliver progress on device), byte-level progress with a smoothed speed readout ("37% · 900 MB of 2.45 GB · 7.4 MB/s"), a 25 s stall watchdog on real delegate bytes, and three fresh retries. User confirmed Qwen 4-bit and OmniVoice downloads with moving progress bars. Phase 5a implemented: family-scoped per-voice settings (overrides keyed per family with legacy-blob migration into the Qwen slot), family-gated sampling UI (OmniVoice shows its diffusion-defaults note instead), and family-aware resolution in generation requests. Phase 5b implemented: OmniVoice diffusion knobs (numSteps/guidanceScale/speed, per-voice persisted, revalidated by a resolver; library defaults keep the streaming path with step fractions, custom knobs use the non-streaming ovParameters overload with section-level progress). Pending in phase 5: A/B listening comparisons and presets closeout, prototype guide refresh. Focused device tests pass 26/26. Pending: on-device download/load/generation and listening checks for the two new families (8 GB classification is experimental); phase 5 (family-scoped settings + A/B closeout) open. Future direction recorded in the plan: Mac-assisted per-user LoRA "studio voices" (mlx-tune/mlx-audio-train ecosystem, Qwen3-TTS stable) — on-device training assessed as not practical; zero-shot remains the default pending listening evidence. Internal LocalQwen* type/file renames explicitly deferred to a mechanical follow-up.
