SpeakTrue

Saved clip metadata

Persistence boundary

The authoritative media index is public.soundboard_clips, scoped by user_id and category_id. Audio and optional transcript objects live in the Soundboard bucket. The row contains the clip ID, file name, storage and transcript paths, transcript, MIME type, byte size and duration when known, order, offline flag, and database timestamps. metadata is JSONB; no new database migration is needed for the additive generation fields below.

Generation settings describe the audio that was produced. Capture them when starting/completing generation, and retain that snapshot when the editor, model, voice, or settings change. Save options describe the exported file. A server-owned generation_artifacts snapshot takes precedence over client generation metadata. Full and streaming tts-generate audio responses expose the same snapshot in X-Speech-Clip-Metadata as percent-encoded JSON. Clients decode the header when present and retain their captured fallback snapshot for older deployments. Web local artifacts use their owner-checked sidecar snapshot. Never substitute current editor settings for a completed result’s generation settings.

Common JSON keys

All keys below are optional for reading historical records. Omit unavailable or inapplicable values rather than inventing defaults. Newly generated snapshots include schema_version: 1, source, generation_type, provider, and generated_at when the generation path supplies them.

Field Meaning
source Originating platform or server workflow, such as ios, android, web, tts-generate, or an STS workflow. Saving must not relabel the generation source.
generation_type tts or sts; existing combined clips retain their separate combined_soundboard_clip structure.
provider Provider identity; local for on-device synthesis, elevenlabs for its hosted/direct output.
generated_at ISO 8601 generation timestamp. Database creation time remains save time.
tts_settings voice, voice_name, model, language_code, stability, similarity, style, speaker_boost, speed, use_pronunciation_dict, latency_optimization, punctuation_break_duration, long_segment_break_duration.
stt_settings model, provider, execution_mode, language_code, when known.
stt_model, tts_model Server-supplied STS model identifiers retained for live-session compatibility.
workflow_mode STS batch, realtime, near_realtime, or live_interpreter.
session_id, duration_ms STS session provenance when supplied by the server; audio row duration remains a separate field.
live_segment_count, live_tts_character_count STS live aggregation counts.
save_options Actual format, optional bitrate_kbps, and normalize.
elevenlabs_voice_settings Provider-native speed, stability, similarity_boost, style, use_speaker_boost, optimize_streaming_latency.
pronunciation_dictionary applied, mode, inline_rule_count, hosted_locator_count; does not contain dictionary text.
generation_metrics Local first-audio latency, total generation time, audio duration, real-time factor, and available phase timings/cache reuse. Durations use seconds; absent phases are omitted.
local_model On-device provenance and applicable settings, below.

Web compatibility metadata may also contain an artifact owner, filenames, saved_format, and a save timestamp for its sidecar consumers. These are not replacements for the row’s ownership or media fields. Existing combined clips retain their ordered source-clip provenance, gaps, and output details. Readers must tolerate metadata objects without a TTS settings block.

On-device model fields

local_model uses the same keys on iOS, Android, and the save endpoint:

The local save action is an explicit cloud upload of the generated audio, output text, and this metadata. Reference recordings, reference transcripts, embeddings, and device filesystem paths are not part of the metadata contract. Device-only history remains separate from Supabase. Do not infer an old result’s model revision from today’s catalog if the generation snapshot did not retain it.

Save and replacement behavior

The shared soundboard-save-generated request accepts optional durationMs and fileSizeBytes. Inline uploads use decoded audio bytes for the stored size. Unknown media dimensions remain null. Invalid/nonfinite values are dropped; duration is constrained to the database integer range. Web measures the final converted file before its temporary file is removed.

Web regeneration publishes the replacement generation metadata together with its copy-on-write audio/transcript paths under the existing compare-and-swap guard. It must not keep stale model settings from the replaced audio.

Web format conversion preserves the source generation metadata, transcript, and combined-clip provenance. It adds converted_from_clip_id, updates export/output details, and measures the converted audio for its new size and duration.

Release boundary

These are additive source changes. Deploy tts-generate and soundboard-save-generated before releasing mobile clients that send the extended local-model fields; the older endpoint strips unknown metadata. Release the web changes alongside the clients as appropriate. Existing saved clips are not backfilled, and missing historical settings cannot be reconstructed reliably. Local tests do not establish hosted Supabase persistence or physical-device save behavior.