The authoritative media index is public.soundboard_clips, scoped by user_id
and category_id. Audio and optional transcript objects live in the Soundboard
bucket. The row contains the clip ID, file name, storage and transcript paths,
transcript, MIME type, byte size and duration when known, order, offline flag,
and database timestamps. metadata is JSONB; no new database migration is needed
for the additive generation fields below.
Generation settings describe the audio that was produced. Capture them when
starting/completing generation, and retain that snapshot when the editor, model,
voice, or settings change. Save options describe the exported file. A server-owned
generation_artifacts snapshot takes precedence over client generation metadata.
Full and streaming tts-generate audio responses expose the same snapshot in
X-Speech-Clip-Metadata as percent-encoded JSON. Clients decode the header when
present and retain their captured fallback snapshot for older deployments.
Web local artifacts use their owner-checked sidecar snapshot. Never substitute
current editor settings for a completed result’s generation settings.
All keys below are optional for reading historical records. Omit unavailable or
inapplicable values rather than inventing defaults. Newly generated snapshots
include schema_version: 1, source, generation_type, provider, and
generated_at when the generation path supplies them.
| Field | Meaning |
|---|---|
source |
Originating platform or server workflow, such as ios, android, web, tts-generate, or an STS workflow. Saving must not relabel the generation source. |
generation_type |
tts or sts; existing combined clips retain their separate combined_soundboard_clip structure. |
provider |
Provider identity; local for on-device synthesis, elevenlabs for its hosted/direct output. |
generated_at |
ISO 8601 generation timestamp. Database creation time remains save time. |
tts_settings |
voice, voice_name, model, language_code, stability, similarity, style, speaker_boost, speed, use_pronunciation_dict, latency_optimization, punctuation_break_duration, long_segment_break_duration. |
stt_settings |
model, provider, execution_mode, language_code, when known. |
stt_model, tts_model |
Server-supplied STS model identifiers retained for live-session compatibility. |
workflow_mode |
STS batch, realtime, near_realtime, or live_interpreter. |
session_id, duration_ms |
STS session provenance when supplied by the server; audio row duration remains a separate field. |
live_segment_count, live_tts_character_count |
STS live aggregation counts. |
save_options |
Actual format, optional bitrate_kbps, and normalize. |
elevenlabs_voice_settings |
Provider-native speed, stability, similarity_boost, style, use_speaker_boost, optimize_streaming_latency. |
pronunciation_dictionary |
applied, mode, inline_rule_count, hosted_locator_count; does not contain dictionary text. |
generation_metrics |
Local first-audio latency, total generation time, audio duration, real-time factor, and available phase timings/cache reuse. Durations use seconds; absent phases are omitted. |
local_model |
On-device provenance and applicable settings, below. |
Web compatibility metadata may also contain an artifact owner, filenames,
saved_format, and a save timestamp for its sidecar consumers. These are not
replacements for the row’s ownership or media fields. Existing combined clips
retain their ordered source-clip provenance, gaps, and output details. Readers
must tolerate metadata objects without a TTS settings block.
local_model uses the same keys on iOS, Android, and the save endpoint:
engine, model_id, model_version, runtime, quantization.voice_id, voice_name, language_code, reference_id,
reference_audio_used.max_tokens, temperature, top_p,
top_k, repetition_penalty, seed when actually known.diffusion_steps, guidance_scale, speed.section_character_limit, paragraph_pause_ms,
join_crossfade_ms, output_speed, sample_rate_hz.The local save action is an explicit cloud upload of the generated audio, output text, and this metadata. Reference recordings, reference transcripts, embeddings, and device filesystem paths are not part of the metadata contract. Device-only history remains separate from Supabase. Do not infer an old result’s model revision from today’s catalog if the generation snapshot did not retain it.
The shared soundboard-save-generated request accepts optional durationMs and
fileSizeBytes. Inline uploads use decoded audio bytes for the stored size.
Unknown media dimensions remain null. Invalid/nonfinite values are dropped;
duration is constrained to the database integer range. Web measures the final
converted file before its temporary file is removed.
Web regeneration publishes the replacement generation metadata together with its copy-on-write audio/transcript paths under the existing compare-and-swap guard. It must not keep stale model settings from the replaced audio.
Web format conversion preserves the source generation metadata, transcript, and
combined-clip provenance. It adds converted_from_clip_id, updates export/output
details, and measures the converted audio for its new size and duration.
These are additive source changes. Deploy tts-generate and soundboard-save-generated before
releasing mobile clients that send the extended local-model fields; the older
endpoint strips unknown metadata. Release the web changes alongside the clients
as appropriate. Existing saved clips are not backfilled, and missing historical
settings cannot be reconstructed reliably. Local tests do not establish hosted
Supabase persistence or physical-device save behavior.