Reported issues for Supertone TTS
Pod holds 16 of 19 problems reported by people outside the maintainer team. Issues filed by the project's own owners, members and collaborators are excluded entirely — a maintainer's release checklist is not a warning to a prospective user.
Back to Supertone TTS.
Most discussed
merge_audio_files: crossfade_ms intermittently truncates short clips (acrossfade bug)
Symptom
crossfade_ms is accurate on normal-length clips but intermittently truncates output on short / similar-length clips. E.g. 1.36s + 1.36s, cf=500 expected ~2.22s but produces ~0.86s (≈ one clip − crossfade) on a fraction of runs. Non-deterministic.
Root cause (confirmed)
Intrinsic ffmpeg acrossfade bug on short inputs — reproduces even on raw inputs with no normalization/aresample. acrossfade's EOF/buffering occasionally drops one stream entirely. Results vary run-to-run (0.
Read the thread · 2026-06-17 · closed · 0 comments
merge_audio_files: WAV duration mis-reported (pipe:1 corrupts RIFF size header)
Symptom
merge_audio_files with WAV output reports a wildly wrong Duration (e.g. 24347.9 seconds for a real ~4.96s file). MP3 output is correct.
Root cause (confirmed)
audio_ops.merge_audio renders to pipe:1 (non-seekable). ffmpeg cannot rewind to patch the RIFF / data chunk size fields when writing WAV to a pipe, so it writes the placeholder 0xFFFFFFFF (4294967295).
tools.calculate_duration (mutagen) then computes 4294967295 / (44100*2*2) ≈ 24347.887s — exactly the repo
Read the thread · 2026-06-17 · closed · 0 comments
[ISSUE-031] Harden _pipe_format to reject unknown output formats (L1)
Goal
Make audio_ops._pipe_format fail loudly on an unknown output format instead of silently mapping anything non-wav to mp3.
Background (review finding L1, from PR #49)
_pipe_format returns "wav" if output_format == "wav" else "mp3". Today output_format is enum-validated upstream (SUPPORTED_FORMATS), so there's no current bug. But MERGE_SUPPORTED_EXTENSIONS exists specifically "so the merge surface can diverge later" — if a third format is added there, _pipe_format woul
Read the thread · 2026-06-16 · open · 0 comments
[ISSUE-030] Real-ffmpeg integration test tier for merge_audio_files (RL-007)
Goal
Add an opt-in (CI-skipped) real-ffmpeg integration test tier for merge_audio_files, so the filter-graph / parameter-compatibility contract is pinned by a test that actually runs the bundled ffmpeg — not just the mocked unit tests.
Background (review lesson RL-007, from PR #49)
tests/test_audio_ops.py mocks asyncio.create_subprocess_exec, so it validates the command string but never ffmpeg's actual behavior. The C1/C2 filter-graph bug in PR #49 (concat/aevalsrc parameter mis
Read the thread · 2026-06-16 · open · 0 comments
[ISSUE-029] Add merge_audio_files tool (ffmpeg-backed audio concatenation)
Goal
Add a merge_audio_files MCP tool backed by a bundled ffmpeg binary that concatenates two or more audio files produced by text_to_speech into a single deliverable, supporting head-to-tail concat, silence-gap insertion, and crossfade blending.
Spec: docs/specs/SPEC-029.md (Option A — bundle ffmpeg via imageio-ffmpeg).
Scope
In: new src/supertone_mcp/audio_ops.py (async ffmpeg subprocess); merge_audio_files handler in tools.py; registration in server.py; `imageio-ffmpeg
Read the thread · 2026-06-16 · closed · 0 comments
[ISSUE-027] New usage tools get_usage_history + get_voice_usage
Goal
Two new MCP tools expose usage data:
get_usage_history()wrapsusage.get_usage_async(advanced analytics; time-bucketed).get_voice_usage(voice_id)reports usage for a single voice.
Each returns a formatted plain-text usage summary.
Scope (In)
- New
SupertoneClient.get_usage_history(...)andSupertoneClient.get_voice_usage(...)methods. - Two tool handlers + formatters in tools.py.
- server.py registration for both.
- Empty-voice_id validation for get_voice_usage.
- Te
Read the thread · 2026-06-05 · closed · 0 comments
[ISSUE-026] New tool get_custom_voice
Closes via PR for ISSUE-026.
Goal
A new get_custom_voice(voice_id) tool returns the detail of a single custom (cloned) voice by wrapping custom_voices.get_custom_voice_async, with empty-voice_id validation.
Scope
In: New async get_custom_voice(voice_id) method on SupertoneClient wrapping custom_voices.get_custom_voice_async; new async get_custom_voice(voice_id) -> str tool handler + format_custom_voice_detail(...) formatter in tools.py; server.py registration; empty voice_id
Read the thread · 2026-06-05 · closed · 0 comments
[ISSUE-025] Expose include_phonemes + normalized_text TTS params
Goal
text_to_speech exposes the SDK 0.2.3 include_phonemes (bool, default false) and normalized_text (str, optional) parameters, passing them through to the client/SDK, with normalized_text documented as effective only for sona_speech_2/sona_speech_2_flash.
Scope
- In: Add
include_phonemes: bool = Falseandnormalized_text: str | None = Nonetotools.text_to_speech,client.synthesizeandsynthesize_stream, andserver.pyregistration/docstrings; pass both through to
Read the thread · 2026-06-05 · closed · 0 comments
Most recent
[ISSUE-024] Relax 300-char hard limit → delegate to SDK auto-chunk
Goal
text_to_speech and predict_duration no longer hard-reject text over 300 characters; long text is delegated to the SDK's internal auto-chunking, while empty-text validation is retained.
Scope
In: Remove the validate_text_max_length hard rejection from predict_duration (confirm text_to_speech does not call it); keep validate_text (non-empty); update text_to_speech and predict_duration tool descriptions/docstrings to note long text is auto-chunked and credit/latency
Read the thread · 2026-06-05 · closed · 0 comments
[ISSUE-023] Add streaming param + route synthesize vs stream + sona_speech_1-only validation
Closes-tracking for ISSUE-023.
Goal
text_to_speech accepts a per-call streaming parameter (default false) that routes to one-shot client.synthesize by default and to client.synthesize_stream when true, with fail-fast validation rejecting streaming=true for any model other than sona_speech_1 before any SDK call.
Scope
In: add streaming: bool = False to tools.text_to_speech + server.py registration/docstring; wire the one-shot client.synthesize() path through existing outpu
Read the thread · 2026-06-05 · closed · 0 comments
[ISSUE-022] Remove behavior env vars → per-call output_mode/autoplay (BREAKING)
Goal
text_to_speech decides output mode and autoplay per call via new output_mode and autoplay parameters; the env vars SUPERTONE_MCP_OUTPUT_MODE and SUPERTONE_MCP_AUTOPLAY are no longer read; autoplay defaults to false. BREAKING.
Scope
In: Add output_mode (str, default files) and autoplay (bool, default false) params to tools.text_to_speech and server.py registration + docstrings; replace resolve_output_mode() env read with validate_output_mode(mode) against
Read the thread · 2026-06-05 · closed · 0 comments
[ISSUE-021] SDK 0.2.3 sync: model enum + default + version pin
Goal
constants.SUPPORTED_MODELS matches the SDK 0.2.3 model enum (all 7 models), the default model is sona_speech_2_flash, and the supertone dependency is pinned to >=0.2.3,<0.3.
Scope
In: Add sona_speech_3t and supertonic_api_3 to SUPPORTED_MODELS and the Model Literal in constants.py; change DEFAULT_MODEL from sona_speech_1 to sona_speech_2_flash; pin supertone>=0.2.3,<0.3 in pyproject.toml; update validate_model tests; correct the stale text_to_speech docstring in server.py. Out: strea
Read the thread · 2026-06-05 · closed · 0 comments
ISSUE-020: Custom voice CRUD tools (search/edit/delete)
Tracking issue for ISSUE-020 (FR-018, FR-019, US-011). See issues.md ISSUE-020 for full spec. Implemented in PR #23.
Read the thread · 2026-05-27 · closed · 0 comments
[ISSUE-019] Add clone_voice tool (single file ≤3MB)
Goal
clone_voice(name, audio_path, description?) creates a custom (cloned) voice from a local audio file (WAV or MP3, ≤3MB) and returns the new custom voice ID.
Scope
- In: New
async create_cloned_voice(name, audio_bytes, file_name, content_type, description?) -> dictmethod onSupertoneClientwrappingcustom_voices.create_cloned_voice_async; newasync clone_voice(name, audio_path, description?) -> strtool handler that reads the file, validates extension/size, builds the SDK `F
Read the thread · 2026-05-27 · closed · 0 comments
[ISSUE-018] Add predict_duration tool (client + handler)
Closes: tracked in issues.md ISSUE-018
Goal
predict_duration(text, voice_id?, language?, model?, output_format?, speed?, pitch_shift?, style?) calls the Supertone duration prediction API and returns the predicted audio length in seconds, which is proportional to credit consumption.
Scope (In/Out)
In:
- New
async predict_duration(...)method onSupertoneClientwrappingtext_to_speech.predict_duration_async - New
async predict_duration(...)tool handler intools.py - Vali
Read the thread · 2026-05-27 · closed · 0 comments
[ISSUE-017] Add preview_voice tool (returns sample URLs)
Goal
`preview_voice(voice_id, language?, style?, model?)` returns matching sample audio URLs from the voice's `samples` array so users (or the client) can listen to a preview before choosing a voice.
Scope
In:
- `async preview_voice(voice_id, language?, style?, model?) -> str` handler in `tools.py`
- New `format_voice_samples(samples, filters) -> str` formatter
- Server registration in `server.py`
- Tests in `tests/test_tools.py` and `tests/test_server.py`
Out:
Read the thread · 2026-05-27 · closed · 0 comments
The remaining reports are on the project's issue tracker.