From d943b6ee495ce9e44934d80dde5163ec580fee62 Mon Sep 17 00:00:00 2001 From: Luke Oliff <289678208+luke-speechify@users.noreply.github.com> Date: Wed, 9 Sep 2026 22:42:33 +0100 Subject: [PATCH 1/4] docs(agents): correct SDK label to v4 in speechify-tts.md [DRG-490] The agents TTS reference labelled its TypeScript and Python snippets @speechify/api v3 / speechify-api v3, but the repo is already on v4 (catalog pins @speechify/api 4.0.1; Python recipes pin speechify-api>=4.0.0). The snippet code already uses the v4 snake_case surface, so this is a label-only fix. Keeps the ask-speechify knowledge engine from surfacing v3 as the current SDK. --- agents/speechify-tts.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/agents/speechify-tts.md b/agents/speechify-tts.md index b7c5f3e..f8ebccd 100644 --- a/agents/speechify-tts.md +++ b/agents/speechify-tts.md @@ -15,7 +15,7 @@ Both SDKs read `SPEECHIFY_API_KEY` from the environment. The deprecated ## Synthesize speech -**TypeScript** (`@speechify/api` v3) +**TypeScript** (`@speechify/api` v4) ```ts import { SpeechifyClient } from "@speechify/api"; @@ -34,7 +34,7 @@ import fs from "node:fs"; fs.writeFileSync("output.mp3", Buffer.from(response.audio_data, "base64")); ``` -**Python** (`speechify-api` v3) +**Python** (`speechify-api` v4) ```python from speechify import Speechify From f30a84210f069fefc477e4987c6092b92b1b813d Mon Sep 17 00:00:00 2001 From: Luke Oliff <289678208+luke-speechify@users.noreply.github.com> Date: Wed, 9 Sep 2026 22:45:50 +0100 Subject: [PATCH 2/4] docs(recipes): correct stale camelCase field names to snake_case v4 [DRG-490] The v3->v4 SDK migration flipped request/response fields to snake_case, but two references were missed: - quickstart README described the call params as voiceId/audioFormat - speech-marks comment referenced speechMarks.chunks The recipe code was already correct (voice_id, audio_format, response.speech_marks); only the prose/comment lagged. --- recipes/audio/typescript/sdk/quickstart/README.md | 2 +- recipes/audio/typescript/sdk/speech-marks/src/index.ts | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/recipes/audio/typescript/sdk/quickstart/README.md b/recipes/audio/typescript/sdk/quickstart/README.md index f225fde..af0ba60 100644 --- a/recipes/audio/typescript/sdk/quickstart/README.md +++ b/recipes/audio/typescript/sdk/quickstart/README.md @@ -25,6 +25,6 @@ You'll get an `output.mp3` in this folder. ## What it does - Creates a `SpeechifyClient` with your API key. -- Calls `client.audio.speech(...)` with `input`, `voiceId`, `audioFormat`, and `model` +- Calls `client.audio.speech(...)` with `input`, `voice_id`, `audio_format`, and `model` (`simba-3.2` for English, lowest latency; `simba-3.0` for multilingual — English plus German, Spanish, French, Italian, and Portuguese). - Decodes the base64 `response.audio_data` and writes it to disk. diff --git a/recipes/audio/typescript/sdk/speech-marks/src/index.ts b/recipes/audio/typescript/sdk/speech-marks/src/index.ts index 799e897..d45e4fc 100644 --- a/recipes/audio/typescript/sdk/speech-marks/src/index.ts +++ b/recipes/audio/typescript/sdk/speech-marks/src/index.ts @@ -28,7 +28,7 @@ async function main() { fs.writeFileSync("output.mp3", Buffer.from(response.audio_data, "base64")); - // `speechMarks.chunks` holds one entry per word, with start/end times in the audio. + // `speech_marks.chunks` holds one entry per word, with start/end times in the audio. const words = response.speech_marks.chunks; // Build a WebVTT file with one cue per word — the basis for karaoke-style highlighting. From 59947915031509cf4dfe933210d59a2b72434f46 Mon Sep 17 00:00:00 2001 From: Luke Oliff <289678208+luke-speechify@users.noreply.github.com> Date: Wed, 9 Sep 2026 22:56:36 +0100 Subject: [PATCH 3/4] docs: fix agent-file claims that contradict the v4 SDK [DRG-490] Adversarial review of agents/*.md against the v4 SDK type defs: - speechify-tts.md: input char limit was stated as ~20,000 for all requests, but audio.speech caps at 2,000; 20,000 is the stream limit. Split the limit per endpoint and link api-limits. - monorepo.md: products table said 'v2 SDKs'; the repo is on v4. - maintenance.md + CONTRIBUTING.md: dropped dangling references to agents/voice-agents.md, which does not exist. --- CONTRIBUTING.md | 3 +-- agents/maintenance.md | 2 +- agents/monorepo.md | 2 +- agents/speechify-tts.md | 2 +- 4 files changed, 4 insertions(+), 5 deletions(-) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index f70052c..b066dca 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -17,8 +17,7 @@ Thanks for adding to the Speechify Cookbook! The full, authoritative checklist l 3. Read the language guide: [`typescript-recipes.md`](./agents/typescript-recipes.md) or [`python-recipes.md`](./agents/python-recipes.md), and the API reference - [`speechify-tts.md`](./agents/speechify-tts.md) / - [`voice-agents.md`](./agents/voice-agents.md). + [`speechify-tts.md`](./agents/speechify-tts.md). 4. Write the recipe + a README following the fixed template. 5. Update [`README.md`](./README.md) and [`COVERAGE.md`](./COVERAGE.md). 6. Run `pnpm format`, then verify the recipe runs from a clean state. diff --git a/agents/maintenance.md b/agents/maintenance.md index 54ba028..b8b17f8 100644 --- a/agents/maintenance.md +++ b/agents/maintenance.md @@ -28,5 +28,5 @@ Keeping the cookbook consistent as it grows. The Speechify API evolves. Before trusting a method name or response field, check the installed SDK or the live docs (`https://docs.speechify.ai/llms.txt` is a good index). If -reality differs from the notes in `agents/speechify-tts.md` or `agents/voice-agents.md`, +reality differs from the notes in `agents/speechify-tts.md`, trust the API and update those files in the same change. diff --git a/agents/monorepo.md b/agents/monorepo.md index fb42510..b0ed436 100644 --- a/agents/monorepo.md +++ b/agents/monorepo.md @@ -46,7 +46,7 @@ speechify-cookbook/ | Folder | What | Status | | -------- | ------------------------------------------- | --------------------------------------- | -| `audio/` | Text-to-Speech, plus future audio products. | Active. TypeScript + Python on v2 SDKs. | +| `audio/` | Text-to-Speech, plus future audio products. | Active. TypeScript + Python on v4 SDKs. | ## Languages and tooling diff --git a/agents/speechify-tts.md b/agents/speechify-tts.md index f8ebccd..38c4327 100644 --- a/agents/speechify-tts.md +++ b/agents/speechify-tts.md @@ -62,7 +62,7 @@ with open("output.mp3", "wb") as f: | Param | Notes | | -------------- | ----------------------------------------------------------------------------------------------------------------------------------- | -| `input` | Text (or SSML) to synthesize. Up to ~20,000 characters per request. | +| `input` | Text (or SSML) to synthesize. Limits are per endpoint: 2,000 characters for `audio.speech`, 20,000 for `audio.stream` / `audio.stream_with_timestamps` (). | | `voice_id` | A voice identifier, e.g. `geffen_32`. | | `model` | `simba-3.2` (English, lowest latency) or `simba-3.0` (multilingual: English plus German, Spanish, French, Italian, and Portuguese). | | `audio_format` | `mp3`, `wav`, `ogg`, `aac`, … | From 3003a1d255f4a61e6edc44f7c53c49daaf8b0593 Mon Sep 17 00:00:00 2001 From: Luke Oliff <289678208+luke-speechify@users.noreply.github.com> Date: Wed, 9 Sep 2026 23:03:42 +0100 Subject: [PATCH 4/4] chore: prettier-format speechify-tts.md table [DRG-490] --- agents/speechify-tts.md | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/agents/speechify-tts.md b/agents/speechify-tts.md index 38c4327..a0ee7b8 100644 --- a/agents/speechify-tts.md +++ b/agents/speechify-tts.md @@ -60,12 +60,12 @@ with open("output.mp3", "wb") as f: ## Parameters -| Param | Notes | -| -------------- | ----------------------------------------------------------------------------------------------------------------------------------- | +| Param | Notes | +| -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `input` | Text (or SSML) to synthesize. Limits are per endpoint: 2,000 characters for `audio.speech`, 20,000 for `audio.stream` / `audio.stream_with_timestamps` (). | -| `voice_id` | A voice identifier, e.g. `geffen_32`. | -| `model` | `simba-3.2` (English, lowest latency) or `simba-3.0` (multilingual: English plus German, Spanish, French, Italian, and Portuguese). | -| `audio_format` | `mp3`, `wav`, `ogg`, `aac`, … | +| `voice_id` | A voice identifier, e.g. `geffen_32`. | +| `model` | `simba-3.2` (English, lowest latency) or `simba-3.0` (multilingual: English plus German, Spanish, French, Italian, and Portuguese). | +| `audio_format` | `mp3`, `wav`, `ogg`, `aac`, … | ## Capabilities to build recipes around