diff --git a/recipes/audio/bash/native/voice-cloning/.env.example b/recipes/audio/bash/native/voice-cloning/.env.example index 70a94b2..cdea4b1 100644 --- a/recipes/audio/bash/native/voice-cloning/.env.example +++ b/recipes/audio/bash/native/voice-cloning/.env.example @@ -1,2 +1,11 @@ # Get a key at https://platform.speechify.ai/api-keys +# Voice cloning must be enabled on your plan. SPEECHIFY_API_KEY= + +# Full name of the person consenting to have their voice cloned. +CONSENT_FULL_NAME=Jane Doe + +# Optional — override where the recipe reads the audio from. +# Defaults to sample.wav and consent.wav in the recipe folder. +# SAMPLE_PATH=./sample.wav +# CONSENT_RECORDING_PATH=./consent.wav diff --git a/recipes/audio/bash/native/voice-cloning/README.md b/recipes/audio/bash/native/voice-cloning/README.md index bd6a429..b6bac54 100644 --- a/recipes/audio/bash/native/voice-cloning/README.md +++ b/recipes/audio/bash/native/voice-cloning/README.md @@ -1,7 +1,8 @@ # Text-to-Speech: voice cloning (Bash, native REST) -Clone a voice from an audio sample, synthesize speech with the clone, then delete it -— same lifecycle as the [TypeScript](../../../typescript/native/voice-cloning) and +Clone a voice from an audio sample, synthesize speech with the clone, then delete it, with +**verified consent** — same lifecycle as the +[TypeScript](../../../typescript/native/voice-cloning) and [Python](../../../python/native/voice-cloning) native recipes, but as a self-contained shell script using `curl` + `jq`. @@ -12,6 +13,9 @@ shell script using `curl` + `jq`. pointing to [Speechify pricing](https://speechify.ai/pricing) (the API returns `402 voice_cloning_not_included`). - `bash`, `curl`, `jq`, and `base64` +- **Two recordings of the same consenting person** (see [Consent](#consent)): + - a voice **sample** to clone — 10–30s of clean speech + - a **consent recording** — that person reading the challenge phrase this script prints ## Setup @@ -22,28 +26,43 @@ chmod +x voice-cloning.sh ## Run +Because consent is verified against a phrase the API generates, this is a two-step run: + ```bash -./voice-cloning.sh +./voice-cloning.sh # 1st run: prints the phrase to read, then exits +# record the speaker reading that phrase → consent.wav, and their voice → sample.wav +./voice-cloning.sh # 2nd run: clones, synthesizes, deletes ``` -Produces an `output.mp3` spoken in the cloned voice, then removes the cloned voice. +Put `sample.wav` and `consent.wav` next to the script, or point `SAMPLE_PATH` / +`CONSENT_RECORDING_PATH` at them. Produces an `output.mp3` in the cloned voice, then +removes the cloned voice. ## What it does +- `POST /v1/voices/consent-challenges` (JSON `{ full_name }`, built with `jq -n`) — returns + a `phrase` the speaker must read aloud and an `id`; cached in `.consent-challenge.json` + between runs (single-use). - `POST /v1/voices` (**multipart/form-data**) via `curl -F`: fields `name`, `gender`, - `consent`, and a `sample=@fixtures/spacewalk.wav` file part. `curl` sets the - `Content-Type` boundary automatically. -- `POST /v1/audio/speech` with the returned `voice_id`. `jq -n` builds the JSON body - safely. -- `DELETE /v1/voices/{id}` runs from an `EXIT` trap, so the cloned voice is removed - even if step 2 fails. + `consent_challenge_id`, and `sample=@…` + `consent_recording=@…` file parts. `curl` sets + the `Content-Type` boundary automatically. +- `POST /v1/audio/speech` with the returned `voice_id`. +- `DELETE /v1/voices/{id}` runs from an `EXIT` trap, so the cloned voice is removed even if + a later step fails. + +All calls pin `Speechify-Version: 2026-09-13`, the API version the verified-consent flow +ships on. ## Consent -`consent` is a **required** JSON string attesting you have the speaker's permission to -clone their voice (`{"fullName": "...", "email": "..."}`). Only clone voices you are -authorized to. The bundled `fixtures/spacewalk.wav` is ~26s of NASA ISS spacewalk audio -(U.S. government work, public domain); replace it with your own consented sample for -real use. +Cloning requires **verified consent**. You create a consent challenge, the speaker records +themselves reading the returned `phrase`, and that recording is sent as `consent_recording` +and retained as the consent record. The `consent_recording` **must be the same person** as +the voice `sample` — so there is no shippable sample that comes with valid consent, and this +recipe is deliberately bring-your-own-audio. Only clone voices you are authorized to. -> Voice cloning reference: https://docs.speechify.ai/tts/guides/voice-cloning +> The old `consent` JSON field (`fullName` + `email`) is removed on `Speechify-Version: +> 2026-09-13` — see the +> [migration guide](https://docs.speechify.ai/build/guides/deprecations/migrating-voice-cloning-consent). +> +> Voice cloning reference: https://docs.speechify.ai/build/guides/voice-cloning/overview diff --git a/recipes/audio/bash/native/voice-cloning/fixtures/spacewalk.wav b/recipes/audio/bash/native/voice-cloning/fixtures/spacewalk.wav deleted file mode 100644 index 498be74..0000000 Binary files a/recipes/audio/bash/native/voice-cloning/fixtures/spacewalk.wav and /dev/null differ diff --git a/recipes/audio/bash/native/voice-cloning/voice-cloning.sh b/recipes/audio/bash/native/voice-cloning/voice-cloning.sh index 78728b3..4f931d5 100755 --- a/recipes/audio/bash/native/voice-cloning/voice-cloning.sh +++ b/recipes/audio/bash/native/voice-cloning/voice-cloning.sh @@ -2,7 +2,7 @@ set -euo pipefail # Speechify TTS voice cloning (Bash + curl + jq). -# Full lifecycle: clone → use → delete. +# Full lifecycle: clone → use → delete, with verified consent. cd "$(dirname "$0")" @@ -16,24 +16,64 @@ fi : "${SPEECHIFY_API_KEY:?Set SPEECHIFY_API_KEY (copy .env.example to .env).}" BASE="https://api.speechify.ai" -# Bundled sample: ~26s of NASA ISS spacewalk audio (public domain). -SAMPLE="fixtures/spacewalk.wav" - -# 1. Clone a voice from an audio sample (10-30s of clean speech works well). -# POST /v1/voices is multipart/form-data — `curl -F` builds the body and sets -# the Content-Type boundary automatically. `consent` is REQUIRED: a JSON -# string attesting you have the speaker's permission to clone their voice. -http_status=0 +# The verified-consent flow ships on this API version — pin it explicitly. +VERSION="2026-09-13" + +# Cloning requires VERIFIED consent: the speaker records themselves reading a +# phrase the API returns, and that recording is kept as the consent record. It +# must be the SAME person as the voice sample, so this recipe is +# bring-your-own-audio — there is no sample that ships with valid consent. +CONSENT_FULL_NAME="${CONSENT_FULL_NAME:-Jane Doe}" +SAMPLE="${SAMPLE_PATH:-sample.wav}" +CONSENT="${CONSENT_RECORDING_PATH:-consent.wav}" +# The challenge is single-use and its phrase is dynamic, so we cache it between +# runs: run once to get the phrase, record it, run again to submit. +CHALLENGE_CACHE=".consent-challenge.json" + +auth=(-H "Authorization: Bearer ${SPEECHIFY_API_KEY}" -H "Speechify-Version: ${VERSION}") + +# 1. Get (or reuse) a consent challenge. Its `phrase` is what the speaker must +# read aloud; `id` ties the recording to this consent on the create. +if [ ! -f "$CHALLENGE_CACHE" ]; then + curl --fail-with-body --silent --show-error \ + -X POST "${BASE}/v1/voices/consent-challenges" \ + "${auth[@]}" \ + -H "Content-Type: application/json" \ + -d "$(jq -n --arg n "$CONSENT_FULL_NAME" '{full_name: $n}')" \ + > "$CHALLENGE_CACHE" +fi +challenge_id=$(jq -r '.id' < "$CHALLENGE_CACHE") +phrase=$(jq -r '.phrase' < "$CHALLENGE_CACHE") +expires_at=$(jq -r '.expires_at' < "$CHALLENGE_CACHE") + +# 2. Make sure we have both recordings before spending the (single-use) challenge. +if [ ! -f "$SAMPLE" ] || [ ! -f "$CONSENT" ]; then + echo "" + echo "Consent required. Have ${CONSENT_FULL_NAME} record themselves reading this phrase, exactly as written:" + echo "" + echo " \"${phrase}\"" + echo "" + echo "Then provide these files and re-run (paths override via SAMPLE_PATH / CONSENT_RECORDING_PATH):" + [ ! -f "$SAMPLE" ] && echo " sample: ${SAMPLE} (${CONSENT_FULL_NAME}'s voice, 10-30s of clean speech)" + [ ! -f "$CONSENT" ] && echo " consent: ${CONSENT} (the SAME person reading the phrase above)" + echo "" + echo "Challenge expires ${expires_at}. If it has expired, delete ${CHALLENGE_CACHE} and re-run." + exit 1 +fi + +# 3. Clone the voice. POST /v1/voices is multipart/form-data — `curl -F` builds +# the body and sets the Content-Type boundary automatically. create_body=$(mktemp) trap 'rm -f "$create_body"' EXIT http_status=$(curl --silent --show-error --output "$create_body" --write-out '%{http_code}' \ -X POST "${BASE}/v1/voices" \ - -H "Authorization: Bearer ${SPEECHIFY_API_KEY}" \ + "${auth[@]}" \ -F "name=cookbook-cloned-voice" \ -F "gender=male" \ - -F 'consent={"fullName":"Jane Doe","email":"jane@example.com"};type=text/plain' \ - -F "sample=@${SAMPLE};type=audio/wav") + -F "consent_challenge_id=${challenge_id}" \ + -F "sample=@${SAMPLE}" \ + -F "consent_recording=@${CONSENT}") if [ "$http_status" = "402" ]; then echo "" @@ -45,19 +85,22 @@ if [ "$http_status" -lt 200 ] || [ "$http_status" -ge 300 ]; then echo "POST /v1/voices → ${http_status}" >&2 cat "$create_body" >&2 echo >&2 + echo "If this is a consent_* error, the challenge may be spent or expired — delete ${CHALLENGE_CACHE} and re-run." >&2 exit 1 fi +rm -f "$CHALLENGE_CACHE" # challenge is spent + voice_id=$(jq -r '.id' < "$create_body") display_name=$(jq -r '.display_name' < "$create_body") voice_type=$(jq -r '.type' < "$create_body") echo "Cloned voice created: ${voice_id} (${display_name}, type=${voice_type})" -# Ensure we always delete the cloned voice, even on failure of step 2. +# Ensure we always delete the cloned voice, even on failure of step 4. cleanup() { del_status=$(curl --silent --show-error --output /dev/null --write-out '%{http_code}' \ -X DELETE "${BASE}/v1/voices/${voice_id}" \ - -H "Authorization: Bearer ${SPEECHIFY_API_KEY}") + "${auth[@]}") if [ "$del_status" -ge 200 ] && [ "$del_status" -lt 300 ]; then echo "Deleted cloned voice ${voice_id}" else @@ -66,10 +109,10 @@ cleanup() { } trap 'cleanup; rm -f "$create_body"' EXIT -# 2. Synthesize speech using the cloned voice — pass its id as voice_id. +# 4. Synthesize speech using the cloned voice — pass its id as voice_id. speech_response=$(curl --fail-with-body --silent --show-error \ -X POST "${BASE}/v1/audio/speech" \ - -H "Authorization: Bearer ${SPEECHIFY_API_KEY}" \ + "${auth[@]}" \ -H "Content-Type: application/json" \ -d "$(jq -n --arg vid "$voice_id" '{ input: "Hello from a voice cloned with the Speechify API.", @@ -81,4 +124,4 @@ speech_response=$(curl --fail-with-body --silent --show-error \ printf '%s' "$speech_response" | jq -r '.audio_data' | base64 -d > output.mp3 echo "Wrote output.mp3" -# 3. Cleanup runs from the EXIT trap. +# 5. Cleanup runs from the EXIT trap. diff --git a/recipes/audio/python/native/voice-cloning/.env.example b/recipes/audio/python/native/voice-cloning/.env.example index 4cb51e4..cdea4b1 100644 --- a/recipes/audio/python/native/voice-cloning/.env.example +++ b/recipes/audio/python/native/voice-cloning/.env.example @@ -1 +1,11 @@ +# Get a key at https://platform.speechify.ai/api-keys +# Voice cloning must be enabled on your plan. SPEECHIFY_API_KEY= + +# Full name of the person consenting to have their voice cloned. +CONSENT_FULL_NAME=Jane Doe + +# Optional — override where the recipe reads the audio from. +# Defaults to sample.wav and consent.wav in the recipe folder. +# SAMPLE_PATH=./sample.wav +# CONSENT_RECORDING_PATH=./consent.wav diff --git a/recipes/audio/python/native/voice-cloning/README.md b/recipes/audio/python/native/voice-cloning/README.md index 862c7dd..493f00b 100644 --- a/recipes/audio/python/native/voice-cloning/README.md +++ b/recipes/audio/python/native/voice-cloning/README.md @@ -1,8 +1,8 @@ # Text-to-Speech: voice cloning (Python, native REST) -The same as [`voice-cloning`](../../sdk/voice-cloning) — clone → synthesize → delete — -but calling the REST API directly with `requests` + multipart form-data instead of the -`speechify-api` SDK. +The same as [`voice-cloning`](../../sdk/voice-cloning) — clone → synthesize → delete, with +**verified consent** — but calling the REST API directly with `requests` + multipart +form-data instead of the `speechify-api` SDK. ## Prerequisites @@ -11,6 +11,9 @@ but calling the REST API directly with `requests` + multipart form-data instead pointing to [Speechify pricing](https://speechify.ai/pricing) (the API returns `402 voice_cloning_not_included`). - Python 3.10+ and [uv](https://docs.astral.sh/uv/) +- **Two recordings of the same consenting person** (see [Consent](#consent)): + - a voice **sample** to clone — 10–30s of clean speech + - a **consent recording** — that person reading the challenge phrase this recipe prints ## Setup @@ -21,26 +24,43 @@ uv sync ## Run +Because consent is verified against a phrase the API generates, this is a two-step run: + ```bash -uv run main.py +uv run main.py # 1st run: prints the phrase to read, then exits +# record the speaker reading that phrase → consent.wav, and their voice → sample.wav +uv run main.py # 2nd run: clones, synthesizes, deletes ``` -Produces an `output.mp3` spoken in the cloned voice, then removes the cloned voice. +Put `sample.wav` and `consent.wav` in the recipe folder, or point `SAMPLE_PATH` / +`CONSENT_RECORDING_PATH` at them. Produces an `output.mp3` in the cloned voice, then +removes the cloned voice. ## What it does -- `POST /v1/voices` (**multipart/form-data**): fields `name`, `gender`, `consent`, and - a `sample` file part. Pass them via `data=` + `files=` and `requests` sets the - `Content-Type` boundary automatically — do **not** set it yourself. +- `POST /v1/voices/consent-challenges` (JSON `{ full_name }`) — returns a `phrase` the + speaker must read aloud and an `id`; cached in `.consent-challenge.json` between runs + (single-use). +- `POST /v1/voices` (**multipart/form-data**): fields `name`, `gender`, + `consent_challenge_id`, and `sample` + `consent_recording` file parts. Pass them via + `data=` + `files=` and `requests` sets the `Content-Type` boundary — do **not** set it + yourself. - `POST /v1/audio/speech` with the returned voice's `id` as `voice_id`. - `DELETE /v1/voices/{id}` — cleans up so personal voices don't accumulate. +All calls pin `Speechify-Version: 2026-09-13`, the API version the verified-consent flow +ships on. + ## Consent -`consent` is a **required** JSON string attesting you have the speaker's permission to -clone their voice (`{"fullName": "...", "email": "..."}`). Only clone voices you are -authorized to. The bundled `fixtures/spacewalk.wav` is ~26s of NASA ISS spacewalk audio -(U.S. government work, public domain); replace it with your own consented sample for -real use. +Cloning requires **verified consent**. You create a consent challenge, the speaker records +themselves reading the returned `phrase`, and that recording is sent as `consent_recording` +and retained as the consent record. The `consent_recording` **must be the same person** as +the voice `sample` — so there is no shippable sample that comes with valid consent, and this +recipe is deliberately bring-your-own-audio. Only clone voices you are authorized to. -> Voice cloning reference: https://docs.speechify.ai/tts/guides/voice-cloning +> The old `consent` JSON field (`fullName` + `email`) is removed on `Speechify-Version: +> 2026-09-13` — see the +> [migration guide](https://docs.speechify.ai/build/guides/deprecations/migrating-voice-cloning-consent). +> +> Voice cloning reference: https://docs.speechify.ai/build/guides/voice-cloning/overview diff --git a/recipes/audio/python/native/voice-cloning/fixtures/spacewalk.wav b/recipes/audio/python/native/voice-cloning/fixtures/spacewalk.wav deleted file mode 100644 index 498be74..0000000 Binary files a/recipes/audio/python/native/voice-cloning/fixtures/spacewalk.wav and /dev/null differ diff --git a/recipes/audio/python/native/voice-cloning/main.py b/recipes/audio/python/native/voice-cloning/main.py index 201cd70..19b4163 100644 --- a/recipes/audio/python/native/voice-cloning/main.py +++ b/recipes/audio/python/native/voice-cloning/main.py @@ -1,6 +1,7 @@ import base64 import json import os +from datetime import datetime, timezone import requests from dotenv import load_dotenv @@ -10,9 +11,31 @@ # form-data instead of the speechify-api SDK. BASE = "https://api.speechify.ai" +# The verified-consent flow ships on this API version — pin it explicitly. +VERSION = "2026-09-13" -# Bundled sample: ~26s of NASA ISS spacewalk audio (public domain). -SAMPLE_PATH = os.path.join(os.path.dirname(__file__), "fixtures", "spacewalk.wav") +# Cloning requires VERIFIED consent: the speaker records themselves reading a +# phrase the API returns, and that recording is kept as the consent record. It +# must be the SAME person as the voice sample, so this recipe is +# bring-your-own-audio — there is no sample that ships with valid consent. +HERE = os.path.dirname(__file__) +CONSENT_FULL_NAME = os.environ.get("CONSENT_FULL_NAME", "Jane Doe") +SAMPLE_PATH = os.environ.get("SAMPLE_PATH", os.path.join(HERE, "sample.wav")) +CONSENT_RECORDING_PATH = os.environ.get("CONSENT_RECORDING_PATH", os.path.join(HERE, "consent.wav")) +# The challenge is single-use and its phrase is dynamic, so we cache it between +# runs: run once to get the phrase, record it, run again to submit. +CHALLENGE_CACHE = os.path.join(HERE, ".consent-challenge.json") + + +def load_challenge(): + if not os.path.exists(CHALLENGE_CACHE): + return None + with open(CHALLENGE_CACHE) as f: + c = json.load(f) + # Python 3.10's fromisoformat doesn't accept a trailing "Z". + if datetime.fromisoformat(c["expires_at"].replace("Z", "+00:00")) <= datetime.now(timezone.utc): + return None # expired -> make a fresh one + return c def main() -> None: @@ -22,22 +45,54 @@ def main() -> None: if not token: raise SystemExit("Set SPEECHIFY_API_KEY (copy .env.example to .env).") - auth = {"Authorization": f"Bearer {token}"} + auth = {"Authorization": f"Bearer {token}", "Speechify-Version": VERSION} + + # 1. Get (or reuse) a consent challenge. Its `phrase` is what the speaker must + # read aloud; `id` ties the recording to this consent on the create. + challenge = load_challenge() + if challenge is None: + resp = requests.post( + f"{BASE}/v1/voices/consent-challenges", + headers={**auth, "Content-Type": "application/json"}, + json={"full_name": CONSENT_FULL_NAME}, + timeout=30, + ) + resp.raise_for_status() + challenge = resp.json() + with open(CHALLENGE_CACHE, "w") as f: + json.dump(challenge, f, indent=2) - # 1. Clone a voice from an audio sample (10-30s of clean speech works well). - # POST /v1/voices is multipart/form-data — pass `files=` to requests and it - # sets the Content-Type boundary automatically. `consent` is REQUIRED: a JSON - # string attesting you have the speaker's permission to clone their voice. - with open(SAMPLE_PATH, "rb") as sample: + # 2. Make sure we have both recordings before spending the (single-use) challenge. + missing = [] + if not os.path.exists(SAMPLE_PATH): + missing.append(f" sample: {SAMPLE_PATH} ({CONSENT_FULL_NAME}'s voice, 10-30s of clean speech)") + if not os.path.exists(CONSENT_RECORDING_PATH): + missing.append(f" consent: {CONSENT_RECORDING_PATH} (the SAME person reading the phrase below)") + if missing: + raise SystemExit( + f"\nConsent required. Have {CONSENT_FULL_NAME} record themselves reading this phrase, " + "exactly as written:\n\n" + f' "{challenge["phrase"]}"\n\n' + "Then provide these files and re-run (paths override via SAMPLE_PATH / CONSENT_RECORDING_PATH):\n" + + "\n".join(missing) + + f"\n\nChallenge expires {challenge['expires_at']}.\n" + ) + + # 3. Clone the voice. POST /v1/voices is multipart/form-data — pass `files=` to + # requests and it sets the Content-Type boundary automatically. + with open(SAMPLE_PATH, "rb") as sample, open(CONSENT_RECORDING_PATH, "rb") as consent: create_resp = requests.post( f"{BASE}/v1/voices", headers=auth, data={ "name": "cookbook-cloned-voice", "gender": "male", - "consent": json.dumps({"fullName": "Jane Doe", "email": "jane@example.com"}), + "consent_challenge_id": challenge["id"], + }, + files={ + "sample": (os.path.basename(SAMPLE_PATH), sample), + "consent_recording": (os.path.basename(CONSENT_RECORDING_PATH), consent), }, - files={"sample": ("spacewalk.wav", sample, "audio/wav")}, timeout=120, ) @@ -47,11 +102,12 @@ def main() -> None: "Upgrade to a plan that includes voice cloning: https://speechify.ai/pricing\n" ) create_resp.raise_for_status() + os.remove(CHALLENGE_CACHE) # challenge is spent voice = create_resp.json() print(f"Cloned voice created: {voice['id']} ({voice['display_name']}, type={voice['type']})") try: - # 2. Synthesize speech using the cloned voice — pass its id as voice_id. + # 4. Synthesize speech using the cloned voice — pass its id as voice_id. speech_resp = requests.post( f"{BASE}/v1/audio/speech", headers={**auth, "Content-Type": "application/json"}, @@ -69,7 +125,7 @@ def main() -> None: f.write(base64.b64decode(speech["audio_data"])) print("Wrote output.mp3") finally: - # 3. Clean up so cloned voices don't accumulate on your account. + # 5. Clean up so cloned voices don't accumulate on your account. # Remove this to keep the voice and reuse it later via voice.id. del_resp = requests.delete(f"{BASE}/v1/voices/{voice['id']}", headers=auth, timeout=30) if del_resp.ok: diff --git a/recipes/audio/typescript/native/voice-cloning/.env.example b/recipes/audio/typescript/native/voice-cloning/.env.example index 70a94b2..cdea4b1 100644 --- a/recipes/audio/typescript/native/voice-cloning/.env.example +++ b/recipes/audio/typescript/native/voice-cloning/.env.example @@ -1,2 +1,11 @@ # Get a key at https://platform.speechify.ai/api-keys +# Voice cloning must be enabled on your plan. SPEECHIFY_API_KEY= + +# Full name of the person consenting to have their voice cloned. +CONSENT_FULL_NAME=Jane Doe + +# Optional — override where the recipe reads the audio from. +# Defaults to sample.wav and consent.wav in the recipe folder. +# SAMPLE_PATH=./sample.wav +# CONSENT_RECORDING_PATH=./consent.wav diff --git a/recipes/audio/typescript/native/voice-cloning/README.md b/recipes/audio/typescript/native/voice-cloning/README.md index 02bb8fe..f46ce1f 100644 --- a/recipes/audio/typescript/native/voice-cloning/README.md +++ b/recipes/audio/typescript/native/voice-cloning/README.md @@ -1,8 +1,8 @@ # Text-to-Speech: voice cloning (TypeScript, native REST) -The same as [`voice-cloning`](../../sdk/voice-cloning) — clone → synthesize → delete — -but calling the REST API directly with `fetch` + multipart `FormData` instead of the -`@speechify/api` SDK. +The same as [`voice-cloning`](../../sdk/voice-cloning) — clone → synthesize → delete, with +**verified consent** — but calling the REST API directly with `fetch` + multipart +`FormData` instead of the `@speechify/api` SDK. ## Prerequisites @@ -11,6 +11,9 @@ but calling the REST API directly with `fetch` + multipart `FormData` instead of pointing to [Speechify pricing](https://speechify.ai/pricing) (the API returns `402 voice_cloning_not_included`). - Node 20+ (for built-in `fetch`, `FormData`, and `Blob`) +- **Two recordings of the same consenting person** (see [Consent](#consent)): + - a voice **sample** to clone — 10–30s of clean speech + - a **consent recording** — that person reading the challenge phrase this recipe prints ## Setup @@ -21,26 +24,43 @@ pnpm install ## Run +Because consent is verified against a phrase the API generates, this is a two-step run: + ```bash -pnpm start +pnpm start # 1st run: prints the phrase to read, then exits +# record the speaker reading that phrase → consent.wav, and their voice → sample.wav +pnpm start # 2nd run: clones, synthesizes, deletes ``` -Produces an `output.mp3` spoken in the cloned voice, then removes the cloned voice. +Put `sample.wav` and `consent.wav` in the recipe folder, or point `SAMPLE_PATH` / +`CONSENT_RECORDING_PATH` at them. Produces an `output.mp3` in the cloned voice, then +removes the cloned voice. ## What it does -- `POST /v1/voices` (**multipart/form-data**): fields `name`, `gender`, `consent`, and a - `sample` file part. Pass a `FormData` instance as `body` and let `fetch` set the - `Content-Type` boundary automatically — do **not** set it yourself. +- `POST /v1/voices/consent-challenges` (JSON `{ full_name }`) — returns a `phrase` the + speaker must read aloud and an `id`; cached in `.consent-challenge.json` between runs + (single-use). +- `POST /v1/voices` (**multipart/form-data**): fields `name`, `gender`, + `consent_challenge_id`, and `sample` + `consent_recording` file parts. Pass a + `FormData` instance as `body` and let `fetch` set the `Content-Type` boundary — do + **not** set it yourself. - `POST /v1/audio/speech` with the returned voice's `id` as `voice_id`. - `DELETE /v1/voices/{id}` — cleans up so personal voices don't accumulate. +All calls pin `Speechify-Version: 2026-09-13`, the API version the verified-consent flow +ships on. + ## Consent -`consent` is a **required** JSON string attesting you have the speaker's permission to -clone their voice (`{"fullName": "...", "email": "..."}`). Only clone voices you are -authorized to. The bundled `fixtures/spacewalk.wav` is ~26s of NASA ISS spacewalk audio -(U.S. government work, public domain); replace it with your own consented sample for -real use. +Cloning requires **verified consent**. You create a consent challenge, the speaker records +themselves reading the returned `phrase`, and that recording is sent as `consent_recording` +and retained as the consent record. The `consent_recording` **must be the same person** as +the voice `sample` — so there is no shippable sample that comes with valid consent, and this +recipe is deliberately bring-your-own-audio. Only clone voices you are authorized to. -> Voice cloning reference: https://docs.speechify.ai/tts/guides/voice-cloning +> The old `consent` JSON field (`fullName` + `email`) is removed on `Speechify-Version: +> 2026-09-13` — see the +> [migration guide](https://docs.speechify.ai/build/guides/deprecations/migrating-voice-cloning-consent). +> +> Voice cloning reference: https://docs.speechify.ai/build/guides/voice-cloning/overview diff --git a/recipes/audio/typescript/native/voice-cloning/fixtures/spacewalk.wav b/recipes/audio/typescript/native/voice-cloning/fixtures/spacewalk.wav deleted file mode 100644 index 498be74..0000000 Binary files a/recipes/audio/typescript/native/voice-cloning/fixtures/spacewalk.wav and /dev/null differ diff --git a/recipes/audio/typescript/native/voice-cloning/src/index.ts b/recipes/audio/typescript/native/voice-cloning/src/index.ts index 8513250..91e7f02 100644 --- a/recipes/audio/typescript/native/voice-cloning/src/index.ts +++ b/recipes/audio/typescript/native/voice-cloning/src/index.ts @@ -12,10 +12,27 @@ if (!token) { } const BASE = "https://api.speechify.ai"; +// The verified-consent flow ships on this API version — pin it explicitly. +const VERSION = "2026-09-13"; +const authHeaders = { Authorization: `Bearer ${token}`, "Speechify-Version": VERSION }; -// Bundled sample: ~26s of NASA ISS spacewalk audio (public domain). -const samplePath = path.resolve(import.meta.dirname, "../fixtures/spacewalk.wav"); +// Cloning requires VERIFIED consent: the speaker records themselves reading a +// phrase the API returns, and that recording is kept as the consent record. It +// must be the SAME person as the voice sample, so this recipe is +// bring-your-own-audio — there is no sample that ships with valid consent. +const dir = import.meta.dirname; +const CONSENT_FULL_NAME = process.env.CONSENT_FULL_NAME ?? "Jane Doe"; +const samplePath = path.resolve(process.env.SAMPLE_PATH ?? path.join(dir, "../sample.wav")); +const consentPath = path.resolve(process.env.CONSENT_RECORDING_PATH ?? path.join(dir, "../consent.wav")); +// The challenge is single-use and its phrase is dynamic, so we cache it between +// runs: run once to get the phrase, record it, run again to submit. +const challengeCache = path.join(dir, "../.consent-challenge.json"); +interface Challenge { + id: string; + phrase: string; + expires_at: string; +} interface CreatedVoice { id: string; display_name: string; @@ -27,23 +44,58 @@ interface SpeechResponse { billable_characters_count: number; } +function loadChallenge(): Challenge | null { + if (!fs.existsSync(challengeCache)) return null; + const c = JSON.parse(fs.readFileSync(challengeCache, "utf8")) as Challenge; + if (new Date(c.expires_at).getTime() <= Date.now()) return null; // expired → make a fresh one + return c; +} + async function main() { - // 1. Clone a voice from an audio sample (10–30s of clean speech works well). - // POST /v1/voices is multipart/form-data — let fetch set the boundary by - // passing a FormData instance directly (do NOT set Content-Type manually). - // `consent` is REQUIRED: a JSON string attesting you have the speaker's - // permission to clone their voice. + // 1. Get (or reuse) a consent challenge. Its `phrase` is what the speaker must + // read aloud; `id` ties the recording to this consent on the create. + let challenge = loadChallenge(); + if (!challenge) { + const res = await fetch(`${BASE}/v1/voices/consent-challenges`, { + method: "POST", + headers: { ...authHeaders, "Content-Type": "application/json" }, + body: JSON.stringify({ full_name: CONSENT_FULL_NAME }), + }); + if (!res.ok) { + throw new Error(`POST /v1/voices/consent-challenges → ${res.status} ${res.statusText}: ${await res.text()}`); + } + challenge = (await res.json()) as Challenge; + fs.writeFileSync(challengeCache, JSON.stringify(challenge, null, 2)); + } + + // 2. Make sure we have both recordings before spending the (single-use) challenge. + const missing = [ + fs.existsSync(samplePath) ? null : ` sample: ${samplePath} (${CONSENT_FULL_NAME}'s voice, 10–30s of clean speech)`, + fs.existsSync(consentPath) ? null : ` consent: ${consentPath} (the SAME person reading the phrase below)`, + ].filter(Boolean); + if (missing.length > 0) { + console.log( + `\nConsent required. Have ${CONSENT_FULL_NAME} record themselves reading this phrase, exactly as written:\n\n` + + ` "${challenge.phrase}"\n\n` + + `Then provide these files and re-run (paths override via SAMPLE_PATH / CONSENT_RECORDING_PATH):\n` + + missing.join("\n") + + `\n\nChallenge expires ${challenge.expires_at}.\n`, + ); + process.exit(1); + } + + // 3. Clone the voice. POST /v1/voices is multipart/form-data — let fetch set the + // boundary by passing a FormData instance directly (do NOT set Content-Type). const form = new FormData(); form.append("name", "cookbook-cloned-voice"); form.append("gender", "male"); - form.append("consent", JSON.stringify({ fullName: "Jane Doe", email: "jane@example.com" })); - // Wrap the file bytes in a Blob with the original filename for the multipart part. - const sampleBytes = fs.readFileSync(samplePath); - form.append("sample", new Blob([sampleBytes], { type: "audio/wav" }), "spacewalk.wav"); + form.append("consent_challenge_id", challenge.id); + form.append("sample", new Blob([fs.readFileSync(samplePath)]), path.basename(samplePath)); + form.append("consent_recording", new Blob([fs.readFileSync(consentPath)]), path.basename(consentPath)); const createRes = await fetch(`${BASE}/v1/voices`, { method: "POST", - headers: { Authorization: `Bearer ${token}` }, + headers: authHeaders, body: form, }); @@ -60,18 +112,16 @@ async function main() { `POST /v1/voices → ${createRes.status} ${createRes.statusText}: ${await createRes.text()}`, ); } + fs.rmSync(challengeCache, { force: true }); // challenge is spent const voice = (await createRes.json()) as CreatedVoice; console.log(`Cloned voice created: ${voice.id} (${voice.display_name}, type=${voice.type})`); try { - // 2. Synthesize speech using the cloned voice — pass its id as voice_id. + // 4. Synthesize speech using the cloned voice — pass its id as voice_id. const speechRes = await fetch(`${BASE}/v1/audio/speech`, { method: "POST", - headers: { - Authorization: `Bearer ${token}`, - "Content-Type": "application/json", - }, + headers: { ...authHeaders, "Content-Type": "application/json" }, body: JSON.stringify({ input: "Hello from a voice cloned with the Speechify API.", voice_id: voice.id, @@ -88,11 +138,11 @@ async function main() { fs.writeFileSync("output.mp3", Buffer.from(speech.audio_data, "base64")); console.log("Wrote output.mp3"); } finally { - // 3. Clean up so cloned voices don't accumulate on your account. + // 5. Clean up so cloned voices don't accumulate on your account. // Remove this to keep the voice and reuse it later via voice.id. const delRes = await fetch(`${BASE}/v1/voices/${encodeURIComponent(voice.id)}`, { method: "DELETE", - headers: { Authorization: `Bearer ${token}` }, + headers: authHeaders, }); if (!delRes.ok) { console.error(