Kurdish TTS & STT API
A simple HTTP API for Kurdish text-to-speech and speech-to-text. Text-to-speech covers Sorani (Central Kurdish), Kurmanji (Northern Kurdish) and Badini (Behdinî — Northern Kurdish in the Arabic script, with six dedicated voices); speech-to-text covers Sorani and Kurmanji. Send text, get natural Kurdish speech from hundreds of voices; or send audio, get an accurate transcript — by file upload or live streaming. There is a free tier, and paid plans start at $12/month, or pay yearly and save 20%.
Machine-readable spec: OpenAPI 3.1 (/openapi.json). Get an API key in Settings → API. See pricing & plans. New here? Start with the step-by-step guides.
Authentication
All endpoints except GET /api/get-speakers require an API key in the x-api-key request header. Base URL: https://www.kurdishtts.com.
TTS and STT use separate keys. A TTS key authenticates the text-to-speech endpoint; an STT key authenticates the speech-to-text endpoints. They are not interchangeable — a TTS key returns 401 against an STT endpoint. Generate both in Settings → API. Keep keys server-side; never ship them in client code.
Text-to-Speech — POST /api/tts-proxy
There is one TTS model: V5. It renders every request. model_version does not pick a model or an engine — it picks which speaker group your speaker_id is looked up in: v3 (198 speakers), v4 (664), or v5 (20 curated, the six badini_ speakers, plus the 13 Cast and Studio speakers on plans bought for the API). The groups share no IDs, so send the one your speaker belongs to. Omitting the field selects v3, which is why every example below states it explicitly. The one exception is a badini_ id: it exists in no other group, so it resolves on v5 whatever you send. Converts Kurdish text to speech and returns audio/wav by default, or JSON with base64 audio and word-level timestamps when include_timestamps is true. The dialect is derived from the speaker_id prefix (sorani_… / kurmanji_… / badini_…).
| Field | Type | Required | Description |
|---|---|---|---|
text | string | Yes | Text to synthesize. Max 500 chars (free) / 5000 (paid). |
speaker_id | string | Yes | Built-in id from /api/get-speakers, or an owned clone_ id from the Voice Cloning page. Clone API use requires a paid subscription. |
dialect | "sorani" | "kurmanji" | "badini" | No | Reading language. Inferred from the speaker_id prefix when it has one. Send it for the Cast and Studio speakers: they read Sorani or Kurmanji and otherwise default to the Sorani of their reference clip. Badini (Behdinî) is Northern Kurdish in the Perso-Arabic script and has its own six badini_ speakers — a badini_ id resolves on v5 whatever model_version you send, and pairing dialect 'badini' with any other voice is refused with a 400. |
model_version | "v3" | "v4" | "v5" | No | Speaker group your speaker_id is resolved in — not a model or engine choice; every value renders on V5. v3 = 198 speakers, v4 = 664, v5 = 20 curated plus the six badini_ speakers and the 13 Cast and Studio ones on plans bought for the API. Defaults to v3 when omitted, and the groups share no IDs, so send it explicitly. |
include_timestamps | boolean | No | Default false. true → JSON with base64 audio + word timestamps. |
format | "wav" | "opus" | "mp3" | No | Default wav. opus = Ogg/Opus at 24 kHz, ~10–15× smaller — ideal for mobile data. mp3 is also 24 kHz. Ignored when include_timestamps is true. |
speed | number | No | 0.25–4.0, higher = faster (industry convention; inverted internally). |
temperature / stability | number | No | Optional generation controls. Mutually exclusive; omit both for the engine default. |
seed | integer | No | Optional reproducibility control. |
pitch, top_p, repetition_penalty… | number | No | Optional advanced and post-processing controls; forwarded to V5 for compatibility. |
Example — cURL
curl -X POST https://www.kurdishtts.com/api/tts-proxy \
-H "x-api-key: YOUR_TTS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"سڵاو، چۆنیت؟","speaker_id":"sorani_1","model_version":"v4"}' \
--output speech.wavExample — Python
import requests
resp = requests.post(
"https://www.kurdishtts.com/api/tts-proxy",
headers={"x-api-key": "YOUR_TTS_API_KEY"},
json={"text": "سڵاو، چۆنیت؟", "speaker_id": "sorani_1", "model_version": "v4"},
)
with open("speech.wav", "wb") as f:
f.write(resp.content)Example — JavaScript
const res = await fetch("https://www.kurdishtts.com/api/tts-proxy", {
method: "POST",
headers: {
"x-api-key": "YOUR_TTS_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({ text: "سڵاو، چۆنیت؟", speaker_id: "sorani_1", model_version: "v4" }),
});
const audio = await res.arrayBuffer(); // audio/wavExample: owned cloned voice
curl -X POST https://www.kurdishtts.com/api/tts-proxy -H "x-api-key: YOUR_TTS_API_KEY" -H "Content-Type: application/json" -d '{"text":"سڵاو، چۆنیت؟","speaker_id":"clone_YOUR_CLONE_ID","dialect":"sorani"}' --output cloned-voice.wavGood to know
speedfollows the industry convention: higher = faster.format: "opus"returnsaudio/ogg(Opus, 24 kHz) — the response is ~10–15× smaller than WAV.formatonly applies to binary responses; the JSON/timestamps response stays base64 PCM.v3andv4are permanent speaker-group selectors, not a deprecated path. Existing speaker IDs and plan entitlements do not change, and will not.- Any paid API plan may select
v5and render its 20 curated speakers and all sixbadini_ones. The 13 Cast and Studio speakers are listed to everyone for discovery but render only on plans bought for the API — Developer, Pro, Business, and the legacystarter,starter-v4andapi-pro. Naming one on another plan returns403withupgrade_required. - What a plan sees inside a group is not uniform. Free gets four speakers each from
v3andv4, plusbadini_story_mandbadini_narrator_fonv5— its only access to that group; the legacy plans are deliberately lopsided (starterhas everyv3speaker but four fromv4;starter-v4is the reverse). Settings → API lists exactly what your key can call. - Creator, Developer, Pro, and Business may send an owned
clone_…ID. Free accounts cannot use cloned voices through the API. - Clone allowances are Creator 3, Developer 10, Pro 20, and Business 30. Free includes one clone for the fixed website preview only.
Text-to-Speech (streaming) — POST /api/tts-stream
Streams speech while it is being generated instead of waiting for the full file — for short sentences the first audio typically arrives in about a second. Same key, limits, and billing as /api/tts-proxy (characters are debited when the stream starts). SSE and PCM are progressive V5 streams. WAV remains accepted for backward compatibility but is buffered before the response begins.
| Field | Type | Required | Description |
|---|---|---|---|
text | string | Yes | Text to synthesize. Max 500 chars (free) / 5000 (paid). Send sentence-by-sentence for the fastest first chunk. |
speaker_id | string | Yes | Built-in id from /api/get-speakers, or an owned clone_ id on a paid subscription. |
stream_format | "sse" | "pcm" | "wav" | No | Default sse (JSON events, base64 audio). pcm = progressive headerless 16-bit little-endian mono PCM at 24 kHz. wav = buffered compatibility response. |
model_version | "v3" | "v4" | "v5" | No | Speaker group your speaker_id is resolved in, exactly as on /api/tts-proxy. Same values, same v3 default, same entitlements. Not a model or engine choice. |
speed | number | No | 0.25–4.0, higher = faster (same convention as /api/tts-proxy). |
include_timestamps | boolean | No | With SSE, emits speech.word_timestamps events and includes the cumulative list in speech.audio.done. |
Example — cURL (raw PCM)
curl -N -X POST https://www.kurdishtts.com/api/tts-stream \
-H "x-api-key: YOUR_TTS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Silav, tu çawa yî?","speaker_id":"kurmanji_236","model_version":"v4","stream_format":"pcm"}' \
--output speech.pcm
# play it: ffplay -f s16le -ar 24000 -ch_layout mono speech.pcmExample — JavaScript (raw PCM)
const res = await fetch("https://www.kurdishtts.com/api/tts-stream", {
method: "POST",
headers: {
"x-api-key": "YOUR_TTS_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
text: "Silav, tu çawa yî?",
speaker_id: "kurmanji_236",
model_version: "v4", // speaker group the id lives in — see "List voices"
stream_format: "pcm", // headerless 16-bit LE mono PCM @ 24000 Hz
}),
});
const reader = res.body.getReader();
while (true) {
const { done, value } = await reader.read();
if (done) break; // normal end of the HTTP response = end of speech
playPcmChunk(value); // Uint8Array of raw samples — feed your audio buffer
}Example — JavaScript (SSE)
const res = await fetch("https://www.kurdishtts.com/api/tts-stream", {
method: "POST",
headers: {
"x-api-key": "YOUR_TTS_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
text: "سڵاو، چۆنیت؟",
speaker_id: "sorani_1",
model_version: "v4", // speaker group the id lives in — see "List voices"
stream_format: "sse",
}),
});
const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
let buf = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buf += value;
const events = buf.split("\n\n");
buf = events.pop() ?? ""; // keep the trailing partial event
for (const line of events) {
if (!line.startsWith("data: ")) continue;
const evt = JSON.parse(line.slice(6));
if (evt.type === "speech.audio.delta") playBase64Pcm(evt.audio); // 16-bit PCM @ 24000 Hz
if (evt.type === "speech.word_timestamps") updateCaptions(evt.words);
if (evt.type === "error") throw new Error(evt.error);
}
}Good to know
- PCM chunks after the first are about 2 seconds of audio (96,000 bytes at the default 24 kHz rate) — do not hard-code the chunk size; read until the response ends. The stream ends with the normal end of the HTTP response — there is no in-band end marker.
- A connection that drops mid-stream means generation failed — discard the audio and retry. In
ssemode you also get an in-banderrorevent. - Response headers describe the audio (
X-Sample-Rate: 24000,X-Channels: 1,X-Bit-Depth: 16) and echo the selected catalog inX-Model-Version.
Speech-to-Text (file upload) — POST /api/stt-proxy
Transcribes an uploaded audio file (WAV/MP3/FLAC/OGG/M4A) sent as multipart/form-data. One credit is debited per successful transcription. Max file size and transcript length depend on your plan (free: 10 MB / 500 chars; starter: 50 MB / unlimited; pro: 100 MB / unlimited).
| Field | Type | Required | Description |
|---|---|---|---|
file | file | Yes | Audio file (WAV/MP3/FLAC/OGG/M4A). |
dialect | "sorani" | "kurmanji" | Yes | Kurdish dialect of the audio. |
Example — cURL
curl -X POST https://www.kurdishtts.com/api/stt-proxy \
-H "x-api-key: YOUR_STT_API_KEY" \
-F "file=@audio.wav" \
-F "dialect=sorani"Example — Python
import requests
resp = requests.post(
"https://www.kurdishtts.com/api/stt-proxy",
headers={"x-api-key": "YOUR_STT_API_KEY"},
files={"file": open("audio.wav", "rb")},
data={"dialect": "sorani"},
)
print(resp.json()["text"])Response JSON includes text, detected_dialect, detected_script, and language. On the free plan a long transcript may be clipped — indicated by truncated: true and truncation_limit.
Speech-to-Text (live streaming) — POST /api/stt-stream-connect
Real-time transcription over a WebSocket, in two steps:
POST /api/stt-stream-connectwith your STT key and{ "dialect": "sorani" }→ returns a temporarywebsocket_url(connect within 5 minutes; it does not carry your key).- Open the WebSocket and stream raw 16-bit PCM, mono, 16 kHz audio as binary frames. Send
{ "type": "control", "event": "finalize" }to flush. The server streams{ "text": "…", "is_final": bool }messages and{ "type": "control", "event": "done" }when complete.
One streaming session is debited per connect. Session limits and max duration depend on your plan (free: 20 sessions / 2 min; starter: 100 / 10 min; pro: 500 / 30 min).
Example — JavaScript (browser)
const API_BASE = "https://www.kurdishtts.com";
const API_KEY = "YOUR_STT_API_KEY";
let ws;
async function connect(dialect) {
const res = await fetch(API_BASE + "/api/stt-stream-connect", {
method: "POST",
headers: { "x-api-key": API_KEY, "Content-Type": "application/json" },
body: JSON.stringify({ dialect }),
});
if (!res.ok) throw new Error((await res.json()).detail || "Failed to connect");
const data = await res.json();
console.log("Sessions remaining:", data.streaming_sessions_remaining);
ws = new WebSocket(data.websocket_url); // temporary URL, no key inside
ws.onopen = () => capture();
ws.onmessage = (event) => {
const msg = JSON.parse(event.data);
if (msg.type === "control" && msg.event === "done") return;
if (msg.text) console.log(msg.is_final ? "Final:" : "Partial:", msg.text);
};
}
async function capture() {
const stream = await navigator.mediaDevices.getUserMedia({ audio: true });
const ctx = new AudioContext({ sampleRate: 16000 }); // 16 kHz required
const source = ctx.createMediaStreamSource(stream);
const processor = ctx.createScriptProcessor(4096, 1, 1);
processor.onaudioprocess = (e) => {
const input = e.inputBuffer.getChannelData(0);
const pcm16 = new Int16Array(input.length); // 16-bit mono PCM
for (let i = 0; i < input.length; i++) {
pcm16[i] = Math.max(-32768, Math.min(32767, input[i] * 32768));
}
if (ws && ws.readyState === WebSocket.OPEN) ws.send(pcm16.buffer);
};
source.connect(processor);
processor.connect(ctx.destination);
}
// Call when the speaker is done:
function finalize() {
if (ws) ws.send(JSON.stringify({ type: "control", event: "finalize" }));
}
connect("sorani");Example — Python
import asyncio, json
import requests, websockets
import numpy as np
import sounddevice as sd
API_BASE = "https://www.kurdishtts.com"
API_KEY = "YOUR_STT_API_KEY"
async def stream_stt(dialect="sorani"):
resp = requests.post(
API_BASE + "/api/stt-stream-connect",
headers={"x-api-key": API_KEY, "Content-Type": "application/json"},
json={"dialect": dialect},
)
resp.raise_for_status()
data = resp.json()
print("Sessions remaining:", data["streaming_sessions_remaining"])
async with websockets.connect(data["websocket_url"]) as ws:
async def receive():
async for message in ws:
msg = json.loads(message)
if msg.get("type") == "control" and msg.get("event") == "done":
return
if "text" in msg:
print("Final:" if msg.get("is_final") else "Partial:", msg["text"])
receiver = asyncio.create_task(receive())
def callback(indata, frames, time, status):
pcm16 = (indata[:, 0] * 32767).astype(np.int16) # 16-bit mono
asyncio.run_coroutine_threadsafe(ws.send(pcm16.tobytes()), asyncio.get_event_loop())
with sd.InputStream(samplerate=16000, channels=1, callback=callback, blocksize=4096):
await asyncio.sleep(30) # record for 30s
await ws.send(json.dumps({"type": "control", "event": "finalize"}))
await receiver
asyncio.run(stream_stt("sorani"))List voices — GET /api/get-speakers
A public, unauthenticated listing of the speaker groups. Use each returned id as the speaker_id, and send the matching model_version with it — the groups share no IDs. Pass ?model_version=v3 (198 speakers), v4 (664), or v5(the 20 curated speakers, the six badini_ ones, plus the 13 Cast and Studio ones).
curl "https://www.kurdishtts.com/api/get-speakers?model_version=v4"Each speaker has id, name, dialect (sorani/kurmanji/badini) and gender. Entries in the v5 group carry two more: origin, which is v5 for a curated speaker, badini for one of the six Badini narrators, and cast or designed for a Cast or Studio one, and moods, the emotions that speaker was recorded with.
This endpoint lists a group in full and cannot see your key, so it is a catalog rather than an entitlement. What your plan may actually render is on the Settings → API page. A Cast or Studio speaker appears twice in the v5 listing, once under Sorani and once under Kurmanji, because it reads both. The six badini_ speakers appear once, under Badini only: they are the only voices auditioned for it, and pairing dialect "badini" with any other speaker returns 400.
MCP server — for AI agents
AI agents can call Kurdish TTS/STT natively over the Model Context Protocol. The MCP endpoint is https://www.kurdishtts.com/api/mcp (Streamable HTTP transport) and exposes six tools: synthesize_speech, transcribe_audio, start_streaming_transcription, list_voices, list_dialects and get_plan.
list_voices, list_dialects and get_plan need no API key, so an agent can connect and browse the voice catalog and the plan ladder before anyone signs up. Call get_plan first whenever a tool returns 401 or 403 — it reports what the current key may use and why the call was refused.
The three billed tools authenticate with Authorization: Bearer <api-key> — a TTS or STT key from Settings → API, or both joined as <ttsKey>:<sttKey>. Every tool call is validated, metered and billed against your plan exactly like the HTTP endpoints above — the MCP server is a thin wrapper, not a separate quota.
MCP payload limits: 4,000 characters per synthesis in the default mp3 container — 600 with format: "wav", and 500 on free plans — and 3MB of decoded audio per transcription, which is about 90 seconds of 16kHz WAV but roughly 25 minutes of 64kbps MP3, so send compressed audio. For longer texts or larger files call the HTTP API directly. Example client config:
{
"mcpServers": {
"kurdish-tts": {
"url": "https://www.kurdishtts.com/api/mcp",
"headers": {
"Authorization": "Bearer YOUR_TTS_KEY:YOUR_STT_KEY"
}
}
}
}Step-by-step setup for Claude Code, Claude Desktop and Cursor: Add Kurdish voice to Claude.
Errors & limits
Successful responses use 200. Common error statuses:
400— bad request: a missing field, an invalid dialect, an invalidspeed, a file too large, anemotion(moods are not on the API yet), aspeaker_idthat is not in thev5speaker group, or a voice paired with a dialect it cannot read (only the sixbadini_speakers read Badini).401— missing or invalid API key (check you are using the right key space).403— plan inactive or expired, characters exhausted, a speaker outside your plan's slice of the group you selected, a Cast or Studio speaker on a plan not bought for the API, or a reserved speaker. Body includeserrorand oftenupgrade_requiredandupgrade_url.422— the engine rejected a generation parameter (e.g.temperature: 0.0); thedetailarray carries specifics.503— transient. The speaker roster could not be reached, so the request was refused rather than validated against a stale list. Retry.
Two that surprise people
- An unknown
model_versionis a403, not a422, and the message names your plan. The check fails closed on any value it does not recognise, so a typo reads like an entitlement problem. Sendv3,v4orv5. - A speaker that is not in the group you selected returns
400underv5but403underv3andv4— the same mistake, two statuses. Handle both if you validate IDs client-side.
Plans, quotas and prices are on the pricing page. The full machine-readable contract is at /openapi.json.