Speech and transcription for your app
Use Audinote’s own model and voice names through a familiar JSON API. Create a private key, call HTTP endpoints, or stream audio and text over WebSockets.
Quickstart
Create a key, keep it on your server, then send a JSON request. This example saves an MP3 and spends your Audinote credits.
- Create an API key in Account → API. Copy the secret while it is shown.
- Set AUDINOTE_API_KEY in your server environment or terminal. Do not add it to client code.
- Run the request below, then play speech.mp3.
export AUDINOTE_API_KEY='<your-key>'curl https://audinote.app/api/v1/audio/speech \
-H "Authorization: Bearer $AUDINOTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"audinote-tts-v1","input":"สวัสดี","voice":"sai-th","response_format":"mp3"}' \
--output speech.mp3To test without writing code, open the API playground
Authentication and keys
Base URL: https://audinote.app/api/v1. Every catalog and audio request needs an Audinote bearer key in the Authorization header. Browser cookies do not authorize these endpoints.
Authorization: Bearer $AUDINOTE_API_KEYMembers can create and revoke up to five active keys in Account → API. A secret is shown only once; the list shows its name, prefix, creation date, and last use. Revocation blocks new requests. Keep keys out of URLs, browser storage, and public repositories.
Optional Idempotency-Key on HTTP transcription and speech accepts 1–128 letters, digits, periods, underscores, or hyphens. Reusing one for the same key and endpoint returns 409 duplicate_request; it does not replay the previous response. Limits are 60 requests per key and 120 per user per minute.
Models and voices
Catalog responses use a data array and contain only currently enabled Audinote aliases. Both endpoints require a bearer key. The speech model lists MP3 for HTTP and PCM for realtime output.
GET /api/v1/models— Enabled model IDs, modalities, and formatsGET /api/v1/voices— Available voice IDs and languages
curl https://audinote.app/api/v1/models -H "Authorization: Bearer $AUDINOTE_API_KEY"{"data":[{"id":"audinote-asr-v1","object":"model","input_modalities":["audio"],"output_modalities":["text"],"formats":["wav"]},{"id":"audinote-tts-v1","object":"model","input_modalities":["text"],"output_modalities":["audio"],"formats":["mp3","pcm"]}]}{"data":[{"id":"sai-th","name":"Sai","language":"th","model":"audinote-tts-v1"},{"id":"rin-en","name":"Rin","language":"en","model":"audinote-tts-v1"}]}audinote-asr-v1 · audinote-tts-v1 · sai-th · rin-en
Transcriptions
POST /api/v1/audio/transcriptions
Send JSON with raw base64 WAV bytes. Input must be mono PCM16 at 16 or 24 kHz, no longer than 10 minutes and no larger than 20 MiB decoded. Do not include a data URL prefix. Unsupported fields are rejected.
AUDIO_B64=$(base64 < audio.wav | tr -d '\n')
curl https://audinote.app/api/v1/audio/transcriptions \
-H "Authorization: Bearer $AUDINOTE_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\":\"audinote-asr-v1\",\"input_audio\":{\"data\":\"$AUDIO_B64\",\"format\":\"wav\"},\"language\":\"th\",\"response_format\":\"verbose_json\"}"| Field | Required | Meaning |
|---|---|---|
| model | Yes | audinote-asr-v1 |
| input_audio | Yes | Object with data (raw base64) and format: wav |
| language | No | th or en; omit for automatic detection |
| response_format | No | json (default) or verbose_json |
Asynchronous transcription
Check file_mode and max_file_seconds in /models. If transcription returns async_required, submit the same audio body to transcription-jobs without response_format. Use a unique Idempotency-Key per job, then poll status_url with the same key until completed or failed. Jobs expire after 24 hours. Timing can cover an entire chunk; speaker is null when diarization is unavailable.
POST /api/v1/audio/transcription-jobs
GET /api/v1/audio/transcription-jobs/:id
{ "id": "...", "status": "pending", "status_url": "/api/v1/audio/transcription-jobs/..." }Response
{
"text": "สวัสดี",
"model": "audinote-asr-v1",
"usage": { "audio_seconds": 2.4, "billed_minutes": 1 },
"duration": 2.4,
"language": "th",
"segments": [{ "text": "สวัสดี", "start": 0, "end": 2.4, "speaker": "speaker_1" }]
}The response always has text, model, and usage.audio_seconds / usage.billed_minutes. verbose_json also has duration, language, and ordered segments with text, start, end, and public speaker labels. When language is omitted, language is null in the verbose response.
Speech
POST /api/v1/audio/speech
Send JSON and receive raw MP3 bytes. Save the response body to a file or play it directly. HTTP speech is capped at 8 MiB of output.
curl https://audinote.app/api/v1/audio/speech \
-H "Authorization: Bearer $AUDINOTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"audinote-tts-v1","input":"สวัสดี","voice":"sai-th","response_format":"mp3"}' \
--output speech.mp3| Field | Required | Meaning |
|---|---|---|
| model | Yes | audinote-tts-v1 |
| input | Yes | 1–2,000 Unicode characters |
| voice | Yes | sai-th / rin-en |
| response_format | No | mp3 (default) or wav, subject to the formats advertised by /models |
A successful response is audio/mpeg, not JSON. X-Audinote-Billed-Minutes reports the charged Audinote credit amount for this request.
Content-Type: audio/mpeg
X-Request-Id: <request-id>
X-Audinote-Billed-Minutes: <credits-charged>
Cache-Control: no-storeRealtime WebSockets
Connect to the ASR or TTS WebSocket with the bearer key in the Authorization header from a non-browser client. Browsers cannot set that header: request a one-use ticket from your backend instead.
curl https://audinote.app/api/v1/realtime/tickets \
-H "Authorization: Bearer $AUDINOTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"kind":"tts"}'POST /realtime/tickets returns { ticket, expires_in: 60 }. Ask for kind asr or tts, pass the ticket to your authenticated browser client, then send audinote-ticket.<ticket> as the WebSocket subprotocol. The /your-app/audinote-ticket URL below represents your own authenticated backend endpoint. The ticket is bound to its route and consumed during upgrade. Never put a key or ticket in a URL.
const { ticket } = await fetch("/your-app/audinote-ticket").then(r => r.json());
const ws = new WebSocket("wss://audinote.app/api/v1/realtime/tts",
[`audinote-ticket.${ticket}`]);
ws.binaryType = "arraybuffer";
ws.onopen = () => {
ws.send(JSON.stringify({ type: "text.append", text: "สวัสดี", voice: "sai-th" }));
ws.send(JSON.stringify({ type: "text.commit" }));
};
ws.onmessage = event => {
if (typeof event.data === "string") console.log(JSON.parse(event.data));
else console.log("PCM16 bytes", event.data);
};WS /api/v1/realtime/asr— Receive session.ready; send binary mono PCM16 at 24 kHz (80 ms frames preferred). Receive speech.start, transcript.partial, transcript.final, speaker, speech.end, and usage. Send input_audio.commit for a boundary or session.close to end.WS /api/v1/realtime/tts— Receive session.ready; send text.append with text and voice, then text.commit. Receive text.committed, audio.chunk metadata followed by binary PCM16 at 44.1 kHz, audio.done, and usage. Send session.close to end early.
{"type":"session.ready","session_id":"...","model":"audinote-asr-v1","audio_format":"pcm_s16le","sample_rate":24000,"max_duration_seconds":3600}
{"type":"transcript.final","turn_id":1,"text":"สวัสดี"}
{"type":"usage","audio_seconds":2.4,"billed_minutes":1}ASR sessions last at most one hour and close after five seconds without audio once streaming starts. TTS sessions last at most five minutes and accept up to 2,000 characters with one voice. Control events and audio frames are size limited; accepted input may be charged after a disconnect.
Credits and usage
API calls use your existing Audinote credits. Transcription is measured per started audio minute; speech per started 1,000 Unicode characters. The credit rate for each active model is set by the administrator and shown in your API settings. Realtime input already accepted for processing may be charged if the connection later drops.
Public model and voice names are Audinote aliases. Upstream providers and their model or voice IDs are not returned by the API. Usage records include units, charges, status, and timestamps, without storing submitted API audio or text.
Open API settings and examplesErrors and limits
Audio route errors return JSON with error.code, error.message, error.type, and request_id. A request rejected before routing for excessive body size uses a top-level code: payload_too_large. Use the code for handling and keep request_id when supplied. Realtime errors emit session.error with a code before the socket closes when possible.
{ "error": { "code": "insufficient_credits", "message": "insufficient credits", "type": "invalid_request_error" }, "request_id": "..." }Common statuses: 400 invalid request/audio/voice; 401 missing or revoked key; 402 insufficient_credits; 409 duplicate_request; 413 payload_too_large; 426 upgrade_required; 429 rate_limited; 502 provider work failed; 503 service unavailable or billing pending.
Routed public HTTP responses include X-Request-Id. Keep it with the status and error code when troubleshooting. Current credit rates are shown in Account → API.