# Audinote audio API v1 Audinote's public API supports HTTP transcription and speech generation, asynchronous transcription jobs, plus realtime ASR and TTS over WebSockets. This is the supported Audinote contract. It uses familiar audio endpoint shapes, but does not support arbitrary OpenRouter models, provider routing, or every OpenRouter option. - Base URL: `https://audinote.app/api/v1` - Machine-readable HTTP schema: [`/openapi.json`](https://audinote.app/openapi.json) ([repository copy](https://audinote.app/openapi.json)); WebSocket events are specified below. - Public reference: [`/developers`](https://audinote.app/developers) - Key management, current credit rates, and browser playground: [`/account/api`](https://audinote.app/account/api) - All `/api/v1` requests require `Authorization: Bearer `, except browser WebSocket upgrades that use a one-use ticket. ## Quickstart Create a key in **Account → API**, copy the secret while it is shown, and set it in your server environment or terminal. The first request below saves an MP3 and spends Audinote credits. ```bash export AUDINOTE_API_KEY='' curl https://audinote.app/api/v1/audio/speech \ -H "Authorization: Bearer $AUDINOTE_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"audinote-tts-v1","input":"สวัสดี","voice":"sai-th","response_format":"mp3"}' \ --output speech.mp3 ``` To try HTTP speech or transcription without writing code, use the playground at `/account/api`. It holds a pasted key in page memory only. Requests in the playground use the same public bearer endpoints and spend credits. ## Authentication and key management Keep keys on your backend. Do not put a key in a URL, client bundle, browser storage, or public repository. Browser cookies do not authorize `/api/v1` requests. A key secret is shown only on creation; listing keys never returns the secret. Members may have up to five active keys and may revoke a key immediately. The following account endpoints use the signed-in browser session, not an API bearer key: | Method and path | Purpose | Response | | --- | --- | --- | | `GET /api/account/keys` | List this member's keys | `{ "data": [{ "id", "name", "prefix", "createdAt", "lastUsedAt", "revokedAt" }] }` | | `POST /api/account/keys` | Create a key from JSON `{ "name": "My app" }` | `201` with key metadata and a one-time `secret` | | `DELETE /api/account/keys/:id` | Revoke a key owned by the member | `204`, or `404` if not found | | `GET /api/account/keys/rates` | Read active customer rates | `{ "asrCreditsPerMinute": number|null, "ttsCreditsPer1000Characters": number|null }` | Key names must be nonblank after trimming and at most 64 characters before trimming. A missing, malformed, unknown, revoked, or erased-account key receives the same generic `401 unauthorized` from `/api/v1`. ## Models and voices Both catalog endpoints require the bearer key. They return a `data` array and list only currently enabled Audinote aliases. | Method and path | Result | | --- | --- | | `GET /api/v1/models` | `audinote-asr-v1` (`audio` → `text`, WAV) and/or `audinote-tts-v1` (`text` → `audio`, supported HTTP formats and separately configured realtime) | | `GET /api/v1/voices` | `sai-th` (Thai) and `rin-en` (English) when TTS is available | ```bash curl https://audinote.app/api/v1/models \ -H "Authorization: Bearer $AUDINOTE_API_KEY" curl https://audinote.app/api/v1/voices \ -H "Authorization: Bearer $AUDINOTE_API_KEY" ``` Public model and voice IDs remain stable Audinote names. Upstream IDs and routing are private and may change. Model entries advertise `operations`, supported `formats`, and, for ASR, `file_mode` (`sync` or `async`) and `max_file_seconds`. File/live defaults are independent. Voice entries depend on configured language voices. Check this catalog rather than assuming every operation or encoding is available. ## HTTP transcription `POST /api/v1/audio/transcriptions` accepts JSON. Encode a mono PCM16 WAV file at 16 or 24 kHz as raw base64; omit any `data:` prefix. The decoded WAV must be at most 20 MiB and at most 10 minutes, and must also fit the selected model duration advertised by `/models`. The HTTP JSON body can be larger because base64 expands the file. ```bash AUDIO_B64=$(base64 < audio.wav | tr -d '\n') curl https://audinote.app/api/v1/audio/transcriptions \ -H "Authorization: Bearer $AUDINOTE_API_KEY" \ -H "Content-Type: application/json" \ -d "{\"model\":\"audinote-asr-v1\",\"input_audio\":{\"data\":\"$AUDIO_B64\",\"format\":\"wav\"},\"language\":\"th\",\"response_format\":\"verbose_json\"}" ``` | JSON field | Required | Accepted value | | --- | --- | --- | | `model` | Yes | `audinote-asr-v1` | | `input_audio.data` | Yes | Raw base64 WAV bytes | | `input_audio.format` | Yes | `wav` | | `language` | No | `th` or `en`; omit for automatic detection | | `response_format` | No | `json` (default) or `verbose_json` | Unsupported fields, formats, and provider passthrough settings are rejected. The default response is JSON with `text`, `model`, and `usage.audio_seconds` / `usage.billed_minutes`. `verbose_json` adds `duration`, `language`, and ordered `segments` with `text`, `start`, `end`, public `speaker` labels, and `timing` (`provider` or `chunk`). `speaker` is `null` when diarization is unavailable; chunk timing covers a transcription chunk and is not precise word alignment. If no language hint was sent, `language` is `null` in the verbose response. ```json { "text": "สวัสดี", "model": "audinote-asr-v1", "usage": { "audio_seconds": 2.4, "billed_minutes": 1 }, "duration": 2.4, "language": "th", "segments": [{ "text": "สวัสดี", "start": 0, "end": 2.4, "speaker": "speaker_1", "timing": "provider" }] } ``` The example response illustrates the shape; text and timing depend on the recording. ## Asynchronous transcription When `/models` reports `file_mode: "async"`, synchronous transcription returns `400 async_required` before reservation. Use `POST /api/v1/audio/transcription-jobs` with the same `model`, `input_audio` and optional `language`; omit `response_format`. The current async adapter is Qwen filetrans. Public responses retain Audinote aliases. ```bash curl https://audinote.app/api/v1/audio/transcription-jobs \ -H "Authorization: Bearer $AUDINOTE_API_KEY" \ -H "Idempotency-Key: upload-001" \ -H "Content-Type: application/json" \ -d "{\"model\":\"audinote-asr-v1\",\"input_audio\":{\"data\":\"$AUDIO_B64\",\"format\":\"wav\"},\"language\":\"th\"}" ``` The `202` response contains `id`, `status: "pending"`, `model`, and `status_url`. Poll `GET /api/v1/audio/transcription-jobs/:id` using the **same API key**. Other keys, including another key owned by the same member, receive `404 not_found`. Polling consumes the normal rate allowance. Status is `pending`, `completed`, or `failed`. Completed responses include `text`, `segments` and `usage` with `audio_seconds`, `billed_minutes`, and `billing_status` (`settled` or `pending`). Failure exposes only `code: "transcription_failed"`. Jobs expire after 24 hours; expired status requests return 404. Credits are held before processing. Durable provider task IDs prevent repeated submission after queue redelivery. A lost enqueue receipt preserves the job, source and hold for scheduled recovery. An ambiguous provider submission becomes failed rather than automatically creating another billable task. Reusing the same idempotency key returns the existing job's ID and current status with 202; the replay response may omit `status_url`, so construct `/api/v1/audio/transcription-jobs/` from the returned ID. An earlier terminal transcript is retrieved through GET. ## HTTP speech `POST /api/v1/audio/speech` accepts JSON and returns raw MP3 or WAV bytes, not JSON, according to the requested supported format. The output is capped at 8 MiB. | JSON field | Required | Accepted value | | --- | --- | --- | | `model` | Yes | `audinote-tts-v1` | | `input` | Yes | 1–2,000 Unicode characters | | `voice` | Yes | `sai-th` or `rin-en` | | `response_format` | No | `mp3` (default) or `wav`, if advertised by `/models` | Use the quickstart curl example above. A successful response has `Content-Type: audio/mpeg` or `audio/wav`, `X-Request-Id`, `X-Audinote-Billed-Minutes` (charged Audinote credits), and `Cache-Control: no-store`. ## Idempotency and rate limits HTTP speech and transcription accept an optional `Idempotency-Key` header containing 1–128 ASCII letters, digits, periods, underscores, or hyphens. Reusing a value with the same key and endpoint returns `409 duplicate_request` before further provider work or billing. The earlier response is not replayed. Asynchronous jobs instead return the existing job ID/status with 202. Requests are limited to 60 per key and 120 per user per minute. The existing daily TTS character quota also applies. `429 rate_limited` may come from either limit or concurrency controls. These limits are for `/api/v1` keys. Sign-in throttling is separate and lives in the D1 `rate_limit` table. ## Realtime ASR Connect to `wss://audinote.app/api/v1/realtime/asr` with `Authorization: Bearer ` from a non-browser client. On `session.ready`, the server reports `model: "audinote-asr-v1"`, `audio_format: "pcm_s16le"`, `sample_rate: 24000`, a `session_id`, and `max_duration_seconds: 3600`. Send binary mono PCM16 little-endian at 24 kHz, preferably in 80 ms frames (3,840 bytes). A frame must be nonempty, even-sized, and no larger than 64 KiB. Buffered audio is capped at 512 KiB. JSON control frames are capped at 8,192 characters. | Direction | Event or frame | Meaning | | --- | --- | --- | | Client → server | Binary PCM16 | Append audio | | Client → server | `{ "type": "input_audio.commit" }` | Mark a boundary; server replies `input_audio.committed` | | Client → server | `{ "type": "session.close" }` | End the session | | Server → client | `speech.start`, `speech.end` | Speech boundary with `turn_id` | | Server → client | `transcript.partial`, `transcript.final` | Text with `turn_id` | | Server → client | `speaker` | Public speaker alias with `turn_id` | | Server → client | `usage` | `audio_seconds` and `billed_minutes` before close, when settlement succeeds | After audio starts, five seconds without more audio closes the session. The maximum duration is one hour. A public failure emits `session.error` with a `code` when possible. ## Realtime TTS Connect to `wss://audinote.app/api/v1/realtime/tts` with the bearer key or a browser ticket. `session.ready` reports `model: "audinote-tts-v1"`, `audio_format: "pcm_s16le"`, `sample_rate: 44100`, a `session_id`, and `max_duration_seconds: 300`. ```json {"type":"text.append","text":"สวัสดี","voice":"sai-th"} ``` Send one or more `text.append` events with the same voice, up to 2,000 characters total, followed by `{ "type": "text.commit" }`. The server acknowledges with `text.committed`. It emits an `audio.chunk` JSON frame with a byte count immediately before each binary PCM16 chunk, then `audio.done` and `usage` (`characters`, `billed_minutes`). Send `{ "type": "session.close" }` to end early. The maximum session duration is five minutes; control JSON is capped at 8,192 characters and output audio at 8 MiB. ## Browser WebSocket tickets Browsers cannot set a WebSocket `Authorization` header. Your backend should request a ticket using the bearer key, then deliver it to your browser through your own authenticated flow: ```bash curl https://audinote.app/api/v1/realtime/tickets \ -H "Authorization: Bearer $AUDINOTE_API_KEY" \ -H "Content-Type: application/json" \ -d '{"kind":"tts"}' ``` The `201` response is `{ "ticket": "audrt_...", "expires_in": 60 }`. Use `kind: "asr"` for the ASR socket. A ticket is bound to that route, valid for 60 seconds, and consumed atomically during upgrade even if the attempt fails. Do not put keys or tickets in URLs. ```js // `ticket` came from your authenticated backend endpoint. const ws = new WebSocket("wss://audinote.app/api/v1/realtime/tts", [ `audinote-ticket.${ticket}`, ]); ws.binaryType = "arraybuffer"; ws.onopen = () => { ws.send(JSON.stringify({ type: "text.append", text: "สวัสดี", voice: "sai-th" })); ws.send(JSON.stringify({ type: "text.commit" })); }; ws.onmessage = (event) => { if (typeof event.data === "string") console.log(JSON.parse(event.data)); else console.log("PCM16 bytes", event.data); }; ``` ## Credits, usage, and privacy Requests spend the member's existing Audinote minute credits. Admins set the active credit rate per started ASR minute and per started 1,000 TTS Unicode characters; the current rates appear at `/account/api` and through the signed-in rates endpoint. The initial rate is one credit per unit, but active rates can change. The selected route and rate are fixed when each request starts. HTTP requests reserve credits before provider work and release them if no work was accepted. Accepted realtime audio or text may be billed after a later disconnect. The credit ledger is the balance source of truth. API usage records store request ID, modality, units, credits, status, and timestamps, not submitted API text or audio. Async jobs separately retain private source audio and normalized transcript results for their 24-hour retrieval window. Random expiring grants allow the provider to read only that job source; terminal state, deletion, account erasure, expiry or key revocation stops grant access. Scheduled cleanup removes expired public-job audio/results; member recordings follow workspace retention. Public responses omit upstream provider names, model and voice IDs, raw upstream errors, and provider headers. ## Errors Routed HTTP failures use this shape, with an `X-Request-Id` header: ```json {"error":{"code":"insufficient_credits","message":"insufficient credits","type":"invalid_request_error"},"request_id":"..."} ``` | Status | Typical code | Action | | --- | --- | --- | | `400` | `invalid_request`, `invalid_audio`, `invalid_voice`, `invalid_input`, `invalid_idempotency_key`, `unsupported_format`, `async_required`, `async_not_supported` | Correct the JSON, audio format, or header | | `401` | `unauthorized` | Check or replace the key | | `402` | `insufficient_credits` | Add credits before retrying | | `409` | `duplicate_request` | Use a new idempotency key only for a new operation | | `413` | `payload_too_large` | Reduce the body size; this pre-route rejection uses a top-level `code` rather than the routed envelope | | `426` | `upgrade_required` | Use a WebSocket upgrade for realtime endpoints | | `404` | `not_found` | Check job ID, owning key and expiry | | `429` | `rate_limited`, `concurrency_limit` | Wait and retry within limits | | `502` | `speech_failed`, `transcription_failed` | Retry later with a new idempotency key if appropriate | | `503` | `service_unavailable`, `billing_pending` | Retry later; keep the request ID for support | Realtime failures emit `{ "type": "session.error", "code": "..." }` before closing when possible. Provider details are not returned. Keep the status, public code, and request ID when asking for support.