AgentAudio

Give your agent ears and a voice

Transcription, speech, and native audio conversations through one router.

Read the launch post →

Paste and ship

Drop this into Cursor, Claude Code, or any agent. It reads the skill manifest and wires up Audio for you.

Agent prompt

Read https://usenaive.ai/skill.md and use the Naïve Audio primitive in my project. Start with: naive audio transcribe meeting.wav

What you get

Three modalitiesSpeech to text, text to speech, speech to speech.
Automatic routingBest model per request, with fallbacks.
Hears toneAudio turns, not transcribe-then-speak.
Route tracesSee every candidate and attempt.

One call to AgentAudio

CLI, SDK, or REST — one bearer token, one credit balance.

POST /v1/audio/transcriptions

Audio

One endpoint for every speech modality. Transcribe recordings, synthesize natural speech, and hold audio-in/audio-out conversations — routed automatically across a managed catalog of speech models, on the same balance as every other primitive.

Three modalities
Speech to text, text to speech, speech to speech.
Automatic routing
Best model per request, with fallbacks.
Hears tone
Audio turns, not transcribe-then-speak.
Route traces
See every candidate and attempt.
# Run with your API key exported
$ naive audio transcribe meeting.wav

Start building with AgentAudio

One API key. One credit balance. Every primitive.