Create Assistant
Create a new AI voice assistant. The assistant is automatically synced to the voice engine on creation.Endpoint
Request Body
Core
Assistant name (max 255 characters)
Assistant description
System instructions for the assistant
LLM configuration
provider—openai,google,groq, orcustommodel— Model name (gpt-4o,llama-3, etc.)temperature— Creativity (0-1)max_tokens— Max response lengthsystem_prompt— System instructions (inline override)urlandserver_credential_id— required as appropriate forcustomfallback_providers— ordered fallback model configurations
Voice (TTS) configuration
provider—11labs/elevenlabs,spitch,inworld,fishaudio,deepgram,groq,cartesia, orcustomvoice_id— Specific voice IDlanguage— Voice languagespeed— Speaking ratefallback_providers— ordered fallback voice configurations
Transcriber (STT) configuration
provider—deepgram,11labs/elevenlabs,spitch, orcustommodel— Recognition modellanguage— Expected languagevocabulary— Optional array of special termsfallback_providers— ordered fallback transcriber configurations
Provider configuration
The API uses the provider IDs above.11labs and elevenlabs are equivalent voice and transcription aliases. A custom provider uses url (or base_url) and may reference server_credential_id; custom STT/TTS also accepts a protocol appropriate to the endpoint.
fallback_providers is an ordered array. The runtime uses the primary configuration first and tries each fallback in sequence. Each fallback is a component-specific object, for example:
Transcription vocabulary
Sendtranscriber.vocabulary as an array of strings. Empty and duplicate terms are removed before use.
Messages
Opening greeting when the assistant picks up (text or URL to an audio file)
assistant-speaks-first, assistant-waits-for-user, or user-speaks-firstWhether the user can interrupt the first message
Message the assistant plays when ending the call
Phrases that trigger an early call end (case-insensitive), e.g.
["goodbye", "bye"]Message played if the call is forwarded to voicemail
Detection & Limits
Voicemail detection configuration (
enabled, provider, thresholds)Maximum call duration in seconds (default 10 minutes)
How long to wait before a call is automatically ended due to inactivity
Audio
Background sound:
office, off, or a custom URLBackground sound playback volume (0.0–1.0)
Use model output in conversation history instead of transcription
Enable audio recording for calls
Post-call Analysis
Enable post-call summary generation
Enable the post-call analysis pipeline
Analysis profile UUID to use (null = organization default)
Analysis plan configuration:
summary, structuredData, successEvaluation. It may produce a summary, extracted structured data, and a success evaluation after a call.Plans
Artifact generation plan configuration
Reply endpointing plan. Use
turn_detection_mode (audio or vad), optional turn_detector_version (v1 or v1-mini), wait_seconds, and smart_endpointing_mode (off or livekit).Interruption plan:
num_words, voice_seconds, and optional backoff_seconds.Voice behavior
voice.character_profile accepts neutral, warm_host, calm_professional, playful_guide, empathetic_support, or confident_expert.
For supported expressive TTS, voice.expressive_delivery accepts enabled, style (restrained, warm, or playful), speech_steering.pace (slow, normal, or fast), speech_steering.disfluencies, speech_steering.nonverbal_sounds, and tts_instructions_append (maximum 4,000 characters). Sensitive, financial, health, safety, legal, error, and complaint contexts remain restrained.
voice.backchanneling controls audio-only listener acknowledgements:
frequency is minimal, natural, or expressive. Backchannel cues are never added to transcripts or LLM context.
Real-time monitoring plan (
listen, control)Avatar behavior controls and portable runtime hints (
emotional_tone, animation, etc.)Background speech denoising configuration (Krisp, Fourier)
Keypad (DTMF) input handling configuration
Observability plan (e.g. Langfuse tracing integration)
Compliance
Compliance controls enforced at runtime. The plan is inert unless
enabled is true.ai_disclosure
recording_consent
The gate listens for the user’s verbal answer (12s timeout per attempt). Declined consent blocks audio recording and is persisted to the session compliance snapshot.
transcript_retention
sensitive_data
Supported categories are
financial, credentials, government_identifiers, contact, names, locations, personal_profile, and telecom_identifiers. When enabled, transcripts, chat context, collected data, summaries, messages, and stored text artifacts are redacted. Redacted values become [REDACTED].
audio_redaction
Audio-redaction processing is fail-closed: if enabled recording redaction cannot complete, the recording is not released as a normal artifact.
Data Collection
Structured data collection configuration.
Field types:
name, email, phone, address, dob, credit_card, dtmf, custom.
dtmf— keypad capture;configsupportsnum_digits,dtmf_input_timeout,dtmf_stop_eventcustom— a plain conversational question; the transcript answer is stored
Transport & Integration
Transport provider configurations (e.g. Twilio)
Inline dynamic credentials for calls, e.g.
[{ "provider": "openai", "api_key": "sk-..." }]Vapi credential UUIDs to use for provider authentication (e.g. your own OpenAI or ElevenLabs API key). Create credentials via the Server Credentials API.
Event hooks configuration
Assistant callback configuration. This is separate from organization outbound webhook subscriptions.
url— Callback URL for call eventstimeout— Request timeout in secondssecret— HMAC signing secret for payload verification
Messages & Metadata
Messages sent to Client SDKs (e.g.
transcript, hang, status-update, speech-update)Events sent to the assistant
server.url: status-update, call.started, assistant.started, call.ended, end-of-call-report, and error.Custom key-value metadata for the assistant
Squad & Attachments
Squad UUID this assistant belongs to
Squad role:
primary, fallback, or specialistUUIDs of knowledge-base files associated with this assistant
IDs of tools associated with this assistant
Example Request
Standard Provider (OpenAI LLM + ElevenLabs Voice + Deepgram STT)
Custom Provider (Spitch Voice + Spitch STT + BYO Credentials)
Spitch is routed through the platform; set its provider name, voice ID, and language. Usecredential_ids to pass provider credentials without exposing keys in client-side code.

