Create Assistant

Create a new AI voice assistant. The assistant is automatically synced to the voice engine on creation.

Endpoint

Request Body

Core

name
string
required
Assistant name (max 255 characters)
description
string
Assistant description
system_prompt
string
System instructions for the assistant
model
object
LLM configuration
  • provider — openai, google, groq, or custom
  • model — Model name (gpt-4o, llama-3, etc.)
  • temperature — Creativity (0-1)
  • max_tokens — Max response length
  • system_prompt — System instructions (inline override)
  • url and server_credential_id — required as appropriate for custom
  • fallback_providers — ordered fallback model configurations
voice
object
Voice (TTS) configuration
  • provider — 11labs/elevenlabs, spitch, inworld, fishaudio, deepgram, groq, cartesia, or custom
  • voice_id — Specific voice ID
  • language — Voice language
  • speed — Speaking rate
  • fallback_providers — ordered fallback voice configurations
transcriber
object
Transcriber (STT) configuration
  • provider — deepgram, 11labs/elevenlabs, spitch, or custom
  • model — Recognition model
  • language — Expected language
  • vocabulary — Optional array of special terms
  • fallback_providers — ordered fallback transcriber configurations

Provider configuration

The API uses the provider IDs above. 11labs and elevenlabs are equivalent voice and transcription aliases. A custom provider uses url (or base_url) and may reference server_credential_id; custom STT/TTS also accepts a protocol appropriate to the endpoint. fallback_providers is an ordered array. The runtime uses the primary configuration first and tries each fallback in sequence. Each fallback is a component-specific object, for example:

Transcription vocabulary

Send transcriber.vocabulary as an array of strings. Empty and duplicate terms are removed before use.

Messages

first_message
string
Opening greeting when the assistant picks up (text or URL to an audio file)
first_message_mode
string
default:"assistant-speaks-first"
assistant-speaks-first, assistant-waits-for-user, or user-speaks-first
first_message_interruptions_enabled
boolean
default:"true"
Whether the user can interrupt the first message
end_call_message
string
Message the assistant plays when ending the call
end_call_phrases
array
Phrases that trigger an early call end (case-insensitive), e.g. ["goodbye", "bye"]
voicemail_message
string
Message played if the call is forwarded to voicemail

Detection & Limits

voicemail_detection
object
Voicemail detection configuration (enabled, provider, thresholds)
max_duration_seconds
number
default:"600"
Maximum call duration in seconds (default 10 minutes)
silence_timeout_seconds
number
default:"30"
How long to wait before a call is automatically ended due to inactivity

Audio

background_sound
string
default:"office"
Background sound: office, off, or a custom URL
background_sound_volume
number
default:"0.8"
Background sound playback volume (0.0–1.0)
model_output_in_messages_enabled
boolean
default:"false"
Use model output in conversation history instead of transcription
audio_recording_enabled
boolean
default:"false"
Enable audio recording for calls

Post-call Analysis

summary_enabled
boolean
default:"false"
Enable post-call summary generation
analysis_enabled
boolean
default:"false"
Enable the post-call analysis pipeline
analysis_profile_id
string
Analysis profile UUID to use (null = organization default)
analysis_plan
object
Analysis plan configuration: summary, structuredData, successEvaluation. It may produce a summary, extracted structured data, and a success evaluation after a call.

Plans

artifact_plan
object
Artifact generation plan configuration
start_speaking_plan
object
Reply endpointing plan. Use turn_detection_mode (audio or vad), optional turn_detector_version (v1 or v1-mini), wait_seconds, and smart_endpointing_mode (off or livekit).
stop_speaking_plan
object
Interruption plan: num_words, voice_seconds, and optional backoff_seconds.

Voice behavior

voice.character_profile accepts neutral, warm_host, calm_professional, playful_guide, empathetic_support, or confident_expert. For supported expressive TTS, voice.expressive_delivery accepts enabled, style (restrained, warm, or playful), speech_steering.pace (slow, normal, or fast), speech_steering.disfluencies, speech_steering.nonverbal_sounds, and tts_instructions_append (maximum 4,000 characters). Sensitive, financial, health, safety, legal, error, and complaint contexts remain restrained. voice.backchanneling controls audio-only listener acknowledgements:
frequency is minimal, natural, or expressive. Backchannel cues are never added to transcripts or LLM context.
monitor_plan
object
Real-time monitoring plan (listen, control)
avatar_behavior_profile
object
Avatar behavior controls and portable runtime hints (emotional_tone, animation, etc.)
background_speech_denoising_plan
object
Background speech denoising configuration (Krisp, Fourier)
keypad_input_plan
object
Keypad (DTMF) input handling configuration
observability_plan
object
Observability plan (e.g. Langfuse tracing integration)

Compliance

compliance_plan
object
Compliance controls enforced at runtime. The plan is inert unless enabled is true.
ai_disclosure recording_consent The gate listens for the user’s verbal answer (12s timeout per attempt). Declined consent blocks audio recording and is persisted to the session compliance snapshot. transcript_retention sensitive_data Supported categories are financial, credentials, government_identifiers, contact, names, locations, personal_profile, and telecom_identifiers. When enabled, transcripts, chat context, collected data, summaries, messages, and stored text artifacts are redacted. Redacted values become [REDACTED]. audio_redaction Audio-redaction processing is fail-closed: if enabled recording redaction cannot complete, the recording is not released as a normal artifact.

Data Collection

data_collection
object
Structured data collection configuration.
Collection item Field types: name, email, phone, address, dob, credit_card, dtmf, custom.
  • dtmf — keypad capture; config supports num_digits, dtmf_input_timeout, dtmf_stop_event
  • custom — a plain conversational question; the transcript answer is stored

Transport & Integration

transport_configurations
array
Transport provider configurations (e.g. Twilio)
credentials
array
Inline dynamic credentials for calls, e.g. [{ "provider": "openai", "api_key": "sk-..." }]
credential_ids
array
Vapi credential UUIDs to use for provider authentication (e.g. your own OpenAI or ElevenLabs API key). Create credentials via the Server Credentials API.
hooks
array
Event hooks configuration
server
object
Assistant callback configuration. This is separate from organization outbound webhook subscriptions.
  • url — Callback URL for call events
  • timeout — Request timeout in seconds
  • secret — HMAC signing secret for payload verification

Messages & Metadata

client_messages
array
Messages sent to Client SDKs (e.g. transcript, hang, status-update, speech-update)
server_messages
array
Events sent to the assistant server.url: status-update, call.started, assistant.started, call.ended, end-of-call-report, and error.
See Webhook event contracts for envelopes and field availability.
metadata
object
Custom key-value metadata for the assistant

Squad & Attachments

squad_id
string
Squad UUID this assistant belongs to
squad_role
string
Squad role: primary, fallback, or specialist
file_ids
array
UUIDs of knowledge-base files associated with this assistant
tool_ids
array
IDs of tools associated with this assistant

Example Request

Standard Provider (OpenAI LLM + ElevenLabs Voice + Deepgram STT)

Custom Provider (Spitch Voice + Spitch STT + BYO Credentials)

Spitch is routed through the platform; set its provider name, voice ID, and language. Use credential_ids to pass provider credentials without exposing keys in client-side code.

Custom LLM Endpoint

Use a custom OpenAI-compatible endpoint by passing inline credentials:

Compliant, Data-Collecting Assistant

Response