Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 7 additions & 5 deletions fern/customization/speech-configuration.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -38,15 +38,17 @@ This plan defines the parameters for when the assistant begins speaking after th
- **End-of-turn prediction** - predicting when the current speaker is likely to finish their turn.
- **Backchannel prediction** - detecting moments where a listener may provide short verbal acknowledgments like "uh-huh", "yeah", etc. to show engagement, without intending to take over the speaking turn. This is better handled by the assistant's stopSpeakingPlan.

We offer different providers that can be audio-based, text-based, or audio-text based:
Vapi supports built-in smart endpointing, endpointing supplied by compatible transcribers, and custom endpointing models:

**Audio-based providers:**
**Custom endpointing:**

- **Krisp**: Audio-based model that analyzes prosodic and acoustic features such as changes in intonation, pitch, and rhythm to detect when users finish speaking. Since it's audio-based, it always notifies when the user is done speaking, even for brief acknowledgments. Vapi offers configurable acknowledgement words and a well-configured stop speaking plan to handle this properly.
- **Custom endpointing model**: Lets your server decide how long Vapi waits before considering the customer's speech finished. Use it when you want your own service to control the endpointing timeout for each turn.

Configure Krisp with a threshold between 0 and 1 (default 0.5), where 1 means the user definitely stopped speaking and 0 means they're still speaking. Use lower values for snappier conversations and higher values for more conservative detection.
Set `startSpeakingPlan.smartEndpointingPlan.provider` to `custom-endpointing-model`. Vapi sends a `call.endpointing.request` containing the conversation history to the configured `server.url` whenever it receives a new transcript.

When interacting with an AI agent, users may genuinely want to interrupt to ask a question or shift the conversation, or they might simply be using backchannel cues like "right" or "okay" to signal they're actively listening. The core challenge lies in distinguishing meaningful interruptions from casual acknowledgments. Since the audio-based model signals end-of-turn after each word, configure the stop speaking plan with the right number of words to interrupt, interruption settings, and acknowledgement phrases to handle backchanneling properly.
Your server returns a `timeoutSeconds` value from 0 to 15. Vapi resets the timeout when it receives the next transcript and sends another request. If `server` is not configured, Vapi uses `assistant.server`, then `org.server`.

See [Custom endpointing model configuration](/customization/voice-pipeline-configuration#custom-endpointing-model-configuration) for configuration details and [Call endpointing request](/server-url/events#call-endpointing-request-custom-endpointing-server) for the request and response contract.

**Audio-text based providers:**

Expand Down
80 changes: 36 additions & 44 deletions fern/customization/voice-pipeline-configuration.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -217,8 +217,8 @@ Uses AI models to analyze speech patterns, context, and audio cues to predict wh
- **livekit**: Advanced model trained on conversation data (English only)
- **vapi**: VAPI-trained model (non-English conversations or LiveKit alternative)

**Audio-based providers:**
- **krisp**: Audio-based model analyzing prosodic features (intonation, pitch, rhythm)
**Custom providers:**
- **custom-endpointing-model**: Sends each endpointing decision to your own server, which returns a `timeoutSeconds` value

**Audio-text based providers:**
- **deepgram-flux**: Deepgram's latest transcriber model with built-in conversational speech recognition. Use `flux-general-en` for English-only conversations or `flux-general-multi` for multilingual conversations.
Expand All @@ -235,7 +235,7 @@ Uses AI models to analyze speech patterns, context, and audio cues to predict wh
- **AssemblyAI**: Use when AssemblyAI is already your transcriber provider and you want integrated end-of-turn detection
- **LiveKit**: English conversations where Deepgram is not the transcriber of choice.
- **Vapi**: Non-English conversations with default stop speaking plan settings
- **Krisp**: Non-English conversations with a robustly configured stop speaking plan
- **Custom endpointing model**: When your own server should decide the endpointing timeout for each turn

### Deepgram Flux configuration

Expand Down Expand Up @@ -357,31 +357,52 @@ The system continuously analyzes the latest user message and applies the first m
- Scenarios requiring predictable, rule-based endpointing behavior
- Fallback option when other smart endpointing providers aren't suitable

### Krisp threshold configuration
### Custom endpointing model configuration

Krisp's audio-base model returns a probability between 0 and 1, where 1 means the user definitely stopped speaking and 0 means they're still speaking.

**Threshold settings:**

- **0.0-0.3:** Very aggressive detection - responds quickly but may interrupt users mid-sentence
- **0.4-0.6:** Balanced detection (default: 0.5) - good balance between responsiveness and accuracy
- **0.7-1.0:** Conservative detection - waits longer to ensure users have finished speaking
Sends each endpointing decision to your own server instead of a built-in model. Vapi POSTs the current transcript to `server.url`; your server responds with a `timeoutSeconds` value, the number of seconds to wait before considering the user's turn finished. The timeout resets each time a new transcript is received.

**Configuration example:**

```json
{
"startSpeakingPlan": {
"smartEndpointingPlan": {
"provider": "krisp",
"threshold": 0.5
"provider": "custom-endpointing-model",
"server": {
"url": "https://your-server.com/endpointing"
}
}
}
}
```

**Important considerations:**
Since Krisp is audio-based, it always notifies when the user is done speaking, even for brief acknowledgments. Configure the stop speaking plan with appropriate `acknowledgementPhrases` and `numWords` settings to handle backchanneling properly.
**Request sent to your server:**

```json
{
"message": {
"type": "call.endpointing.request",
"messages": [
{
"role": "user",
"message": "Hello, how are you?",
"time": 1234567890,
"secondsFromStart": 0
}
]
}
}
```

**Expected response:**

```json
{
"timeoutSeconds": 0.5
}
```

If `server` is not provided, the request is sent to `assistant.server`, then `org.server` if that isn't set either.

### Assembly turn detection

Expand Down Expand Up @@ -665,35 +686,6 @@ User Interrupts → Assistant Audio Stopped → backoffSeconds Blocks All Output

**Optimized for:** Text-based endpointing with longer timeouts for different speech patterns and international support.

### Audio-based endpointing (Krisp example)

```json
{
"startSpeakingPlan": {
"waitSeconds": 0.4,
"smartEndpointingPlan": {
"provider": "krisp",
"threshold": 0.5
}
},
"stopSpeakingPlan": {
"numWords": 2,
"voiceSeconds": 0.2,
"backoffSeconds": 1.0,
"acknowledgementPhrases": [
"okay",
"right",
"uh-huh",
"yeah",
"mm-hmm",
"got it"
]
}
}
```

**Optimized for:** Non-English conversations with robust backchanneling configuration to handle audio-based detection limitations.

### Audio-text based endpointing (Assembly example)

```json
Expand Down
Loading