Skip to main content

What is Speaker Diarization?

Speaker diarization automatically labels who is speaking at each moment in an audio file. Example Output:

Available Providers

Best for: Speaker identification with high accuracy

Deepgram (Alternative)

Best for: Speed and real-time performance

Pyannote Audio (Advanced)

Best for: Open-source, self-hosted

Use Cases

Customer Service

Scenario: Monitoring agent performance and call quality

Multi-Department Calls

Scenario: Call involves multiple agents/departments

Compliance & Recording

Scenario: Regulatory requirements for call recording

Call Center Analytics

Scenario: Performance tracking and improvement

Setup

Step 1: Configure STT AssemblyAI diarization requires using AssemblyAI as your STT provider.
  1. Go to Settings → Developer Settings
  2. Add AssemblyAI API key (if not already done)
  3. Save configuration
Step 2: Enable Diarization in Agent
  1. Create or Edit Agent
  2. Go to “Speech-to-Text” settings
  3. Select AssemblyAI as provider
  4. Enable “Speaker Diarization” option
  5. Set speaker count:
    • Auto-detect (default)
    • Or specify 2, 3, 4+ speakers
  6. Save
Step 3: Test
  1. Make test call with multiple participants
  2. Check transcript for speaker labels
  3. Verify accuracy
  4. Adjust if needed

Deepgram Setup

Deepgram includes diarization in standard pricing: Step 1: Ensure Deepgram is Configured
  1. Go to Settings → Developer Settings
  2. Add Deepgram API key (if not done)
  3. Save
Step 2: Enable in Agent
  1. Create or Edit Agent
  2. Select Deepgram as STT provider
  3. Enable “Diarization” option
  4. Configure speaker count
  5. Save
Step 3: Test
  1. Make test call
  2. Review transcript for speaker identification
  3. Verify labels are accurate

Pyannote Audio Setup (Advanced)

For self-hosted or advanced use: Step 1: Install Pyannote
Step 2: Download Model
Step 3: Integrate with CallIntel Requires custom implementation. Contact support for guidance.

Configuration

Speaker Count

Auto-Detect:
Specify Count:

Minimum Speaker Duration

Some systems allow configuring minimum speech duration:

Clustering Method

Algorithm used to identify speakers:

Output Format

Transcript with Speaker Labels

Standard Format:

Labeled JSON Output

Custom Speaker Names

Option 1: Generic Labels Use SPEAKER_00, SPEAKER_01, etc. (default) Option 2: Role-Based Configure custom names:
Option 2: Actual Names If known, specify:

Accuracy Considerations

Factors Affecting Accuracy

Positive Factors:
  • Clear audio quality
  • Distinct speakers (different voices)
  • Normal speaking volume
  • Minimal background noise
  • Correct speaker count specified
Negative Factors:
  • Poor audio quality
  • Overlapping speakers
  • Similar voices
  • Background noise
  • Incorrect speaker count

Improving Accuracy

1. Use High-Quality Audio
2. Specify Correct Speaker Count
3. Minimize Overlapping Speech
4. Use Silence Gaps

Cost Analysis

Pricing Comparison

AssemblyAI Diarization:
Deepgram (Diarization Included):
Recommendation: Use Deepgram if diarization needed frequently (included cost).

Use Cases & Examples

Sales Call Recording

Setup:
Output:
Benefits:
  • Quality assurance
  • Sales coaching
  • Compliance
  • Training

Support Escalation

Setup:
Benefits:
  • Track escalation handling
  • Training for supervisors
  • Document decisions
  • Quality control

Conference Call Transcription

Setup:
Benefits:
  • Complete transcript
  • Know who said what
  • Meeting notes
  • Action item tracking

Troubleshooting

Check audio quality, specify correct speaker count, verify that speakers are clearly distinct.
Increase minimum speaker duration setting, or verify audio quality isn’t causing segmentation.
Diarization has limitations with simultaneous speech. Encourage natural turn-taking in your process.
Switch to Deepgram (includes diarization), or only enable for calls where needed.
Specify exact speaker count (don’t use auto-detect) for faster processing.

Best Practices

1. Use Auto-Detect When Unknown

2. Enable Selectively

3. Monitor Accuracy

4. Clear Audio Quality

5. Document Speaker Roles

Advanced Features

Custom Models (Enterprise)

Some providers offer custom diarization models:

Continuous Learning

Performance Metrics

Key Metrics to Track

See Also

Speech-to-Text

Configure STT providers

Call History

View and analyze call transcripts

Quality Assurance

Use diarization for call analysis and coaching

Support

Contact Support