跳至主要內容

ElevenLabs

ElevenLabs 提供高品質的 AI 語音技術,包括透過其轉錄 API 提供的語音轉文字功能。

屬性詳細資訊
說明ElevenLabs 提供先進的 AI 語音技術,具備語音轉文字轉錄與文字轉語音功能,支援多種語言與說話者分離。
LiteLLM 上的提供者路由elevenlabs/
提供者文件ElevenLabs API ↗
支援的端點/audio/transcriptions, /audio/speech

快速開始

LiteLLM Python SDK

Basic audio transcription with ElevenLabs
import litellm

# Transcribe audio file
with open("audio.mp3", "rb") as audio_file:
response = litellm.transcription(
model="elevenlabs/scribe_v1",
file=audio_file,
api_key="your-elevenlabs-api-key" # or set ELEVENLABS_API_KEY env var
)

print(response.text)

LiteLLM Proxy

1. 設定您的 proxy

ElevenLabs configuration in config.yaml
model_list:
- model_name: elevenlabs-transcription
litellm_params:
model: elevenlabs/scribe_v1
api_key: os.environ/ELEVENLABS_API_KEY

general_settings:
master_key: your-master-key

2. 啟動 proxy

Start LiteLLM proxy server
litellm --config config.yaml

# Proxy will be available at http://localhost:4000

3. 發出轉錄請求

Audio transcription with curl
curl http://localhost:4000/v1/audio/transcriptions \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: multipart/form-data" \
-F file="@audio.mp3" \
-F model="elevenlabs-transcription" \
-F language="en" \
-F temperature="0.3"

回應格式

ElevenLabs 會以 OpenAI 相容格式回傳轉錄回應:

Example transcription response
{
"text": "Hello, this is a sample transcription with multiple speakers.",
"task": "transcribe",
"language": "en",
"words": [
{
"word": "Hello",
"start": 0.0,
"end": 0.5
},
{
"word": "this",
"start": 0.5,
"end": 0.8
}
]
}

常見問題

  1. 無效的 API 金鑰:請確保 ELEVENLABS_API_KEY 已正確設定

文字轉語音(TTS)

ElevenLabs 透過其 TTS API 提供高品質的文字轉語音功能,支援多種聲音、語言與音訊格式。

概覽

屬性詳細資訊
說明使用 ElevenLabs 的進階 TTS 模型將文字轉換為自然發聲的語音
LiteLLM 上的提供者路由elevenlabs/
支援的操作/audio/speech
提供者文件連結ElevenLabs TTS API ↗

支援的模型

模型路由說明
Eleven v3elevenlabs/eleven_v3最具表現力的模型。支援 70+ 種語言,並可透過 audio tags 支援音效與停頓。
Eleven Multilingual v2elevenlabs/eleven_multilingual_v2預設 TTS 模型。支援 29 種語言,穩定且可用於正式環境。

快速開始

LiteLLM Python SDK

ElevenLabs Text-to-Speech with SDK
import litellm
import os

os.environ["ELEVENLABS_API_KEY"] = "your-elevenlabs-api-key"

# Basic usage with voice mapping
audio = litellm.speech(
model="elevenlabs/eleven_multilingual_v2",
input="Testing ElevenLabs speech from LiteLLM.",
voice="alloy", # Maps to ElevenLabs voice ID automatically
)

# Save audio to file
with open("test_output.mp3", "wb") as f:
f.write(audio.read())

使用帶有 Audio Tags 的 Eleven v3

Eleven v3 支援 audio tags,可直接在文字中加入音效與停頓:

Eleven v3 with audio tags
import litellm
import os

os.environ["ELEVENLABS_API_KEY"] = "your-elevenlabs-api-key"

audio = litellm.speech(
model="elevenlabs/eleven_v3",
input='Welcome back. <sfx>applause</sfx> Today we have a special guest. <pause duration="1.5s"/> Let me introduce them.',
voice="alloy",
)

with open("eleven_v3_output.mp3", "wb") as f:
f.write(audio.read())

進階用法:覆寫參數與 ElevenLabs 專屬功能

Advanced TTS with custom parameters
import litellm
import os

os.environ["ELEVENLABS_API_KEY"] = "your-elevenlabs-api-key"

# Example showing parameter overriding and ElevenLabs-specific parameters
audio = litellm.speech(
model="elevenlabs/eleven_multilingual_v2",
input="Testing ElevenLabs speech from LiteLLM.",
voice="alloy", # Can use mapped voice name or raw ElevenLabs voice_id
response_format="pcm", # Maps to ElevenLabs output_format
speed=1.1, # Maps to voice_settings.speed
# ElevenLabs-specific parameters - passed directly to API
pronunciation_dictionary_locators=[
{"pronunciation_dictionary_id": "dict_123", "version_id": "v1"}
],
model_id="eleven_multilingual_v2", # Override model if needed
)

# Save audio to file
with open("test_output.mp3", "wb") as f:
f.write(audio.read())

聲音對應

LiteLLM 會自動將常見的 OpenAI 聲音名稱對應到 ElevenLabs 的聲音 ID:

OpenAI 聲音ElevenLabs 聲音 ID說明
alloy21m00Tcm4TlvDq8ikWAMRachel - 中性且平衡
amber5Q0t7uMcjvnagumLfvZiPaul - 溫暖且友善
ashAZnzlk1XvdvUeBnXmlldDomi - 有活力
augustD38z5RcWu1voky8WS1jaFin - 專業
blue2EiwWnXFnvU5JabPnv8nClyde - 深沉且權威
coral9BWtsMINqrJLrRacOk9xAria - 富有表現力
lilyEXAVITQu4vr4xnSDxMaLSarah - 友善
onyx29vD33N1CtxCmqQRPOHJDrew - 強而有力
sageCwhRBWXzGAHq8TQ4Fs17Roger - 平靜
verseCYw3kZ02Hs0563khs1FjDave - 對話式

使用自訂聲音 ID:您也可以直接傳入任何 ElevenLabs 的聲音 ID。如果聲音名稱不在對應表中,LiteLLM 會原樣使用:

Using custom ElevenLabs voice ID
audio = litellm.speech(
model="elevenlabs/eleven_multilingual_v2",
input="Testing with a custom voice.",
voice="21m00Tcm4TlvDq8ikWAM", # Direct ElevenLabs voice ID
)

回應格式對應

LiteLLM 會將 OpenAI 的回應格式對應到 ElevenLabs 的輸出格式:

OpenAI 格式ElevenLabs 格式
mp3mp3_44100_128
pcmpcm_44100
opusopus_48000_128

您也可以直接使用 output_format 參數傳入 ElevenLabs 專屬的輸出格式。

支援的參數

All Supported Parameters
audio = litellm.speech(
model="elevenlabs/eleven_multilingual_v2", # Required
input="Text to convert to speech", # Required
voice="alloy", # Required: Voice selection (mapped or raw ID)
response_format="mp3", # Optional: Audio format (mp3, pcm, opus)
speed=1.0, # Optional: Speech speed (maps to voice_settings.speed)
# ElevenLabs-specific parameters (passed directly):
model_id="eleven_multilingual_v2", # Optional: Override model
voice_settings={ # Optional: Voice customization
"stability": 0.5,
"similarity_boost": 0.75,
"speed": 1.0
},
pronunciation_dictionary_locators=[ # Optional: Custom pronunciation
{"pronunciation_dictionary_id": "dict_123", "version_id": "v1"}
],
)

LiteLLM Proxy

1. 設定您的 proxy

ElevenLabs TTS configuration in config.yaml
model_list:
- model_name: elevenlabs-tts
litellm_params:
model: elevenlabs/eleven_multilingual_v2
api_key: os.environ/ELEVENLABS_API_KEY

general_settings:
master_key: your-master-key

2. 發出 TTS 請求

簡單用法(OpenAI 參數)

您可以使用標準的 OpenAI 相容參數,而無需任何提供者專屬設定:

Simple TTS request with curl
curl http://localhost:4000/v1/audio/speech \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "elevenlabs-tts",
"input": "Testing ElevenLabs speech via the LiteLLM proxy.",
"voice": "alloy",
"response_format": "mp3"
}' \
--output speech.mp3
Simple TTS with OpenAI SDK
from openai import OpenAI

client = OpenAI(
base_url="http://localhost:4000",
api_key="your-litellm-api-key"
)

response = client.audio.speech.create(
model="elevenlabs-tts",
input="Testing ElevenLabs speech via the LiteLLM proxy.",
voice="alloy",
response_format="mp3"
)

# Save audio
with open("speech.mp3", "wb") as f:
f.write(response.content)
進階用法(ElevenLabs 專屬參數)

注意:使用 proxy 時,提供者專屬參數(例如 pronunciation_dictionary_locatorsvoice_settings 等)必須傳入 extra_body 欄位。

Advanced TTS request with curl
curl http://localhost:4000/v1/audio/speech \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "elevenlabs-tts",
"input": "Testing ElevenLabs speech via the LiteLLM proxy.",
"voice": "alloy",
"response_format": "pcm",
"extra_body": {
"pronunciation_dictionary_locators": [
{"pronunciation_dictionary_id": "dict_123", "version_id": "v1"}
],
"voice_settings": {
"speed": 1.1,
"stability": 0.5,
"similarity_boost": 0.75
}
}
}' \
--output speech.mp3
Advanced TTS with OpenAI SDK
from openai import OpenAI

client = OpenAI(
base_url="http://localhost:4000",
api_key="your-litellm-api-key"
)

response = client.audio.speech.create(
model="elevenlabs-tts",
input="Testing ElevenLabs speech via the LiteLLM proxy.",
voice="alloy",
response_format="pcm",
extra_body={
"pronunciation_dictionary_locators": [
{"pronunciation_dictionary_id": "dict_123", "version_id": "v1"}
],
"voice_settings": {
"speed": 1.1,
"stability": 0.5,
"similarity_boost": 0.75
}
}
)

# Save audio
with open("speech.mp3", "wb") as f:
f.write(response.content)