跳至主要內容

v1.90.0 - 六個新提供者、OpenTelemetry v2 對等功能與串流可靠性

部署此版本

docker run \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
docker.litellm.ai/berriai/litellm:1.90.0

重點摘要

  • 六個新提供者 - ModelScope、LibertAI、Parasail、Pinstripes、TinyFish(搜尋)與 FastCRW(搜尋)- 另外還有新的 e2b 程式碼執行 sandbox 基元。
  • 91 個新模型 涵蓋 Fireworks AI、Scaleway、Tensormesh、LibertAI、Azure AI(包含 gpt-5.5 與 DeepSeek V4)以及 Bedrock Mantle。
  • OpenTelemetry v2 在指標上達到與 v1 對等,會發出六個 gen_ai.client.* 指標、標記輸入/輸出訊息內容,並依租戶範圍隔離 OTLP 憑證。
  • 廣泛的串流可靠性檢視:當用戶端在串流中途中斷連線時,會釋放上游連線(Gemini、aiohttp),請求會被乾淨地取消,且中斷的串流會記錄部分消費。
  • 兩個新的防護欄(Cisco AI Defense、Repello Argus)以及一項大規模的 Next.js App Router UI 遷移,涵蓋 models、teams、users、organizations、api-keys 與 usage 頁面。

App Router 路由

我們正在把 Admin UI 從基於查詢參數的路由遷移到 Nextjs App Router。其動機是路由現在位於 URL 中,因此任何檢視(特定團隊、篩選過的 usage 報告、單一 key)都會變成可分享的連結,您可以傳給同事或加入書籤,而不是只存在於記憶體中的用戶端狀態。

這麼做有雙重動機:它為許多高度期待的 UI 功能/改善奠定基礎,也為更容易為 LiteLLM 貢獻、讓維護者更容易以人類可讀的方式審查程式碼奠定基礎。

其中最大的好處將會是能夠分享不同頁面的連結,例如特定 logs 頁面、teams 頁面等等。

新提供者與端點

新提供者(6 個新提供者)

提供者支援的 LiteLLM 端點說明
ModelScope (modelscope)聊天補全與 OpenAI 相容的 ModelScope 託管模型提供者 - PR #28460
LibertAI (libertai)Chat Completions, Embeddings以 JSON 設定的 OpenAI 相容提供者;提供 12 個目錄模型,包含 bge-m3 embeddings - PR #30203
TinyFish (tinyfish)搜尋網頁搜尋提供者 - PR #30634
FastCRW (fastcrw)搜尋網頁搜尋提供者 - PR #30434
Parasail (parasail)聊天補全與 OpenAI 相容的提供者
Pinstripes (pinstripes)聊天補全新的聊天提供者;提供 6 個目錄模型

新 LLM API 端點

功能說明文件
程式碼執行 (e2b)用於執行模型生成程式碼的新 sandbox / code-interpreter 基元 - PR #30898Sandbox

新模型 / 更新模型

新模型支援(91 個新模型)

提供者模型上下文輸入 ($/1M)輸出 ($/1M)功能
Azure AIazure_ai/gpt-5.51,050,000$5$30reasoning, function calling, prompt caching, pdf, vision
Azure AIazure_ai/gpt-5.5-2026-04-231,050,000$5$30reasoning, function calling, prompt caching, pdf, vision
Azure AIazure_ai/deepseek-v4-flash1,000,000$0.19$0.51reasoning, function calling
Azure AIazure_ai/deepseek-v4-pro1,000,000$1.74$3.48reasoning, function calling
Azure AIazure_ai/deepseek-v3.1131,072$1.23$4.94reasoning, function calling
Azure AIazure_ai/MAI-Image-2.5-$5-影像生成
Azure AIazure_ai/MAI-Image-2.5-Flash-$1.75-影像生成
Azure AIazure_ai/MAI-Image-2e-$5-影像生成
Azureazure/gpt-realtime-whisper---音訊轉錄
OpenAIgpt-realtime-whisper---音訊轉錄
DeepSeekdeepseek-v4-flash / deepseek/deepseek-v4-flash1,000,000$0.14$0.28function calling, prompt caching
DeepSeekdeepseek-v4-pro / deepseek/deepseek-v4-pro1,000,000$0.43$0.87function calling, prompt caching
Mistralmistral/mistral-medium-3-5262,144$1.50$7.50function calling, vision
GitHub Copilotgithub_copilot/mai-code-1-flash128,000$0.75$4.50函式呼叫
Fireworks AI24 models incl. deepseek-v4-pro, glm-5p2, kimi-k2p6/kimi-k2p7-code, minimax-m3, qwen3p7-plus, gpt-oss-120b/gpt-oss-20bup to 1,048,576$0.07-$2.80$0.28-$8.80function calling, reasoning, vision
Bedrock Mantlebedrock_mantle/google.gemma-4-26b-a4b / gemma-4-31b / gemma-4-e2b128k-256k$0.04-$0.14$0.08-$0.40function calling, reasoning, vision
LibertAI12 models incl. qwen3.6-35b-a3b(-thinking), gemma-4-31b-it(-thinking), deepseek-v4-flash, bge-m3up to 262,144$0.01-$0.25free-$1.75function calling, reasoning, vision, embedding
Pinstripes6 models incl. ps/minimax-m2.7, ps/qwen3.6-35b-a3b, ps/glm-4.5-air, ps/deepseek-v4-flashup to 1,000,192$0.09-$0.30$0.20-$0.60function calling, reasoning
Scaleway17 models incl. qwen3.5-397b-a17b, mistral-medium-3.5-128b, gemma-4-26b-a4b-it, gpt-oss-120b, whisper-large-v3up to 256,000free-$1.50free-$7.50function calling, reasoning, vision, audio, embedding
Tensormesh10 models incl. Qwen3-Coder-480B-A35B-FP8, Qwen3.5-397B-A17B-FP8, Kimi-K2.6, DeepSeek-V4-Flash, gpt-oss-120b/gpt-oss-20bup to 262,144$0.07-$1.40$0.28-$4.40function calling, reasoning, prompt caching
Sonioxsoniox/stt-async-v58,000--音訊轉錄
TinyFishtinyfish/search---search

這 91 個新項目也包含完整的 fireworks_ai/accounts/... 模型與 router 路徑。Claude Fable 5 已在 v1.89.0 發布,因此不計入此處。完整 diff:model_prices_and_context_window.json

功能

  • Anthropic
    • 顯示 compaction 使用迭代資料 - PR #27065
    • 提供 Anthropic 原生 /v1/models 供 Claude Code gateway 探索 - PR #30273
  • OpenRouter
    • 將 reasoning max 層級對應到 xhigh - PR #28881
  • Bedrock
    • 在 AgentCore InvokeAgentRuntime 中可選擇性轉送多模態內容區塊 - PR #28885
    • 支援批次輸出檔案的檔案內容擷取 - PR #30595
    • 讓 Bedrock Mantle Responses routing 對每個模型採用資料驅動 - PR #30700
  • DashScope
  • OCI
    • 讓 Cohere {{trace}} judges 可運作(tool 參數型別 + agentic tool-calling continuation)- PR #30646

錯誤修正

  • Anthropic
    • /v1/messages 路徑上套用 cache_control_injection_points - PR #30341
    • /v1/messages 回應中移除 LiteLLM 注入的 total_tokens - PR #30382
    • 將 cache_control 注入上限設為 4 個區塊 - PR #30480
    • 在通用 OpenAI 用戶端的多輪重播中,移除孤立的 server_tool_use - PR #30486
    • 不要將工具 type 洩漏到 OpenAI function parameters schema 中 - PR #30618
  • Bedrock
    • /v1/messages adapter 中為 ARN models 保留 cache_control - PR #29823
    • 處理 /v1/messages 上 messages array 內的 role: "system" - PR #30443
    • 為 Bedrock Mantle responses->chat tool calls 使用唯一的 function-call id - PR #30426
    • 為 Bedrock Mantle chat completions 驗證新增 SigV4 備援 - PR #30714
  • Gemini / Vertex AI
    • cachedContents host 使用 get_vertex_base_url - PR #29707
    • 緩衝原生 Gemini SSE frames - PR #30225
    • 將 Gemini upstream-error body code 429 對應為 RateLimitError - PR #30417
    • 確保檢查顯示 gemini-3-flash-preview 支援 responseJsonSchema - PR #30696
  • OpenAI-compatible
    • 為 OpenAI-compatible 自訂 endpoints 保留 cache_control - PR #30387
    • hosted_vllm: 移除 thinking_blocks 並將 list content 轉換為字串 - PR #30475
    • 不要在具有自訂 prefix 的 wildcard models 上堆疊 provider prefix - PR #30360
  • WatsonX
    • 為 WatsonX API 將字串 embedding input 包裝在 array 中 - PR #30897
  • 價格/成本對照表
    • deepseek-v4-flash/deepseek-v4-pro 新增成本對應 - PR #27056
    • mistral-medium-3-5 新增至成本對應 - PR #29303
    • azure_ai/gpt-5.5 新增至 model cost map - PR #30428
    • 新增 GitHub Copilot MAI Code Flash 定價 - PR #30415
    • 將 Fireworks AI model registry 與目前的平台目錄同步 - PR #30616
    • 新增 soniox/stt-async-v5 - PR #30672
    • 更正 command-r7b-12-2024 的輸入/輸出 token 成本對調 - PR #30413
    • 為 Anthropic Sonnet 4.5/4.6 新增 1h cache-write 成本 - PR #30474
    • 將 Volcengine (Doubao) tiered-pricing models 路由至 tiered cost handler - PR #30357;按數值排序 tiered thresholds - PR #30375;將 DashScope 明確的 0.0 tier cost 視為真實價格 - PR #30653
    • 捨棄 register_model 中合成的零成本,以保留稀疏項目 - PR #30201

LLM API 端點

功能

  • Responses API
    • completed_response 透過 FallbackResponsesStreamWrapper 傳遞,以進行串流 /v1/responses container ownership - PR #30213
  • /v1/models
    • /v1/models 上顯示 max_input_tokens/max_output_tokens - PR #30272
    • 在 v1 model info 中包含 model group aliases - PR #30626
  • Realtime
    • 允許非管理員 virtual keys 呼叫 GA Realtime WebRTC HTTP routes - PR #30089
  • Files
    • 附加既有的 OpenAI file ids - PR #30628

錯誤

  • 一般
    • Token counter:處理 Anthropic tool_reference blocks,以停止遺失的 spend logs - PR #30302
    • Streaming:保護 raise_on_model_repetition 免於空的 choices - PR #30485
    • Audio:不要以 verbose_json 覆寫明確的 response_format - PR #30599
    • 驗證 /realtime/client_secrets 中非轉錄 sessions 的已解析模型 - PR #30710

管理端點 / UI

功能

  • App Router 遷移 - models - PR #30677, teams - PR #30343, users - PR #30334, organizations - PR #30336, api-keys - PR #30699, usage report - PR #30694, agents + router-settings - PR #30323
  • UI cleanup - 移除無法抵達的 /chat 頁面 - PR #30178、失效的 UI components - PR #30340、孤立的 pass-through-settings route - PR #30692;移除產品內問卷與回饋提醒 - PR #30773
  • 虛擬金鑰 - 在 /key/info 中顯示每個模型的 budget 使用量 - PR #30394;grace-period key rotation 在 401 時回傳 deprecated-key lookup 的結果 - PR #30327
  • Teams / Orgs - 新增 key_limit query param 至 /team/info - PR #30006;在 /v1/models 中列出公開 team model names - PR #30588
  • Proxy CLI Auth - 將 verification_uri_complete 新增至 CLI SSO device flow - PR #30571
  • Proxy - 可設定的 response headers 與登入頁面提示 - PR #30792;在 env flag 後方將 /ui/login 上的 "Default Credentials" 提示設為閘道 - PR #30234

錯誤

  • 存取控制 / 金鑰
    • /key/list 現在預設採用精確的 user_id/key_alias 比對,防止跨使用者金鑰外洩 - PR #30593
    • /customer/daily/activity 限制為僅管理員可用 - PR #28849
    • 當 UI 傳送自己的 user_id 時,org_admin 會看到所有組織團隊 - PR #30247
    • 允許內部角色存取向量儲存 CRUD 路由 - PR #30503
    • 僅在啟用 premium 中繼資料欄位時才要求 premium - PR #30506
    • 保護 check_and_fix_namespace 不受 None 金鑰影響 - PR #30435
    • custom_auth 跳過 common_checks 強制執行時,在啟動時發出警告 - PR #30665
    • 從團隊 BYOK 部署解析 list-files 憑證 - PR #30495;在 azure_ad_tokenCredentialLiteLLMParams 中保留 /v1/files + 批次 - PR #30241
    • 對不在成本對照表中的模型強制執行預算 - PR #24949
  • UI
    • 停止 Virtual Keys 頁面的無限重新渲染迴圈 - PR #30397
    • useAuthorized 取得 api-keys 身分,以避免顯示「User ID is not set」 - PR #30903
    • 在自訂 server_root_path 下正確顯示標誌 - PR #31156
    • 在刪除團隊對話框中警告團隊模型將被刪除 - PR #29990
    • 三個小修正 - Gemini api_base、憑證表單重設、Mode 徽章 - PR #30419
    • 將失效的 usage-guide 連結重新指向成本追蹤文件 - PR #30859
  • Proxy
    • 支援 SMTP 隱式 SSL(465 埠) - PR #30395

AI 整合

記錄

  • OpenTelemetry
    • 在 v2 中以與 v1 相同的標準發出六項 gen_ai.client.* 指標 - PR #30326
    • 單一 v2 記錄器擁有全域 provider;依 exporter 範圍限定租戶 OTLP 憑證 - PR #30590
    • 將 v2 gen_ai client 指標匯出到已設定的 meter provider - PR #30549
    • 在 v2 spans 上加上 gen_ai.input/output.messages - PR #30548
    • 使用 include/exclude 清單限制指標屬性的基數 - PR #30257
    • 在 v2 的標準 exception event 上記錄完整錯誤訊息 - PR #30380
    • 接受 v2 中的 UPPER_SNAKE_CASE OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT - PR #30562
  • 一般
    • 在支出記錄中,ProxyException 失敗時保留 error_message - PR #30381

防護欄

  • Cisco AI Defense - 新整合 - PR #28249
  • Repello Argus - 新整合 - PR #30465
  • Presidio - 新增缺少的 UK PII 實體類型 - PR #30537;當防護欄為 logging_only 時,不要遮罩即時請求 - PR #30461
  • AIM - 當 AIM 封鎖請求時,回傳 400 而非 500 - PR #30573
  • 一般
    • 停止在每次輪詢時重新初始化 DB 防護欄 - PR #30542
    • 對模型層級防護欄只執行一次 pre_call hook - PR #30543
    • disable_global_guardrails 會覆寫團隊清單 - PR #28563
    • 在防護欄追蹤中顯示 OpenAI moderation violation_categories - PR #30659

機密管理器

支出追蹤、預算與速率限制

  • 服務等級定價 - 將 service_tier 後綴套用於超過門檻的快取費率,並在 ModelInfo 中顯示 priority+threshold 金鑰 - PR #30450;在成本追蹤中為 Anthropic 回應 service_tier 定價並顯示 - PR #30558;停止非字串 service_tier 悄悄略過成本追蹤 - PR #30690, PR #30706
  • 預算 - 當跨 pod 計數器過期時,根據具權威性的 DB 支出強制執行預算 - PR #30684;當請求在傳輸途中取消時釋放預算保留額度 - PR #30522;當 budget_duration 變更時重新計算 budget_reset_at - PR #30555
  • 速率限制 - 防止內部 parallel_request_limiter 欄位洩漏到上游提供者 - PR #30545
  • 支出準確性 - 在中斷的串流之失敗列上記錄部分支出 - PR #30788;回復中斷的 Anthropic 串流輸出 token - PR #30787;停止 Perplexity 在手動成本備援中對推理 token 進行雙重計費 - PR #30488;修正搭配 ChatCompletionUsageBlock 的快取 token 使用量 - PR #30422
  • 用量彙總 - 在每個 flush 週期排空所有 daily-spend 批次 - PR #30505;在請求記錄中顯示 session-aggregate 成本與持續時間 - PR #30507;合併無支出金鑰的空值彙總 - PR #29945;移除 daily-activity 彙總中的時區日期展開 - PR #29569

MCP 閘道

  • 透過 env vars 使 MCP gateway 名稱與描述可設定 - PR #30473
  • 當範圍篩選器解析為沒有伺服器時,採取封閉失敗 - PR #30353
  • 重新拋出,而不是悄悄丟棄 MCP 團隊權限 - PR #30477
  • 移除委派 OAuth2 tool calls 上虛構的 401 span - PR #30494
  • 使用上游 resource_metadata 向 delegate-auth OAuth servers 發出挑戰 - PR #31255
  • 將 Linear MCP registry 項目預設為可串流 HTTP - PR #30396
  • 在 semantic filter hook 中保留原生 tools - PR #26650

效能 / 負載平衡 / 可靠性改進

  • 串流連線衛生 - 在用戶端中斷連線時取消上游 Gemini 請求並釋放 httpx 連線 - PR #30075; 在用戶端於串流中途中斷時關閉上游 LLM 串流 - PR #30245; 在串流迭代異常結束時釋放 aiohttp 連線 - PR #30271; 在 ModifyResponseException 串流透傳中對 logging_obj 使用 e.request_data - PR #30800
  • 快取 - 新增 valkey-semantic 快取後端並修正 semantic-cache 作用域鍵 - PR #30675; 在 GCS 快取 GET 路徑中對物件名稱進行 URL 編碼 - PR #30378; 允許沒有 Redis 快取的 use_redis_transaction_buffer - PR #28764
  • 路由 / 備援 - 解決模型別名上的 list-unhashable 當機問題 - PR #30464; 在 upsert/delete 時清理 pattern_router 狀態 - PR #29601; 在 SDK 備援回應中保留備援模型 - PR #28260; 新增 expose_router_debug_in_errors(預設為 True)以遮罩內部 model_group/fallback 名稱 - PR #30418
  • 啟動 / 工作程序 - 對非 PostgreSQL 的 DATABASE_URL 失敗即快速結束,而不是卡住 - PR #30366; 新增 --max_requests_before_restart_jitter 以錯開工作程序重啟 - PR #30601; 修正 IAM refresh-engine 監看程式競態條件 - PR #30183; 透過匹配 async_set_cache JSON 編碼來釋放 cron pod-lock - PR #30600
  • 健康檢查 - 修正 Bedrock embedding 健康檢查 - PR #30583; 將健康檢查 max_tokens 預設值提高到 16,以符合 GPT-5 相容性 - PR #30708, PR #26610
  • 開發者體驗 / CI - 約 30 個 PR 強化 lint 與 type-check 閘門(統一採用 basedpyright、移除 mypy、逐步提高任何規範的預算),一個 osv-scanner lockfile 工作流程、zizmor PR 閘門、以本機 fake-OpenAI 測試端點取代共用 mock、相依性升級,以及固定的建置工具鏈。

文件更新

  • 新增 1-click AWS/GCP Terraform 部署按鈕並修正 README 部署按鈕渲染 - PR #29879
  • 強化 CLAUDE.md 中的程式碼慣例 - PR #30333
  • 釐清 PR 範本中 Linear 的部分 - PR #30766

新貢獻者

@hannahmadison, @ayushh0110, @Dotify71, @munnr, @V-3604, @yrk111222, @Silvenga, @djmaze, @apshada, @HumphreySun98, @Harshxth, @tomoyat1, @S0ngRu1, @habonlaci, @moshemalawach, @nahrinoda, @Vedant-Agarwal, @lollinng, @anneheartrecord, @hdt12a1, @vineethsaivs, @krishvsoni, @rvishwas26, @santino18727-debug, @darktheorys, @songkuan-zheng, @Thijmen, @Kropiunig, @jay-tau, @KnyazSh, @koztkozt, @us, @Anuj7411, @zkryakgul, @lavish619, @EugeneLugovtsov, @Bochenski, @menardorama, @factnn, @semmons99, @nitishagar, @FadelT, @jho1-godaddy, @yucheng-berri, @ad1269, @shzdehmd, @vanika02, @Nithish-Yenaganti, @simantak-dabhade, @devYRPauli, @clpatterson, @tcconnally

完整變更記錄

v1.89.0...v1.90.0