跳至主要內容

v1.80.11-stable - Google Interactions API

部署此版本

docker run litellm
docker run \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
docker.litellm.ai/berriai/litellm:v1.80.11-stable

主要亮點


UI 上的 Cloudzero 整合

使用者現在可以直接在 UI 上設定他們的 Cloudzero 整合。


效能:LiteLLM SDK 的記憶體使用量與匯入延遲降低 50%

我們已完全重新架構 litellm.__init__.py,將耗費資源的匯入延後到實際需要時才執行,為 109 個元件 實作了延遲載入。

此重構包含 41 個提供者設定類別40 個工具函式、快取實作(Redis、DualCache、InMemoryCache)、HTTP 處理器、記錄、型別,以及其他耗費資源的相依性。像 tiktoken 和 boto3 這類大型函式庫現在改為按需載入,而不是在匯入時就預先載入。

這使 LiteLLM 對於伺服器無狀態函式、Lambda 部署與容器化環境特別有利,因為冷啟動時間與記憶體占用都很重要。


新提供者與端點

新提供者(5 個新提供者)

提供者支援的 LiteLLM 端點說明
Stability AI/images/generations, /images/editsStable Diffusion 3、SD3.5、影像編輯與生成
Venice.ai/chat/completions, /messages, /responses透過 providers.json 的 Venice.ai API 整合
Pydantic AI Agents/a2a用於 A2A 協定工作流程的 Pydantic AI 代理程式
VertexAI Agent Engine/a2a適用於 agentic 工作流程的 Google Vertex AI Agent Engine
LinkUp Search/searchLinkUp 網頁搜尋 API 整合

新 LLM API 端點(2 個新端點)

端點方法說明文件
/interactionsPOST用於對話式 AI 的 Google Interactions API文件
/searchPOST具 reranker 的 RAG Search API文件

新模型 / 更新模型

新模型支援(55+ 個新模型)

提供者模型上下文視窗輸入($/百萬 tokens)輸出($/百萬 tokens)功能
Geminigemini/gemini-3-flash-preview1M$0.50$3.00推理、視覺、音訊、影片、PDF
Vertex AIvertex_ai/gemini-3-flash-preview1M$0.50$3.00推理、視覺、音訊、影片、PDF
Azure AIazure_ai/deepseek-v3.2164K$0.58$1.68推理、函式呼叫、快取
Azure AIazure_ai/cohere-rerank-v4.0-pro32K$0.0025/query-Rerank
Azure AIazure_ai/cohere-rerank-v4.0-fast32K$0.002/query-Rerank
OpenRouteropenrouter/openai/gpt-5.2400K$1.75$14.00推理、視覺、快取
OpenRouteropenrouter/openai/gpt-5.2-pro400K$21.00$168.00推理、視覺
OpenRouteropenrouter/mistralai/devstral-2512262K$0.15$0.60函式呼叫
OpenRouteropenrouter/mistralai/ministral-3b-2512131K$0.10$0.10函式呼叫、視覺
OpenRouteropenrouter/mistralai/ministral-8b-2512262K$0.15$0.15函式呼叫、視覺
OpenRouteropenrouter/mistralai/ministral-14b-2512262K$0.20$0.20函式呼叫、視覺
OpenRouteropenrouter/mistralai/mistral-large-2512262K$0.50$1.50函式呼叫、視覺
OpenAIgpt-4o-transcribe-diarize16K$6.00/audio-含 diarization 的音訊轉錄
OpenAIgpt-image-1.5-2025-12-16-多種多種影像生成
Stabilitystability/sd3-large--$0.065/image影像生成
Stabilitystability/sd3.5-large--$0.065/image影像生成
Stabilitystability/stable-image-ultra--$0.08/image影像生成
Stabilitystability/inpaint--$0.005/image影像編輯
Stabilitystability/outpaint--$0.004/image影像編輯
Bedrockstability.stable-conservative-upscale-v1:0--$0.40/image影像放大
Bedrockstability.stable-creative-upscale-v1:0--$0.60/image影像放大
Vertex AIvertex_ai/deepseek-ai/deepseek-ocr-maas-$0.30$1.20OCR
LinkUplinkup/search-$5.87/1K queries-網頁搜尋
LinkUplinkup/search-deep-$58.67/1K queries-深層網路搜尋
GitHub Copilot20+ models多種--聊天補全

功能

  • Gemini
    • 新增 Gemini 3 Flash Preview Day 0 支援,包含 reasoning - PR #18135
    • 在批次 embeddings 中支援 extra_headers - PR #18004
    • 產生圖片時傳遞 token 使用量 - PR #17987
    • 圖片編輯請求改用 JSON,而非 form-data - PR #18012
    • 修正 web search 請求計數 - PR #17921
  • Anthropic
    • 根據 model 使用動態 max_tokens - PR #17900
    • 將 claude-3-7-sonnet 的 max_tokens 預設修正為 64K - PR #17979
    • 新增與 OpenAI 相容、支援 modify_params=True 的 API - PR #17106
  • Vertex AI
    • 新增 Gemini 3 Flash Preview 支援 - PR #18164
    • 新增 gemini-3-flash-preview 的 reasoning 支援 - PR #18175
    • 修正圖片編輯憑證來源 - PR #18121
    • 對自訂端點將憑證傳遞給 PredictionServiceClient - PR #17757
    • 修正文字 + base64 圖片組合的多模態 embeddings - PR #18172
    • 新增 DeepSeek model 的 OCR 支援 - PR #17971
  • Azure AI
    • 新增 Azure Cohere 4 reranking models - PR #17961
    • 新增 Azure DeepSeek V3.2 版本 - PR #18019
    • 在 get_provider_chat_config 中,針對 Claude models 回傳 AzureAnthropicConfig - PR #18086
  • Fireworks AI
    • 新增 Fireworks AI models 的 reasoning 參數支援 - PR #17967
  • Bedrock
    • 在 get_bedrock_model_id 中新增 Qwen 2 和 Qwen 3 - PR #18100
    • 路由至 bedrock 時移除 ttl 欄位 - PR #18049
    • 新增 Bedrock Stability 圖片編輯 models - PR #18254
  • Perplexity
    • 使用 API 提供的成本,而非手動計算 - PR #17887
  • OpenAI
    • 為音訊轉錄新增 diarize model - PR #18117
    • 在 model cost map 中新增 gpt-image-1.5-2025-12-16 - PR #18107
    • 修正 gpt-image-1 model 的成本計算 - PR #17966
  • GitHub Copilot
    • 新增 github_copilot model 資訊 - PR #17858
  • 自訂 LLM
    • 新增 image_edit 與 aimage_edit 支援 - PR #17999

錯誤修正

  • Gemini
    • 修正 Vertex AI 上 Gemini 3 Flash 的定價 - PR #18202
    • 為 gemini-2.5-flash-image models 新增 output_cost_per_image_token - PR #18156
    • 修正 OBJECT 類型的 properties 應為非空 - PR #18237
  • Qwen
    • 新增 qwen3-embedding-8b 每 token 輸入價格 - PR #18018
  • 一般
    • 修正圖片 URL 處理 - PR #18139
    • 在 Image Processing 中支援附帶查詢參數的 Signed URLs - PR #17976
    • 將 encoding_format 設為 none,而不是省略 - PR #18042

LLM API 端點

功能

錯誤

  • 一般
    • 修正 guardrail translation 中的 basemodel 匯入 - PR #17977
    • 修正 No module named 'fastapi' 錯誤 - PR #18239

管理端點 / UI

功能

  • 虛擬金鑰
    • 為 credentials table 新增 master key 旋轉 - PR #17952
    • 修正 tag 管理以保留 litellm_params 中的加密欄位 - PR #17484
    • 修正 key 刪除與重新產生權限 - PR #18214
  • 模型 + 端點
    • 在 UI 中新增 Models Conditional Rendering - PR #18071
    • 在 UI 中新增 Wildcard Model 的 Health Check Model - PR #18269
    • 自動解析 Vector Store Embedding Model 設定 - PR #18167
  • 向量儲存
    • 新增 Milvus Vector Store UI 支援 - PR #18030
    • 在 Team Update 中持久化 Vector Store 設定 - PR #18274
  • 記錄與支出
    • 在 Logs 中新增 LiteLLM Overhead - PR #18033
    • 在 Logs UI 中顯示 LiteLLM Overhead - PR #18034
    • 在 Usage Page 將 Team ID 解析為 Team Alias - PR #18275
    • 修正 Usage Page Top Key View 按鈕可見性 - PR #18203
  • SSO 與健康狀態
    • 新增 SSO Readiness Health Check - PR #18078
    • 修正 /health/test_connection 以解析如 /chat/completions 的環境變數 - PR #17752
  • CloudZero
  • 一般
    • 更新非 root Docker 的 UI 路徑處理 - PR #17989

錯誤

  • UI 修正
    • 修正登入頁面 Failed To Parse JSON Error - PR #18159
    • 修正 new user 路由的 user_id 衝突處理 - PR #17559
    • 修正 Callback Environment Variables 大小寫 - PR #17912

AI 整合

記錄

防護欄

秘密管理器


支出追蹤、預算與速率限制

  • 電子郵件預算警示 - 當達到預算時傳送電子郵件通知 - PR #17995

MCP 閘道

  • 驗證標頭傳遞 - 新增 MCP 驗證標頭傳遞 - PR #17963
  • 修正 deepcopy 錯誤 - 修正處理請求時的 MCP 工具呼叫 deepcopy 錯誤 - PR #18010
  • 修正 list tool - 修正沒有資料庫連線時 MCP list_tools 無法運作的問題 - PR #18161

代理程式閘道 (A2A)

  • 新提供者:Agent Gateway - 新增對 pydantic ai agents 的支援 - PR #18013
  • VertexAI Agent Engine - 新增 Vertex AI Agent Engine 提供者 - PR #18014
  • 修正模型擷取 - 修正 get_model_from_request() 以從 Vertex AI passthrough URLs 擷取 model ID - PR #18097

效能 / 負載平衡 / 可靠性改進

  • 延遲匯入 - 使用按屬性延遲匯入並擷取共用常數 - PR #17994
  • 延遲載入 HTTP 處理器 - 延遲載入 http handlers - PR #17997
  • 延遲載入快取 - 延遲載入快取 - PR #18001
  • 延遲載入類型 - 延遲載入 bedrock types、.types.utils、GuardrailItem - PR #18053, PR #18054, PR #18072
  • 延遲載入設定 - 延遲載入 41 個設定類別 - PR #18267
  • 延遲載入用戶端裝飾器 - 延遲載入大型用戶端裝飾器匯入 - PR #18064
  • Prisma 建置時間 - 在建置時而非執行時下載 Prisma binaries,以供安全受限環境使用 - PR #17695
  • Docker Alpine - 為 ARM64 音訊處理在 Alpine image 中新增 libsndfile - PR #18092
  • 安全性 - 防止 LiteLLM API key 在 /health endpoint 失敗時外洩 - PR #18133

文件更新

  • SAP 文件 - 更新 SAP 文件 - PR #17974
  • Pydantic AI Agents - 新增關於在 LiteLLM A2A gateway 中使用 pydantic ai agents 的文件 - PR #18026
  • Vertex AI Agent Engine - 新增 Vertex AI Agent Engine 文件 - PR #18027
  • Router 順序 - 新增 router order 參數文件 - PR #18045
  • 秘密管理器設定 - 改善秘密管理器設定文件 - PR #18235
  • Gemini 3 Flash - 在 Gemini 3 Flash 部落格中新增版本需求 - PR #18227
  • README - 擴充 Responses API 區段並更新 endpoints - PR #17354
  • Amazon Nova - 在側邊欄和支援模型中新增 Amazon Nova - PR #18220
  • Benchmarks - 在 benchmarks 文件中新增基礎設施建議 - PR #18264
  • 損壞的連結 - 修正損壞連結的更正 - PR #18104
  • README 修正 - 各種 README 改進 - PR #18206

基礎設施 / CI/CD

  • PR 範本 - 新增 LiteLLM 團隊 PR 範本與 CI/CD 規則 - PR #17983, PR #17985
  • Issue 標記 - 透過元件下拉選單和更多提供者關鍵字改善 issue 標記 - PR #17957
  • PR 範本清理 - 移除 PR 範本中多餘欄位 - PR #17956
  • 相依套件 - 將 altcha-lib 從 1.3.0 升級到 1.4.1 - PR #18017

新貢獻者

  • @dongbin-lunark 在 PR #17757 中完成了第一次貢獻
  • @qdrddr 在 PR #18004 中完成了第一次貢獻
  • @donicrosby 在 PR #17962 中完成了第一次貢獻
  • @NicolaivdSmagt 在 PR #17992 中完成了第一次貢獻
  • @Reapor-Yurnero 在 PR #18085 中完成了第一次貢獻
  • @jk-f5 在 PR #18086 中完成了第一次貢獻
  • @castrapel 在 PR #18077 中完成了第一次貢獻
  • @dtikhonov 在 PR #17484 中完成了第一次貢獻
  • @opleonnn 在 PR #18175 中完成了第一次貢獻
  • @eurogig 在 PR #18084 中完成了第一次貢獻

完整變更記錄

在 GitHub 上檢視完整變更記錄