跳至主要內容

提示快取

支援的提供者:

  • OpenAI (openai/)
  • Anthropic API (anthropic/)
  • Google AI Studio (gemini/)
  • Vertex AI (vertex_ai/, vertex_ai_beta/)
  • Bedrock (bedrock/, bedrock/invoke/, bedrock/converse) (Bedrock 支援提示快取的所有模型)
  • Deepseek API (deepseek/)
  • xAI (xai/)
最低 token 要求

當輸入低於提供者的最小值時,提示快取會被靜默略過 — 不會回傳錯誤。請務必透過檢查回應中的 cache_creation_input_tokens 來確認是否已發生快取。

提供者最低輸入 token
OpenAI1,024
Anthropic (Claude 3.x)1,024
Anthropic (Claude Sonnet/Opus 4.x)2,048
Anthropic (Claude Haiku 4.5+, Opus 4.5+)4,096
Bedrock (Claude 3.5, 3.7)1,024
Bedrock (Claude Sonnet 4.x)2,048
Google Gemini1,024

對於支援的提供者,LiteLLM 會遵循 OpenAI 提示快取的 usage 物件格式:

"usage": {
"prompt_tokens": 2006,
"completion_tokens": 300,
"total_tokens": 2306,
"prompt_tokens_details": {
"cached_tokens": 1920
},
"completion_tokens_details": {
"reasoning_tokens": 0
}
# ANTHROPIC_ONLY #
"cache_creation_input_tokens": 0
}
  • prompt_tokens:這些是所有提示 token,包含 cache-miss 與 cache-hit 的輸入 token。
  • completion_tokens:這些是模型產生的輸出 token。
  • total_tokens:prompt_tokens + completion_tokens 的總和。
  • prompt_tokens_details:包含 cached_tokens 的物件。
    • cached_tokens:這次呼叫中屬於 cache-hit 的 token。
  • completion_tokens_details:包含 reasoning_tokens 的物件。
  • ANTHROPIC_ONLYcache_creation_input_tokens 是寫入快取的 token 數量。(Anthropic 會針對此收費)

快速開始

注意:OpenAI 快取僅適用於包含 1024 個 token 或以上的提示

from litellm import completion 
import os

os.environ["OPENAI_API_KEY"] = ""

for _ in range(2):
response = completion(
model="gpt-4o",
messages=[
# System Message
{
"role": "system",
"content": [
{
"type": "text",
"text": "Here is the full text of a complex legal agreement"
* 400,
}
],
},
{
"role": "user",
"content": [
{
"type": "text",
"text": "What are the key terms and conditions in this agreement?",
}
],
},
{
"role": "assistant",
"content": "Certainly! the key terms and conditions are the following: the contract is 1 year long for $10/mo",
},
{
"role": "user",
"content": [
{
"type": "text",
"text": "What are the key terms and conditions in this agreement?",
}
],
},
],
temperature=0.2,
max_tokens=10,
)

print("response=", response)
print("response.usage=", response.usage)

assert "prompt_tokens_details" in response.usage
assert response.usage.prompt_tokens_details.cached_tokens > 0

OpenAI prompt_cache_keyprompt_cache_retention

OpenAI 提示快取是 自動 的 — 不需要 cache_control 訊息註解。任何包含 1024+ 提示 token 的請求都符合快取資格。

OpenAI 也支援兩個可選參數,可更精細控制快取行為:

  • prompt_cache_key(string)— 一個路由提示,可提升共享長共同前綴之請求的快取命中率。具有相同 cache key 的請求會被路由到相同的後端,提高快取命中的可能性。
  • prompt_cache_retention"in_memory""24h")— 控制快取 TTL。預設為 "in_memory"(5–10 分鐘)。設為 "24h" 可啟用延伸快取,將 KV tensors 卸載到 GPU 本地儲存。
from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = ""

response = completion(
model="gpt-4o",
messages=[
{
"role": "system",
"content": "You are an AI assistant tasked with analyzing legal documents. "
+ "Here is the full text of a complex legal agreement " * 400,
},
{
"role": "user",
"content": "What are the key terms and conditions?",
},
],
prompt_cache_key="legal-doc-analysis",
prompt_cache_retention="24h",
)
print(response.usage)

Anthropic 範例

Anthropic 會針對寫入快取收費。

使用 "cache_control": {"type": "ephemeral"} 指定要快取的內容。

這個相同格式也適用於 Gemini / Vertex AI。對於其他提供者,會被忽略。

from litellm import completion 
import litellm
import os

litellm.set_verbose = True # 👈 SEE RAW REQUEST
os.environ["ANTHROPIC_API_KEY"] = ""

response = completion(
model="anthropic/claude-3-5-sonnet-20240620",
messages=[
{
"role": "system",
"content": [
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing legal documents.",
},
{
"type": "text",
"text": "Here is the full text of a complex legal agreement" * 400,
"cache_control": {"type": "ephemeral"},
},
],
},
{
"role": "user",
"content": "what are the key terms and conditions in this agreement?",
},
]
)

print(response.usage)
最低 token 數(Anthropic)

低於最低值的提示會在不使用快取的情況下處理 — 不會回傳錯誤。請檢查回應中的 cache_creation_input_tokens

模型最低 token 數
Claude 3 Haiku, 3 Sonnet, 3 Opus1,024
Claude 3.5 Sonnet, 3.7 Sonnet1,024
Claude 3.5 Haiku2,048
Claude Sonnet 4.5, Sonnet 4.6, Opus 42,048
Claude Haiku 4.5, Opus 4.5+4,096

Bedrock 範例

LiteLLM 會自動將 OpenAI 格式的 cache_control 標記轉換為 Bedrock 原生的 cachePoint 格式 — 如果您已經在使用 cache_control,現有程式碼不需要修改。

最低 token 數(Bedrock)

低於最低值的提示會在不使用快取的情況下處理 — 不會回傳錯誤。請檢查回應中的 cache_creation_input_tokens

模型家族每次請求最低 token 數
Claude 3.5 Sonnet v2, Claude 3.7 Sonnet1,024
Claude Sonnet 4.5, Sonnet 4.62,048
import litellm

response = litellm.completion(
model="bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0",
messages=[
{
"role": "system",
"content": [
{
"type": "text",
"text": "<your large system prompt here — min 1,024 tokens for Claude 3.x, 2,048 for Claude Sonnet 4.x>",
"cache_control": {"type": "ephemeral"}
}
]
},
{"role": "user", "content": "What is prompt caching?"}
]
)

print(response.usage)
# cache_creation_input_tokens > 0 on first call (cache written)
# cache_read_input_tokens > 0 on subsequent calls (cache hit)

支援的 Bedrock 模型:

模型Bedrock Model ID最低 TokenTTL 選項
Claude 3.5 Sonnet v2anthropic.claude-3-5-sonnet-20241022-v2:01,0245 分鐘、1 小時
Claude 3.7 Sonnetanthropic.claude-3-7-sonnet-20250219-v1:01,0245 分鐘、1 小時
Claude Opus 4anthropic.claude-opus-4-20250514-v1:01,0245 分鐘、1 小時
Claude Sonnet 4.5, 4.6us.anthropic.claude-sonnet-4-5-*, us.anthropic.claude-sonnet-4-6-*2,0485 分鐘、1 小時

也支援上述模型的跨區域推論設定檔。

請參閱 AWS Bedrock prompt caching 文件 以取得完整支援模型與區域清單。

Google AI Studio / Vertex AI (Gemini) 範例

使用相同的 Anthropic 風格 cache_control 格式 — LiteLLM 會自動將其轉換為 Google 的 context caching API

其底層運作方式:

  1. 含有 cache_control 的訊息會被分離並送至 Google 的 cachedContents API
  2. 接著會將快取內容 ID 作為 cachedContent 傳入 Gemini 請求本文
  3. 可跨三種提供者運作:gemini/(Google AI Studio)、vertex_ai/vertex_ai_beta/
  4. 快取內容至少需要 1024 tokens — 低於此值時,快取會被靜默略過
from litellm import completion
import os

os.environ["GEMINI_API_KEY"] = ""

response = completion(
model="gemini/gemini-2.5-flash",
messages=[
{
"role": "system",
"content": [
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing legal documents.",
},
{
"type": "text",
"text": "Here is the full text of a complex legal agreement" * 400,
"cache_control": {"type": "ephemeral"},
},
],
},
{
"role": "user",
"content": "what are the key terms and conditions in this agreement?",
},
],
)

print(response.usage)

Vertex AI

對於 Vertex AI,請使用 vertex_ai/ 前綴:

from litellm import completion

response = completion(
model="vertex_ai/gemini-2.5-flash",
vertex_project="my-gcp-project",
vertex_location="us-central1",
messages=[
{
"role": "system",
"content": [
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing legal documents.",
},
{
"type": "text",
"text": "Here is the full text of a complex legal agreement" * 400,
"cache_control": {"type": "ephemeral"},
},
],
},
{
"role": "user",
"content": "what are the key terms and conditions in this agreement?",
},
],
)

print(response.usage)

Deepeek 範例

與 OpenAI 的運作方式相同。

from litellm import completion 
import litellm
import os

os.environ["DEEPSEEK_API_KEY"] = ""

litellm.set_verbose = True # 👈 SEE RAW REQUEST

model_name = "deepseek/deepseek-chat"
messages_1 = [
{
"role": "system",
"content": "You are a history expert. The user will provide a series of questions, and your answers should be concise and start with `Answer:`",
},
{
"role": "user",
"content": "In what year did Qin Shi Huang unify the six states?",
},
{"role": "assistant", "content": "Answer: 221 BC"},
{"role": "user", "content": "Who was the founder of the Han Dynasty?"},
{"role": "assistant", "content": "Answer: Liu Bang"},
{"role": "user", "content": "Who was the last emperor of the Tang Dynasty?"},
{"role": "assistant", "content": "Answer: Li Zhu"},
{
"role": "user",
"content": "Who was the founding emperor of the Ming Dynasty?",
},
{"role": "assistant", "content": "Answer: Zhu Yuanzhang"},
{
"role": "user",
"content": "Who was the founding emperor of the Qing Dynasty?",
},
]

message_2 = [
{
"role": "system",
"content": "You are a history expert. The user will provide a series of questions, and your answers should be concise and start with `Answer:`",
},
{
"role": "user",
"content": "In what year did Qin Shi Huang unify the six states?",
},
{"role": "assistant", "content": "Answer: 221 BC"},
{"role": "user", "content": "Who was the founder of the Han Dynasty?"},
{"role": "assistant", "content": "Answer: Liu Bang"},
{"role": "user", "content": "Who was the last emperor of the Tang Dynasty?"},
{"role": "assistant", "content": "Answer: Li Zhu"},
{
"role": "user",
"content": "Who was the founding emperor of the Ming Dynasty?",
},
{"role": "assistant", "content": "Answer: Zhu Yuanzhang"},
{"role": "user", "content": "When did the Shang Dynasty fall?"},
]

response_1 = litellm.completion(model=model_name, messages=messages_1)
response_2 = litellm.completion(model=model_name, messages=message_2)

# Add any assertions here to check the response
print(response_2.usage)

計算成本

快取命中時的提示 token 成本可能與快取未命中時不同。

使用 completion_cost() 函式來計算成本(也會處理提示快取成本計算)。查看更多輔助函式

cost = completion_cost(completion_response=response, model=model)

用法

from litellm import completion, completion_cost
import litellm
import os

litellm.set_verbose = True # 👈 SEE RAW REQUEST
os.environ["ANTHROPIC_API_KEY"] = ""
model = "anthropic/claude-3-5-sonnet-20240620"
response = completion(
model=model,
messages=[
{
"role": "system",
"content": [
{
"type": "text",
"text": "You are an AI assistant tasked with analyzing legal documents.",
},
{
"type": "text",
"text": "Here is the full text of a complex legal agreement" * 400,
"cache_control": {"type": "ephemeral"},
},
],
},
{
"role": "user",
"content": "what are the key terms and conditions in this agreement?",
},
]
)

print(response.usage)

cost = completion_cost(completion_response=response, model=model)

formatted_string = f"${float(cost):.10f}"
print(formatted_string)

檢查模型支援

使用 supports_prompt_caching() 檢查模型是否支援提示快取

from litellm.utils import supports_prompt_caching

supports_pc: bool = supports_prompt_caching(model="anthropic/claude-3-5-sonnet-20240620")

assert supports_pc

這會檢查我們維護的 模型資訊/成本對照表

進一步閱讀

自動注入提示快取

想讓 LiteLLM 自動加入 cache_control 指令,而不修改您的程式碼嗎?

請參閱 自動注入提示快取教學,了解如何使用 cache_control_injection_points 來自動快取系統訊息、依索引指定的特定訊息,或自訂注入模式。