跳至主要內容

Letta 整合

Letta(前稱 MemGPT)是一個用於建構具備持久記憶之有狀態 LLM 代理程式的框架。本指南說明如何將 LiteLLM SDK 與 LiteLLM Proxy 皆整合至 Letta,以便在建構具備記憶功能的代理程式時運用多個 LLM 提供者。

什麼是 Letta?

Letta 可讓您建構能夠:

  • 在對話之間保留長期記憶
  • 使用函式呼叫進行工具互動
  • 有效率地處理大型上下文視窗
  • 持久化代理程式狀態與記憶

前置需求

uv add letta litellm

快速開始

1. 啟動 LiteLLM Proxy

首先,為您的 LiteLLM proxy 建立設定檔:

# config.yaml
model_list:
- model_name: gpt-4
litellm_params:
model: openai/gpt-4
api_key: os.environ/OPENAI_API_KEY

- model_name: claude-3-sonnet
litellm_params:
model: anthropic/claude-3-sonnet-20240229
api_key: os.environ/ANTHROPIC_API_KEY

- model_name: gpt-3.5-turbo
litellm_params:
model: azure/gpt-35-turbo
api_key: os.environ/AZURE_API_KEY
api_base: os.environ/AZURE_API_BASE
api_version: "2023-07-01-preview"

啟動 proxy:

litellm --config config.yaml --port 4000

2. 將 Letta 設定為使用 LiteLLM Proxy

設定 Letta 使用您的 LiteLLM proxy 端點:

import letta
from letta import create_client

# Configure Letta to use LiteLLM proxy
client = create_client()

# Configure the LLM endpoint
client.set_default_llm_config(
model="gpt-4", # This should match a model from your LiteLLM config
model_endpoint_type="openai",
model_endpoint="http://localhost:4000", # Your LiteLLM proxy URL
context_window=8192
)

# Configure embedding endpoint (optional)
client.set_default_embedding_config(
embedding_endpoint_type="openai",
embedding_endpoint="http://localhost:4000",
embedding_model="text-embedding-ada-002"
)

3. 建立並使用 Letta 代理程式

import letta
from letta import create_client

# Create Letta client
client = create_client()

# Create a new agent
agent_state = client.create_agent(
name="my-assistant",
system="You are a helpful assistant with persistent memory.",
llm_config=client.get_default_llm_config(),
embedding_config=client.get_default_embedding_config()
)

# Send a message to the agent
response = client.user_message(
agent_id=agent_state.id,
message="Hi! My name is Alice and I love reading science fiction books."
)

print(f"Agent response: {response.messages[-1].text}")

# Send another message - the agent will remember previous context
response = client.user_message(
agent_id=agent_state.id,
message="What did I tell you about my interests?"
)

print(f"Agent response: {response.messages[-1].text}")

進階設定

為不同代理程式使用不同模型

from letta import LLMConfig, EmbeddingConfig

# Create different LLM configurations pointing to your proxy
gpt4_config = LLMConfig(
model="gpt-4",
model_endpoint_type="openai",
model_endpoint="http://localhost:4000",
context_window=8192
)

claude_config = LLMConfig(
model="claude-3-sonnet",
model_endpoint_type="openai", # Using OpenAI-compatible endpoint
model_endpoint="http://localhost:4000",
context_window=200000
)

# Create agents with different configurations
research_agent = client.create_agent(
name="research-agent",
system="You are a research assistant specialized in analysis.",
llm_config=claude_config # Use Claude for research tasks
)

creative_agent = client.create_agent(
name="creative-agent",
system="You are a creative writing assistant.",
llm_config=gpt4_config # Use GPT-4 for creative tasks
)

搭配工具使用函式呼叫

# Define custom tools for your agent
def search_web(query: str) -> str:
"""Search the web for information"""
# Your web search implementation
return f"Search results for: {query}"

def save_note(content: str) -> str:
"""Save a note to persistent storage"""
# Your note saving implementation
return f"Note saved: {content}"

# Create agent with tools (using proxy endpoint)
agent_state = client.create_agent(
name="research-assistant",
system="You are a research assistant that can search the web and save notes.",
llm_config=client.get_default_llm_config(),
embedding_config=client.get_default_embedding_config(),
tools=[search_web, save_note]
)

# The agent can now use these tools
response = client.user_message(
agent_id=agent_state.id,
message="Search for recent developments in AI and save important findings."
)

驗證

如果您的 LiteLLM proxy 需要驗證:

import os
from letta import LLMConfig

# Set up authenticated configuration
llm_config = LLMConfig(
model="gpt-4",
model_endpoint_type="openai",
model_endpoint="http://localhost:4000",
model_wrapper="openai",
context_window=8192
)

# If using API keys with your proxy
os.environ["OPENAI_API_KEY"] = "your-litellm-proxy-api-key"

client = create_client()
client.set_default_llm_config(llm_config)

對於已啟用驗證的 proxy:

# config.yaml with auth
general_settings:
master_key: "your-master-key"

model_list:
- model_name: gpt-4
litellm_params:
model: openai/gpt-4
api_key: os.environ/OPENAI_API_KEY
# Configure Letta with authenticated proxy
llm_config = LLMConfig(
model="gpt-4",
model_endpoint_type="openai",
model_endpoint="http://localhost:4000",
context_window=8192,
api_key="your-master-key" # Proxy master key
)

負載平衡與備援

LiteLLM proxy 的負載平衡與備援功能可與 Letta 無縫搭配:

# config.yaml with fallbacks
model_list:
- model_name: gpt-4
litellm_params:
model: openai/gpt-4
api_key: os.environ/OPENAI_API_KEY
tpm: 40000
rpm: 500

- model_name: gpt-4 # Same model name for fallback
litellm_params:
model: azure/gpt-4
api_key: os.environ/AZURE_API_KEY
api_base: os.environ/AZURE_API_BASE
api_version: "2023-07-01-preview"
tpm: 80000
rpm: 800

router_settings:
routing_strategy: "usage-based-routing"
fallbacks: [{"gpt-4": ["azure/gpt-4"]}]

proxy 會為 Letta 透明地處理所有路由、負載平衡與備援。

監控與可觀測性

啟用記錄,以透過 proxy 追蹤您的 Letta 代理程式的 LLM 使用情況:

# config.yaml with logging
model_list:
# ... your models

litellm_settings:
success_callback: ["langfuse"] # or other observability tools

environment_variables:
LANGFUSE_PUBLIC_KEY: "your-key"
LANGFUSE_SECRET_KEY: "your-secret"

在 proxy 儀表板中檢視指標:

# Start proxy with UI
litellm --config config.yaml --port 4000 --detailed_debug

範例:多代理程式系統

import letta
from letta import create_client, LLMConfig

client = create_client()

# Create specialized agents using proxy endpoints
agents = {}

# Research agent using Claude for analysis
agents['researcher'] = client.create_agent(
name="researcher",
system="You are a research specialist. Analyze information thoroughly.",
llm_config=LLMConfig(
model="claude-3-sonnet",
model_endpoint="http://localhost:4000",
model_endpoint_type="openai"
)
)

# Writer agent using GPT-4 for content creation
agents['writer'] = client.create_agent(
name="writer",
system="You are a content writer. Create engaging, well-structured content.",
llm_config=LLMConfig(
model="gpt-4",
model_endpoint="http://localhost:4000",
model_endpoint_type="openai"
)
)

# Coordinator workflow
def research_and_write_workflow(topic: str):
# Research phase
research_response = client.user_message(
agent_id=agents['researcher'].id,
message=f"Research the topic: {topic}. Provide key insights and data."
)

research_results = research_response.messages[-1].text

# Writing phase
write_response = client.user_message(
agent_id=agents['writer'].id,
message=f"Based on this research: {research_results}\n\nWrite an article about {topic}."
)

return write_response.messages[-1].text

# Execute workflow
article = research_and_write_workflow("The future of AI in healthcare")
print(article)

最佳做法

  1. 模型選擇:針對不同任務使用適當的模型:

    • 用 Claude 進行分析與推理
    • 用 GPT-4 進行創意任務
    • 用 GPT-3.5-turbo 處理簡單互動
  2. Proxy 設定

    • 設定適當的速率限制與逾時
    • 使用備援以提升可靠性
    • 為正式環境啟用驗證
  3. 記憶管理:Letta 會自動處理記憶,但在大型上下文時請監控使用量

  4. 成本優化

    • 使用 proxy 的預算功能來控制成本
    • 針對每位使用者/團隊設定速率限制
    • 透過 proxy 儀表板監控 token 使用量
  5. 監控:啟用可觀測性以追蹤代理程式效能與 token 使用量

疑難排解

連線問題

# Test your LiteLLM proxy
curl -X POST http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4",
"messages": [{"role": "user", "content": "Hello"}]
}'

設定除錯

# Enable verbose logging
import logging
logging.basicConfig(level=logging.DEBUG)

# Test Letta configuration
client = create_client()
print(client.get_default_llm_config())

常見 Proxy 問題

  • 連接埠衝突:請確認 4000 連接埠未被使用
  • 找不到模型:驗證模型名稱是否與您的 config.yaml 相符
  • 驗證錯誤:檢查 master key 設定
  • 速率限制:監控 proxy 記錄中是否有觸發 rate limit

資源