Chatbot with Memory
A LangGraph agent that remembers a conversation across turns and processes, with a token budget so it doesn't grow unbounded — thread persistence and trimming applied together rather than explained again here.
from langchain.agents import create_agent
from langchain_core.messages.utils import trim_messages, count_tokens_approximately
from langgraph.checkpoint.sqlite import SqliteSaver
def search_faq(query: str) -> str:
"""Search the product FAQ."""
return f"FAQ result for '{query}': see docs.example.com/faq"
with SqliteSaver.from_conn_string("chat_history.db") as checkpointer:
agent = create_agent(
model="anthropic:claude-sonnet-4-6",
tools=[search_faq],
checkpointer=checkpointer,
pre_model_hook=lambda state: {
"llm_input_messages": trim_messages(
state["messages"],
strategy="last",
token_counter=count_tokens_approximately,
max_tokens=2000,
start_on="human",
end_on=("human", "tool"),
)
},
)
config = {"configurable": {"thread_id": "user_42"}}
agent.invoke(
{"messages": [{"role": "user", "content": "Hi, I'm Ana."}]},
config,
)
agent.invoke(
{"messages": [{"role": "user", "content": "What's my name?"}]},
config,
) # -> "Your name is Ana." — resumed from the same thread_id
Two thread ids per user (one per conversation) keep separate chats from bleeding into each other; a pre_model_hook runs the trim on every turn without touching what the checkpointer actually stores, so the full history is still there if you need to inspect or export it later.
See also
- Thread Persistence — the checkpointer mechanics behind
thread_id. - Trimming and Summarization — the
trim_messagescall used here. - React Agent — the base
create_agentpattern this extends.