LangChainGuide
Isometric dark scene of a server with a chip sending glowing cyan token cubes through clear tubes via a hub and small nodes to a monitor, like streamed model output
how-to

How to Stream LangChain Responses: Models, Agents, APIs

Stream LangChain output with model.stream(), agent stream_mode messages, and astream_events, then ship tokens over SSE without proxy buffering.

By LangChainGuide Editorial · · 6 min read

Here is how to stream LangChain responses: call model.stream() (or astream()) on a chat model and print each chunk’s .text. For agents built with create_agent, call agent.stream(..., stream_mode="messages") to get (token, metadata) tuples. To serve it, wrap astream() in a FastAPI StreamingResponse with text/event-stream and turn off proxy buffering.

What follows is the operational detail: which method fits which layer, what to measure, and where streaming silently breaks between Python and the browser.

Which LangChain streaming method should you use?

Use stream()/astream() on a bare chat model or LCEL chain, stream_mode="messages" on an agent or LangGraph graph, and astream_events() only when you need typed lifecycle events (on_chat_model_start, on_chat_model_stream, on_chat_model_end) from nested components. Every Runnable exposes stream(), so the choice is about how much structure you need around the tokens.

If you are still deciding between a chain and an agent, LangChain’s building blocks lays out when an agent loop is worth the extra latency.

Streaming a chat model directly

The LangChain models documentation shows the minimal loop. stream() returns an iterator of AIMessageChunk objects instead of the single AIMessage that invoke() returns:

from langchain.chat_models import init_chat_model

model = init_chat_model("openai:gpt-5.4-mini")

for chunk in model.stream("Why do parrots have colorful feathers?"):
    print(chunk.text, end="", flush=True)

Chunks are designed to be summed. If you need the full message afterwards (to append to history, log it, or read tool calls), accumulate as you go rather than making a second call:

full = None
for chunk in model.stream("What color is the sky?"):
    full = chunk if full is None else full + chunk
    print(chunk.text, end="", flush=True)

# full is now usable anywhere an invoke() result would be

For reasoning models and tool calls, iterate chunk.content_blocks and branch on block["type"]: "text", "reasoning", or "tool_call_chunk". Tool-call arguments arrive as partial JSON fragments, so do not try to parse them until the stream ends and the chunks are merged.

The same loop works for local models. ChatOllama implements the same interface, so the setup in LangChain with Ollama streams without code changes.

Chains

An LCEL chain streams if every step can pass chunks through. prompt | model | StrOutputParser() does: the parser emits string fragments as the model produces them.

from langchain_core.prompts import ChatPromptTemplate
from langchain_core.output_parsers import StrOutputParser

prompt = ChatPromptTemplate.from_template("Summarize: {text}")
chain = prompt | model | StrOutputParser()

for piece in chain.stream({"text": doc_text}):
    print(piece, end="", flush=True)

In a RAG chain the retriever step runs to completion first, then generation streams. That means your first token waits on the vector search. If the retrieval half is slow, fix it there; the RAG pipeline setup guide covers the retriever side.

How do you stream an agent in LangChain v1?

Call agent.stream() with stream_mode="messages" for LLM tokens, "updates" for per-step state changes, or "custom" for payloads you emit from tools. Pass a list to get several at once. Agents from create_agent run on LangGraph, so these are the LangGraph stream modes, which also include values, checkpoints, tasks, and debug.

The LangChain streaming docs describe version="v2" as a unified output format, where every chunk is a dict with type, ns, and data keys regardless of mode. That removes the tuple unpacking the default format requires when you combine modes:

from langchain.agents import create_agent
from langgraph.config import get_stream_writer

def get_weather(city: str) -> str:
    """Get weather for a given city."""
    writer = get_stream_writer()
    writer({"status": f"looking up {city}"})
    return f"It's always sunny in {city}!"

agent = create_agent(model="openai:gpt-5.4-mini", tools=[get_weather])

for chunk in agent.stream(
    {"messages": [{"role": "user", "content": "What is the weather in SF?"}]},
    stream_mode=["messages", "updates", "custom"],
    version="v2",
):
    if chunk["type"] == "messages":
        token, metadata = chunk["data"]
        if metadata["langgraph_node"] == "model" and token.text:
            print(token.text, end="", flush=True)
    elif chunk["type"] == "custom":
        print(f"\n[{chunk['data']['status']}]")
    elif chunk["type"] == "updates":
        pass  # step boundaries: log, trace, or drive a progress UI

The langgraph_node filter matters. In messages mode the docs show chunks tagged "model" for generation and "tools" for tool results. Without the filter, raw tool output gets printed into the user’s answer. In custom graphs with several LLM nodes, filter on node name or on metadata["tags"] after attaching tags to a model with init_chat_model(..., tags=["answer"]).

To keep an internal model (a classifier, a router, a sub-agent) from streaming to the client at all, construct it with streaming=False, or disable_streaming=True for integrations that do not accept the streaming parameter. If your agent is looping or failing on tool calls, the stream just shows you the failure faster; agent loop and parsing errors covers the actual fixes.

One async gotcha from the LangGraph docs: on Python below 3.11, pass config explicitly to model.ainvoke(..., config) inside async nodes and take writer: StreamWriter as a parameter instead of calling get_stream_writer(). Otherwise messages mode stays silent.

Wiring it up: serving tokens over HTTP

Streaming inside a Python loop is the easy half. The hard half is getting chunks across ASGI, a reverse proxy, and a load balancer without something buffering them into one blob. Server-sent events (SSE) are the usual transport: MDN’s SSE guide specifies a text/event-stream content type with data: lines, each message terminated by a blank line.

FastAPI’s StreamingResponse accepts an async generator directly:

import json
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse
from langchain.chat_models import init_chat_model

app = FastAPI()
model = init_chat_model("openai:gpt-5.4-mini")

@app.get("/chat")
async def chat(q: str, request: Request):
    async def event_stream():
        async for chunk in model.astream(q):
            if await request.is_disconnected():
                break
            if chunk.text:
                yield f"data: {json.dumps({'token': chunk.text})}\n\n"
        yield "data: [DONE]\n\n"

    return StreamingResponse(
        event_stream(),
        media_type="text/event-stream",
        headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
    )

JSON-encode each token. A raw token containing a newline would otherwise split an SSE frame and corrupt the client’s parse. The is_disconnected() check stops the upstream generation when the user closes the tab, so you stop paying for tokens nobody reads.

The X-Accel-Buffering: no header exists for nginx. Per the nginx proxy module docs, proxy_buffering defaults to on, and a yes/no value in that response header toggles it per response. If you control the nginx config, set it on the location instead:

location /chat {
    proxy_pass http://app:8000;
    proxy_http_version 1.1;
    proxy_buffering off;
    proxy_read_timeout 300s;
}

The metric that matters: time to first token

Once responses stream, total request latency stops describing what users feel. Track time to first token (TTFT): the wall-clock gap between request receipt and the first non-empty chunk leaving your server. Pair it with inter-token latency, (last token time minus first token time) divided by (tokens minus one), which shows whether generation stalls mid-answer.

TTFT beats end-to-end p99 because a long answer that starts almost immediately feels fast, while a short answer behind a long spinner does not. End-to-end latency mostly measures output length.

import time
from prometheus_client import Histogram

TTFT = Histogram(
    "llm_ttft_seconds", "Time to first streamed token",
    ["route"], buckets=(0.1, 0.25, 0.5, 1, 2, 4, 8),
)

async def timed_stream(route: str, q: str):
    start = time.perf_counter()
    first = True
    async for chunk in model.astream(q):
        if first and chunk.text:
            TTFT.labels(route=route).observe(time.perf_counter() - start)
            first = False
        yield chunk

Keep label cardinality to route or model name. Never label by user or thread ID.

What you’ll see

Healthy: server-side TTFT sits well below end-to-end latency, and client-measured TTFT tracks it closely. Broken buffering looks like client TTFT roughly equal to total latency while server TTFT stays low: tokens are leaving Python and piling up in a proxy. A TTFT that climbs with no change in output length usually points upstream, at a slow retriever, a tool call that runs before generation, or provider queueing.

Caveats

  • Buffering outside nginx. CDNs, API gateways, and compression middleware can also hold chunks. Test through the full production path.
  • Connection limits. MDN notes that over HTTP/1.x, browsers cap SSE at six open connections per domain, shared across tabs. Serve over HTTP/2.
  • Structured output. A parser that needs the whole JSON object collapses the stream into one final chunk.
  • Cancellation. FastAPI’s docs warn that an async generator is only cancellable at an await.
  • Mid-stream errors. The HTTP 200 is already sent when a provider fails mid-answer. Emit an explicit error event so the client can tell failure from completion.

FAQ

does langchain invoke support streaming

No, invoke() returns one complete AIMessage after generation finishes. Use stream() or astream() on the same model, chain, or agent to receive AIMessageChunk objects as they are produced. Every Runnable exposes these methods, so switching usually means changing one call and adding a loop. If you still need the final message, sum the chunks rather than making a second request.

how to stream langchain agent output token by token

Call agent.stream() with stream_mode="messages" and iterate the (token, metadata) pairs. Filter on metadata["langgraph_node"] == "model" so tool results do not print into the answer. Passing version="v2" returns uniform dicts with type and data keys, which makes combining several stream modes in a single loop much cleaner.

why is my langchain stream not streaming in production

A reverse proxy or gateway is almost always buffering the response. nginx buffers proxied responses by default, so set proxy_buffering off or send X-Accel-Buffering: no. Also confirm the endpoint returns text/event-stream, that compression middleware skips it, and that your generator yields each chunk instead of collecting them first.

how to get the full message after streaming in langchain

Add the chunks together as they arrive: full = chunk if full is None else full + chunk. LangChain designs AIMessageChunk objects to merge by summation, so the result behaves like an invoke() response. You can append it to conversation history or read merged tool calls without making a second model request.

Sources

  1. LangChain documentation: Streaming
  2. LangChain documentation: Models (streaming section)
  3. LangGraph documentation: Streaming
  4. FastAPI: Custom Response - StreamingResponse
  5. MDN: Using server-sent events
  6. nginx: ngx_http_proxy_module (proxy_buffering)
#langchain #streaming #langgraph #fastapi#server-sent-events

Related