LangGraph Production Playbook: Prototype to Production-Grade in 2026
Key Takeaways
- →Six pillars separate production from prototype: checkpointing, evals, observability, guardrails, failover and SLOs.
- →Moving from SQLite to PostgresSaver is the line between a demo and a system that survives a worker restart.
- →Every prompt, model, tool or graph change runs against a golden eval set of 100-300 cases before merge.
- →Budget five cost drivers separately: inference, checkpoint storage, observability, guardrail inference and eval overhead.
- →The rollout runs in 90 days: foundation, agent v1, guardrails and failover, then production hardening.

On this page⌄
Running LangGraph in production means six things a prototype never needs: durable checkpointing, a continuous eval framework, full observability, layered guardrails, failover, and measurable SLOs. A LangGraph prototype takes a weekend to build. A production system that survives real traffic, real failure modes, and a real audit takes weeks, not days, and most teams underestimate that gap and ship the prototype anyway, then spend the next quarter firefighting hallucinations, silent failures, runaway token spend, and customer-trust incidents.
This is the playbook worth having before you ship production agent systems in regulated domains like healthcare, legal, finance, insurance, and real estate. Six engineering pillars, the 2026 stack we deploy against each, the architectural patterns that survive production traffic, and the gotchas that are not in the official docs. If you want an outside check on where your own build stands against these six pillars before you ship, that is what a production readiness review is for.
If you are still picking a framework, our agent framework comparison covers why LangGraph wins over CrewAI and AutoGen, now in maintenance mode with Microsoft Agent Framework as its stated successor, for the production tier, and how it stacks against the OpenAI Agents SDK, Claude Agent SDK, and Microsoft Agent Framework. This guide assumes you already chose LangGraph. The question now is how to operate it.
Why LangGraph specifically, and when not
LangGraph is the right choice when your agent has any of these requirements:
- Multi-step workflows with branching logic and conditional retries
- Durable state that must survive process restarts and worker scaling events
- Human-in-the-loop approval checkpoints (medical sign-off, contract approval, compliance review)
- Multi-vendor LLM routing (Claude for reasoning, GPT-5.6 for tool use, an open-weight model such as Qwen3.8 Max or GLM-5.2 for self-hosted bulk)
- Per-tenant isolation with auditable per-tenant configuration
- Audit-trail requirements (EU AI Act Article 12, HIPAA, NY DFS, SOX)
LangGraph is the wrong choice when:
- The agent is a one-shot prompt or a 2-3 step chain that does not need state. LangChain expression language or a vanilla Python script is fine.
- You are single-vendor and committed (Claude only, or GPT only). The Claude Agent SDK or OpenAI Agents SDK ships natively integrated observability and may be cleaner.
- The team has zero Python depth. LangGraph's superpower is its programmability. If you cannot write Python and read tracebacks, you will struggle in production.
Assuming you are still here, the six pillars.
Pillar 1: durable state with checkpointing
The single feature that separates LangGraph from a glorified prompt chain is its checkpointer. Every state mutation is persisted to a backing store (Postgres, Redis, SQLite for prototypes). If a worker crashes mid-execution, the next worker resumes from the last checkpoint. If a human-in-the-loop step pauses for hours, the state survives the pause.
A minimal production checkpointer setup looks like this:
```python
from langgraph.checkpoint.postgres import PostgresSaver
with PostgresSaver.fromconnstring(DATABASE_URL) as checkpointer:
checkpointer.setup()
graph = builder.compile(checkpointer=checkpointer)
```
That single change, from the in-memory or SQLite checkpointer to a Postgres-backed one, is the line between a demo and something that survives a worker restart.
What to use
- Prototyping: SQLite checkpointer. Ships in the library, zero setup.
- Production single-tenant: PostgresSaver. Use your existing managed Postgres (RDS, Cloud SQL, Supabase). Plan for roughly 1KB-10KB per checkpoint, 10-50 checkpoints per agent run.
- Production multi-tenant or high-throughput: PostgresSaver with per-tenant database (cleanest) or per-tenant schema with row-level security. Redis as an L1 cache in front for hot reads.
- High-availability: Postgres with read replicas for the checkpoint reads, primary for writes. Avoid SQLite in production, full stop.
The gotchas
- State size grows quickly. Every step appends to the conversation history by default. After 50 steps your checkpoint blob is 100KB-1MB. Use the `summarize` pattern to compress old turns, or split long-running agents into sub-graphs with their own checkpoint scope.
- Checkpoint cleanup is your problem. LangGraph does not garbage-collect old checkpoints. Write a scheduled job to archive checkpoints older than your retention window (typically 90 days for ops, longer for audit trails).
- Migration is painful. State schema changes (you add a new field, change a tool signature) break in-flight agents. Version your state schema, write migrators, and run them before deploying breaking changes.
Pillar 2: evals before every prompt or model change
The rule: every change to a prompt, model, tool signature, or graph structure runs against a golden eval set before merge. Without this, you have no way to know whether a change improved or regressed the system. "It works on my example" is not eval data.
The eval framework
Golden dataset. 100-300 representative cases covering happy paths, edge cases, and known failure modes. Hand-curated initially, then grown by adding a case for every production incident. Stored in version control alongside the agent code.
Scoring rubric. 3-5 criteria specific to your use case: factual accuracy, format compliance, tone match, tool-call correctness, citation quality. Each scored 1-5.
Judge. Three options. Human review of every eval run is the gold standard but slow and expensive. LLM-as-judge using a different vendor than the system under test (Claude judging GPT output or vice versa) is fast but carries same-family bias. A hybrid, LLM-as-judge for fast iteration plus a human spot-check of a sample of cases each release, is a common compromise.
Tooling. Promptfoo for lightweight YAML-defined evals, LangSmith for evals integrated with LangGraph tracing, or OpenAI Evals if you are OpenAI-primary. Promptfoo for fast prompt iteration, LangSmith for end-to-end graph evals against production traces, is a reasonable default pairing.
The gotchas
- LLM-as-judge is not free. Budget the judge token cost separately, and do the arithmetic rather than guessing. Multiply the number of cases by the input and output tokens per judgment and by your judge model's per-token price; a longer case or a chain-of-thought rubric pushes that further. Check your judge model's current pricing before committing to running it on every merge.
- Eval bias is real. LLMs prefer outputs from the same family. Cross-family judging (Claude judging GPT, Gemini judging Claude) reduces bias but does not eliminate it. Human spot-check is the calibration.
- Regression detection requires a baseline. Score your current production agent against the eval set today. Every future change is delta-vs-baseline. Without a baseline, "the new version is good" is just an assertion.
- Test the tool calls, not just the text. For agents that call functions, score whether the right tool was called with the right arguments, not whether the final natural-language output looks good.
Pillar 3: observability and audit trail
If you cannot replay a production failure, you cannot fix it. Observability is the difference between an outage you debug and an outage you guess about.
The 2026 stack
Langfuse. A common default for production LangGraph. Captures every node execution, every LLM call, every tool call, every state mutation. Self-hostable. The trace UI is built for agentic workloads, not generic spans. Integrates with LangGraph via a single callback handler.
Helicone. Strong if you are OpenAI-heavy. Proxy-based, so it captures every LLM call regardless of orchestration layer. Less LangGraph-native than Langfuse, but its async-by-default model has lower latency overhead.
OpenTelemetry GenAI Semantic Conventions. The emerging standard. Your traces emit standard span attributes that any OTel-compatible backend (Honeycomb, Datadog, Grafana Tempo) can consume. Use this if you already run a central observability stack and want LangGraph traces in the same pane of glass.
What to capture
At minimum, for every agent run:
- Input payload (with PII redaction policy applied)
- Every LLM call with model, prompt, temperature, response, latency, token counts
- Every tool call with name, arguments, response, duration
- Every state mutation
- Every guardrail decision (pass, warn, block)
- Every retry and the reason
- Every human intervention or override
- Final output
For EU AI Act Article 12 compliance, these traces double as the audit log. Store them in tamper-evident storage (S3 Object Lock, Azure immutable Blob) for the retention window your regulator requires.
The gotchas
- Async-by-default is non-negotiable. Synchronous trace export adds real latency per call, on the order of tens to hundreds of milliseconds. For real-time chat or voice agents this is a deal-breaker. Use async exporters.
- PII redaction must happen pre-trace. If user PII flows into your trace store, that store becomes a regulated data system. Redact before the trace handler, not after.
- Sampling at high volume. Capture every failure, every human-flagged output, and a small sample of successful runs. Storing every trace for a high-volume system is expensive and rarely useful.
Pillar 4: guardrails
Guardrails are the layer between the LLM's output and your downstream system. They catch prompt injection, sensitive-data exfiltration, jailbreaks, factual hallucinations, and policy violations.
The layers
Input guardrails. Run on user input before it reaches the agent. Detect prompt injection (instructions to ignore prior context, prompt-leak attempts), jailbreaks, PII you do not want in the agent's context.
Output guardrails. Run on agent output before it ships to the user or downstream system. Detect policy violations, factual inconsistency with retrieved context (for RAG agents), tone violations, sensitive-data exfiltration, hallucinated tool calls.
Tool-call guardrails. Run between an LLM's tool-call decision and the actual tool execution. Validate arguments against schemas, check authorization for the requested resource, enforce rate limits, sandbox dangerous tools (filesystem access, code execution).
What to use
- NeMo Guardrails (NVIDIA). Flexible policy-as-code framework. Best for complex multi-rule policies.
- Llama Guard 4 (Meta). Open-weight, 12 billion parameters, natively multimodal, and consolidates the earlier Llama Guard 3 text (8B) and vision (11B) variants into one classifier that scores text and image content together (huggingface.co/meta-llama/Llama-Guard-4-12B). Strong baseline for content-policy enforcement. Pair it with Prompt Guard 2 for injection and jailbreak detection specifically, and LlamaFirewall if you want an orchestration layer over both.
- Lakera Guard. Managed prompt-injection and jailbreak detection. Faster to deploy than self-hosted, lower flexibility.
- Custom Pydantic validators. For structured output (JSON schema enforcement, type coercion, range checks). Cheap, fast, no LLM involvement.
The gotchas
- Guardrail latency adds up. Each guardrail adds real latency, roughly tens to a few hundred milliseconds per check. Three layers compound. Run input and output guardrails in parallel where logically safe. Reserve the slowest guardrails for high-risk operations.
- Guardrails fail closed for safety, open for availability. Decide the policy explicitly and test it. A guardrail timeout that blocks a valid request looks like an outage to your customer.
- Test guardrails against red-team data, not happy-path data. Use the OWASP Top 10 for LLM Applications as your red-team checklist. In the current published ranking, prompt injection (LLM01) and sensitive information disclosure (LLM02) hold the top two slots (genai.owasp.org/llm-top-10, checked September 2026). Excessive agency (LLM06) sits further down the list but matters disproportionately for anyone reading a LangGraph playbook, because agent frameworks are exactly what hands a model real tools to misuse. Vector and embedding weaknesses (LLM08) matters directly if your agent is RAG-backed. If your agent accepts a file upload, treat cross-modal injection (instructions hidden inside an image or an audio track, where a text filter never looks) as in scope even though it is not yet its own numbered entry.
- Guardrail eval is its own eval set. Your accuracy eval and your guardrail eval are different datasets. Mix them and you cannot tell whether a regression is a model issue or a guardrail issue.
Pillar 5: failover and resilience
LLMs fail. APIs go down. Rate limits trigger. Tool calls time out. Production agents must degrade gracefully, not crash.
The failover patterns
Model fallback ladder. Primary model (a frontier reasoning model for high-stakes decisions) falls back to a faster mid-tier model, then to a self-hosted open-weight model for bulk or degraded-mode traffic. Trigger conditions: rate limit, timeout above N seconds, error response. Each fallback step is logged for post-incident analysis.
Tool-call retry with exponential backoff. Idempotent tools (read operations) can retry aggressively. Non-idempotent tools (writes, payments, sends) need idempotency keys plus a careful retry policy. Without idempotency keys, you risk double-charging customers, double-sending emails, double-writing records.
Circuit breakers. When a downstream tool fails repeatedly, open the circuit and skip the tool for a cool-down period. Surface the degraded state to the agent's reasoning loop so it can decide what to do without the tool's data.
Deterministic fallback paths. For critical workflows, design a non-LLM fallback. The LLM-driven path is the happy path. A rule-based path handles the case where the LLM is unavailable or the eval set indicates the model is unreliable on this input. This is also the EU AI Act human-oversight obligation (Article 14) in practice.
The gotchas
- "Just retry" creates cost spirals. Aggressive retries on transient errors can multiply your token bill many times over in an incident. Cap retries per request and per agent run.
- Fallback model drift. If you fall back to a smaller model, the prompt that worked on your primary model may not work on the fallback the same way. Test the fallback in your eval set, not just in production when it is already too late.
- Silent failure is the worst failure. A guardrail that blocks output silently looks like a successful run with no response. Always log every block decision and surface it in observability.
Pillar 6: SLAs and SLOs
The step from "we ship an agent" to "we sell a managed agent service" is the SLA. Pick a small number of measurable commitments, define the math, monitor in production, and report monthly.
The four SLOs to track
Recommended starting targets for the four SLOs below; adjust them to your own risk tolerance.
- Availability. Percentage of agent runs that complete (success or graceful failure, not crash). Target: 99.5% for most production tiers, 99.9% for tier-1 customer-facing.
- Latency. P50, P95, P99 of end-to-end agent run time. Target: P95 under 5 seconds for chat-style agents, P95 under 30 seconds for complex multi-step agents. Measure on the user-perceived clock, not just the LLM-call clock.
- Accuracy. Percentage of agent outputs that pass the eval scoring rubric. Run a continuous sample of production outputs through the eval pipeline. Target: typically 90-95% depending on use-case sensitivity.
- Hallucination rate. For RAG agents, percentage of outputs that contradict the retrieved context. Measured via citation-grounding eval. Target: under 5% for most workloads, under 1% for high-stakes (legal, medical).
Monthly reporting
The customer-facing artifact is a one-page SLA report. Each SLO with the target, the actual, the breach windows (with root-cause one-liners), and the remediation actions. This is what differentiates a managed AI service from a software license.
The gotchas
- Accuracy SLOs are expensive to monitor. Continuous LLM-as-judge sampling on production output adds cost. Budget a real percentage of your inference cost for eval overhead rather than treating it as free.
- Latency SLOs trip on guardrails first. Long-tail latency in production is usually a slow guardrail call or a slow tool call, not the LLM. Trace breakdown is essential.
- SLA breaches need a remediation budget. Decide in advance whether breaches credit the customer's invoice or trigger an incident review. Without policy, every breach becomes a one-off negotiation.
Deployment options for LangGraph in 2026
LangGraph Platform, managed (billed through LangSmith). Built by the LangChain team. Hosted runtime, hosted checkpointer, integrated LangSmith tracing. LangChain renamed the managed deployment product to LangSmith Deployment after the 1.0 release in October 2025, so both names appear in the docs. Published pricing as of September 2026 (langchain.com/pricing): Developer at $0 for a single seat with up to 5k base traces a month, Plus at $39 per seat per month with up to 10k base traces and one free small serverless deployment, Enterprise on custom terms with self-hosted and hybrid options. Compute is metered on top at $1.50 per LangChain Compute Unit and $1.00 per LangChain Storage Unit, which is where a production bill actually comes from. The open-source framework itself stays MIT-licensed with no usage limit.
Self-hosted on Kubernetes. Run the LangGraph runtime as your own service. Bring your own Postgres for checkpointing, your own Langfuse for observability, your own scaling and security policies. Best for teams with regulated workloads, custom compliance requirements, or significant existing Kubernetes investment.
Serverless (AWS Lambda, Cloud Run, Azure Functions). LangGraph runs cleanly in serverless if your agents are short-lived (under 15-minute execution caps) and stateless between runs (state in external Postgres). Best for bursty workloads where you do not want idle cost.
Hybrid. Tier-1 customer-facing agents on managed cloud for the SLA. Internal agents on self-hosted Kubernetes for cost control. Both share the same Langfuse and Postgres backends. This is the pattern most production-scale deployments converge to.
For a deeper take on private deployment specifically, our private LLM infrastructure guide covers the on-prem option.
What drives the cost at production scale
For a single tier-1 production agent, the cost has five real drivers, each worth budgeting separately rather than lumping into one number: LLM inference (scales with model choice and prompt size), checkpoint storage (modest on managed Postgres), observability (self-hosted Langfuse is cheaper than SaaS-tier, but neither is free at volume), guardrail inference (Llama Guard 4 or a managed service, plus any custom validators), and eval overhead (continuous LLM-as-judge sampling, typically a meaningful percentage on top of inference cost). Engineering ops time for monitoring and incident response is the driver most teams forget to budget until the first incident.
Our AI agents and workflow automation services can scope a build against these six pillars for a specific use case; pricing depends on the agent's scope, compliance requirements, and how much of the platform infrastructure (checkpointing, observability, guardrails) already exists versus needs to be built from scratch.
The 90-day production rollout plan
If you are starting today and want a tier-1 production agent live in 90 days, this is the sequence to run.
Days 1-14: foundation. Pick the agent's bounded scope. Build the golden eval set (start with 50 cases, grow to 200). Stand up checkpointer (Postgres) and observability (Langfuse). Define the SLOs.
Days 15-30: agent v1. Build the LangGraph graph. Wire the eval pipeline. Hit baseline accuracy on the eval set. Ship to internal users only.
Days 31-60: guardrails and failover. Add input, output, and tool-call guardrails. Add a model fallback ladder. Add a deterministic fallback path for critical workflows. Red-team against the OWASP LLM Top 10. Run accuracy regression evals.
Days 61-90: production hardening. Ship to a small share of traffic via feature flag. Monitor SLOs. Iterate on observed failure modes. Ramp to all traffic over weeks 11-12. Publish the first monthly SLA report.
This is aggressive but doable for a focused team. The teams that miss are the ones that try to build all six pillars in parallel without sequencing. Foundation first, agent v1 second, hardening third.
When LangGraph is not enough
Three scenarios where LangGraph alone leaves gaps:
Voice agents with sub-300ms latency requirements. LangGraph's overhead is small but real. For real-time voice, layer a voice-specific framework (Vapi, ElevenLabs Conversational AI) and use LangGraph for the off-path reasoning that does not block the voice turn.
Massive parallelism (1M-plus concurrent agents). LangGraph's checkpointer model assumes per-agent state. At true massive scale, you may want a custom event-sourced architecture where state is reconstructed on demand. Most teams do not need this.
Heavy multi-modal. Image, video, and audio agentic workflows can be modeled in LangGraph, but you will often want a multi-modal-native framework for the modality-specific paths. LangGraph still drives the top-level planning.
Keep reading
For the framework decision, see multi-agent systems for our take on LangGraph vs OpenAI Agents SDK vs Claude Agent SDK vs Microsoft Agent Framework. For the model selection inside your LangGraph nodes, see our LLM decision framework. For RAG-specific patterns, see our enterprise RAG pipeline guide. And when you are ready to scope a tier-1 production build, let us talk.
Sources
- LangSmith Pricing (LangChain, official) (primary source, checked 2026-09-14)
Frequently Asked Questions
Should I use LangGraph or CrewAI for production?+
What is the difference between LangGraph and LangChain?+
How do I deploy LangGraph in production?+
What is the right checkpointer for production LangGraph?+
What evals should I run before every prompt or model change?+
How do I add guardrails to a LangGraph agent?+
What SLAs should a managed LangGraph agent service offer?+
How long does a production LangGraph rollout take?+
When should I not use LangGraph?+
Need a tier-1 production agent built on LangGraph with evals, observability, guardrails, and SLAs? Let us scope it.
Explore AI Agent ServicesAbout the Author

Rajat Gautam
AI Engineer and Consultant
My work goes far beyond recommending tools - I design AI systems that integrate directly into your workflows, eliminate inefficiencies, and deliver measurable business impact. Every solution I build is tailored, practical, and built with long-term scalability in mind.
Need help with this?
Related Topics
Related Articles



Ready to transform your business with AI? Let's talk strategy.
Book a Free Strategy Call