Chatbot Backend Tools for the JVM: Spring AI vs LangChain4j
The short answer
For a chatbot backend on the JVM, pick the framework that matches the rest of your stack. As of October 11, 2026:
- Spring AI — if the service is a Spring Boot app. The 2.x line targets Spring Boot 4; 2.1.0-M1 (September 25, 2026) moves the baseline to Spring Boot 4.2 and adds the OpenAI Responses API.
- LangChain4j — if you want framework independence (Quarkus, Micronaut, Helidon, plain Java). Current release 1.22.0, which added OpenAI Decisions API support and Anthropic automatic prompt caching.
- JetBrains Koog — if the “chatbot” is really an agent with long-running state. Koog 1.0 (KotlinConf 2026) promises no breaking changes in stable modules for at least one year.
The comparison
| Spring AI | LangChain4j | Koog | |
|---|---|---|---|
| Maintainer | Spring team (Broadcom) | LangChain4j project (community, Apache 2.0) | JetBrains (open source) |
| Current version | 2.0.x GA; 2.1.0-M1 milestone | 1.22.0 | 1.0 stable core |
| Runs on | Spring Boot 4.x | Any JVM framework or plain Java | JVM, Android, iOS, JS/Wasm via Kotlin Multiplatform |
| Programming model | ChatClient + advisors (memory, RAG, tools) | AI Services: annotated Java interfaces | Agents with tools, strategies and persistence |
| Model providers | Anthropic, OpenAI, Microsoft, Amazon, Google, Ollama and more | 20+ providers | Anthropic, OpenAI, Google, Bedrock, DeepSeek, Mistral, Ollama, OpenRouter |
| Vector stores | PGVector, Redis, Qdrant, Milvus, MongoDB Atlas, Neo4j, Pinecone, Weaviate and more | 30+ embedding stores | Via integrations |
| Observability | Micrometer metrics and traces | Micrometer listeners | OpenTelemetry across all targets; Langfuse adapter |
| Stability promise | Spring release train | Fast cadence, occasional breaking changes | No breaking changes in stable modules for 12+ months |
Picks by situation
An existing Spring Boot service: Spring AI. You add a starter, set a model key in application.yml, and inject a ChatClient. Advisors handle chat memory and retrieval, and output maps to Java records. The 2.1 milestone introduces ordered message parts — text, reasoning, tool calls and media in the order the model returned them — which matters for current reasoning models whose thinking signatures must be replayed unchanged. Only the new OpenAiResponsesChatModel uses parts natively in M1; other providers follow in RC1.
Quarkus, Micronaut or a framework-free service: LangChain4j. You declare interface Assistant { String chat(@UserMessage String msg); } and LangChain4j generates the implementation with memory, tools and RAG wired in. Since 1.20 the return type picks the execution mode — blocking, async, reactive stream or token streaming. Quarkus ships a first-party LangChain4j extension, which makes it the usual pick for native-image, fast-start services. Watch the release notes: version 1.19.1 was published from the wrong branch and LangChain4j tells users on it to move to 1.20.0 or later.
A long-running assistant or agent: Koog. Koog gives you agent persistence (restore state after a crash), retries, history compression to save tokens, and OpenTelemetry traces on every target. It also runs LiteRT models locally on Android. Koog integrates with Spring Boot, so a Spring team can keep Spring AI for simple endpoints and use Koog for the agent.
What a production chatbot backend needs regardless of framework
- Streaming to the client (Server-Sent Events or WebSocket) — all three support token streaming.
- Conversation memory stored outside the JVM (Redis, Postgres) so replicas can scale.
- Retrieval against your own documents; add a reranker for quality (see how to add a reranker).
- Tracing of every model call with tokens and cost — export to an LLM observability tool.
- A cheap routing step for intent or escalation; a decision model can do this without generating text (see decision model APIs compared).
What it costs
The frameworks are free. Cost is the model: a Spring AI or LangChain4j bot on Claude Haiku 5.5 or GPT-6 Luna pays $0.10 per million input tokens and $0.50 per million output tokens, as of October 11, 2026. See our current API prices. TypeScript teams facing the same choice should read Mastra vs LangChain vs Vercel AI SDK.
Last verified: October 11, 2026.