Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To build an AI agent in Java, connect a language model to narrowly defined Java tools, run a loop that executes the model’s tool requests, and return each result to the model until it can answer. Add chat memory, retrieval (RAG), planning, or multiple agents only when the task requires them. Keep credentials and side effects in application code; the model should request an action, never call your databases or services directly.

Two established routes are LangChain4j, a Java-first library with AI Services and an agentic module, and Spring AI, which provides Spring-centric ChatClient, advisors, tool-calling, and MCP APIs. Neither has an evidence-based universal performance or quality advantage. Choose the route that fits the framework your service already uses.

What makes a Java application an agent?

A single prompt followed by one model response is a model integration, not necessarily an agent. An agent can ask the model what to do next, expose application tools, execute a requested tool, feed the result back, and repeat until the model produces a final answer. Memory and planning are optional capabilities.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google Developers Codelabs describes agentic AI as systems in which language models are equipped with “tools, memory, and planning capabilities to autonomously accomplish complex, multi-step goals.” In practice, the important boundary is tool execution: your Java process validates arguments, checks authorization, performs the operation, and controls what data is returned.

Choose a workflow when the path is known

For a fixed sequence such as validate an order, reserve inventory, charge a payment, and send a receipt, ordinary Java orchestration is usually easier to test and audit. Spring AI’s reference guidance notes that workflows often provide better predictability and consistency for well-defined tasks. Let the model fill in limited parameters or classify the request, while code controls the stages.

Choose an agent loop when the path is uncertain

An agent is useful when the model must choose among tools, inspect intermediate results, or decide whether another step is necessary—for example, answering a support question by searching several systems and then composing a response.

The core architecture

  1. Accept a user goal. Define limits such as maximum turns, time, and output size.
  2. Send the goal and tool descriptions to the model. Descriptions should state arguments, units, permissions, and failure behavior.
  3. Inspect the response. It should be either a final message or a structured tool request.
  4. Validate and authorize. Parse arguments into typed Java objects, reject unknown fields, enforce tenant and user scope, and require approval for sensitive actions.
  5. Execute the tool in Java. Apply timeouts, idempotency, rate limits, and logging. Never give the model a database credential or unrestricted HTTP client.
  6. Return a bounded result. Remove secrets and irrelevant fields before sending the result back.
  7. Repeat or finish. Stop on a final answer, a limit, cancellation, or an unrecoverable error.

A useful first agent has one read-only tool. This isolates model behavior from destructive side effects and makes failures diagnosable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prerequisites and a sensible build order

Google Developers Codelabs’ LangChain4j and Google GenAI tutorial lists JDK 17 or later, Maven 3.5 or later, and a Gemini API key. Those requirements apply to that tutorial, not to every Java agent framework. Keep the key in an environment variable or secret manager, never in source control.

  1. Make one ordinary model call and log request IDs, latency, and token usage when the provider exposes them.
  2. Add one narrow, read-only tool and verify that invalid arguments are rejected.
  3. Return a typed result rather than an unstructured paragraph.
  4. Implement or enable the tool-request/result loop.
  5. Add memory only if a later turn genuinely needs earlier context.
  6. Add retrieval, parallel work, or multiple agents only for a demonstrated requirement.

A runnable Java agent loop without a framework

The following program is a complete, deterministic demonstration of the control loop. It uses a fake model so you can run it without a provider account, then replace the DemoModel implementation with a LangChain4j or Spring AI client. The model requests a weather tool, Java executes it, and the returned result produces a final answer.

import java.util.*;

public class JavaAgentDemo {
  record ToolRequest(String name, Map<String, String> arguments) {}
  record ModelReply(String text, ToolRequest request) {
    static ModelReply tool(ToolRequest r) { return new ModelReply(null, r); }
    static ModelReply done(String s) { return new ModelReply(s, null); }
  }

  interface Model {
    ModelReply next(List<String> transcript, Set<String> tools);
  }

  interface Tool {
    String call(Map<String, String> arguments);
  }

  static final class DemoModel implements Model {
    private int turn;
    public ModelReply next(List<String> transcript, Set<String> tools) {
      if (turn++ == 0) {
        return ModelReply.tool(new ToolRequest("weather",
            Map.of("city", "Berlin")));
      }
      return ModelReply.done("The weather tool returned: " + transcript.get(transcript.size() - 1));
    }
  }

  static String run(String goal, Model model, Map<String, Tool> registry,
                    int maxTurns) {
    List<String> transcript = new ArrayList<>();
    transcript.add("USER: " + goal);
    for (int turn = 0; turn < maxTurns; turn++) {
      ModelReply reply = model.next(transcript, registry.keySet());
      if (reply.request() == null) return reply.text();
      Tool tool = registry.get(reply.request().name());
      if (tool == null) throw new IllegalArgumentException("Unknown tool");
      String result = tool.call(reply.request().arguments());
      transcript.add("TOOL " + reply.request().name() + ": " + result);
    }
    throw new IllegalStateException("Agent turn limit reached");
  }

  public static void main(String[] args) {
    Map<String, Tool> tools = Map.of(
      "weather", a -> "18 C, cloudy (sample data)"
    );
    System.out.println(run("What is the weather in Berlin?",
        new DemoModel(), tools, 4));
  }
}

Compile and run it with javac JavaAgentDemo.java && java JavaAgentDemo. In production, replace the demo model with a provider client, use a structured response format for tool calls, and treat malformed model output as an error rather than guessing.

Using LangChain4j

LangChain4j offers low-level primitives and AI Services: Java interfaces implemented through proxies. AI Services can format inputs, parse outputs, retain chat memory, expose Java methods as tools, and connect retrieval components. Its agentic module documents sequential and other workflows and uses an AgenticScope to share outputs; that state is transient unless you configure persistence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A typical shape is:

interface SupportAssistant {
  String answer(String question);
}

class SupportTools {
  // Annotate this method with LangChain4j's current @Tool annotation.
  public String findOrder(String orderId) {
    // Validate tenant and authorization before querying.
    return "status returned by application code";
  }
}

// Build a ChatLanguageModel with your chosen provider and API key.
// Then create an AI Service, register SupportTools, and call answer(...).
// Pin dependency versions from the current LangChain4j documentation.

The exact builder and provider modules change over time, so pin the version you select and follow that version’s AI Services and agentic documentation. LangChain4j labels Chains as legacy and says it does not plan to add more Chains; new code should use AI Services or the agentic abstractions instead.

LangChain4j design choices

  • Use AI Services for a concise interface around one assistant and its tools.
  • Use the agentic module when you need explicit sequential, parallel, or goal-oriented orchestration.
  • Use ChatMemory for conversational continuity, and configure persistence if a conversation must survive a process restart.
  • Use its embedding and retrieval integrations only when answers must be grounded in a private corpus.
  • Use its MCP tool-agent support when tools are hosted by an MCP server.

Using Spring AI

Spring AI is a natural fit for a Spring Boot service. ChatClient and Advisors compose model calls, memory, retrieval, and tools. In Spring AI 2.0.1, the tool loop is driven through the advisor chain: the model requests a tool, application code invokes the callback, the result is sent back, and the cycle continues until there are no more tool calls. Calling ChatModel directly does not automatically execute that loop.

@Component
class AccountTools {
  @Tool(description = "Read the current account balance for the authenticated user")
  public Balance balance() {
    // Obtain identity from the server-side security context, not model input.
    return accountService.currentBalance();
  }
}

@Service
class AssistantService {
  private final ChatClient client;

  AssistantService(ChatClient.Builder builder, AccountTools tools) {
    this.client = builder
      // Register the tool and the tool-calling advisor according to
      // the Spring AI version pinned by your project.
      .defaultTools(tools)
      .build();
  }

  String answer(String question) {
    return client.prompt().user(question).call().content();
  }
}

Check the versioned Spring AI API before compiling: 1.x and 2.0 APIs are not interchangeable. Keep tool implementations in ordinary Spring beans so transactions, authorization, timeouts, and observability remain under your control.

Memory, retrieval, planning, and MCP

Conversation memory

Memory preserves prior turns, but it also increases context size and can retain sensitive data. Assign a conversation or user identifier, cap the history, summarize old turns, and define deletion and retention rules. Do not treat transient agent state as durable storage.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

RAG for private knowledge

Use retrieval-augmented generation when the answer depends on documents or records not present in the model’s training data. Retrieve only data the current principal may read, include source identifiers in the internal result, and reject requests that return no trustworthy evidence instead of allowing the model to invent one.

Planning and multiple agents

Planning can split a complex objective into subtasks, while multiple agents can separate roles such as research and verification. Both add coordination, latency, state, and failure modes. Start with one agent and code-defined stages; introduce delegation only when a measured requirement justifies it.

MCP interoperability

MCP can let a Java application consume tools from an MCP server or expose Spring services for other clients. LangChain4j documents wrapping MCP tools for agentic systems, while Spring AI provides APIs for consuming servers and exposing services. Treat an external MCP server like any other dependency: authenticate it, restrict available tools, validate returned data, and set network and execution timeouts.

Security and reliability checklist

  • Keep provider keys, database credentials, and signing secrets outside prompts and source files.
  • Authorize every tool call against the authenticated user, tenant, and resource.
  • Use allow-lists for tool names, argument schemas, domains, file paths, and SQL operations.
  • Require human approval for payments, deletions, account changes, messages, or other irreversible effects.
  • Make write tools idempotent and attach an operation ID so retries cannot duplicate an action.
  • Set maximum turns, wall-clock time, tool calls, response size, and spend where the provider supports budgets.
  • Redact secrets and personal data from prompts, tool results, traces, and error messages.
  • Log each model request, tool request, validation decision, result status, and final response with correlation IDs.
  • Test prompt injection, malformed JSON, tool timeouts, duplicate requests, partial outages, and cancellation.

Performance and cost considerations

Every extra model turn adds latency and usually consumes more tokens. Keep tool descriptions short, return compact structured results, avoid sending entire documents when a filtered slice is enough, and parallelize independent read-only calls in code. Cache stable lookups with an explicit time-to-live, but never cache data across tenants without a security boundary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Measure end-to-end latency, model latency, tool latency, turn count, input and output tokens, error rate, and human-approval wait time. The available documentation does not establish a universal benchmark for LangChain4j versus Spring AI, nor does it establish a production-readiness guarantee or model-price comparison; evaluate your own workload.

Troubleshooting common failures

Symptom Likely cause Fix
The model describes a tool instead of calling it Tool schema is missing, malformed, or not registered. Inspect the outgoing tool definitions, use typed arguments, and verify the framework’s tool annotation and registration path.
Tool calls repeat forever The result does not answer the model’s request or the loop has no stop condition. Return a clear success or failure object, detect duplicate calls, and enforce a maximum-turn limit.
Spring application returns after the first request ChatModel was called directly, bypassing the tool-calling advisor. Use ChatClient with the version-appropriate tool-calling advisor configuration.
Arguments fail to parse Free-form text or a schema mismatch. Use strict DTOs, reject unknown fields, and send a structured validation error back for one retry.
Answers ignore private documents Retrieval was not invoked or results exceeded the context budget. Log retrieved IDs, reduce and rank the result set, and require citations or an explicit “not found” response.
Requests become slow or expensive Too many turns, oversized memory, or serial independent tools. Cap history and turns, summarize, cache safe reads, and run independent reads concurrently.
A restarted process loses the conversation Only in-memory state was configured. Persist chat memory and AgenticScope data with an appropriate retention policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If one of your Java agent’s tools needs a current webpage image or PDF, ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Java can call the same endpoint with the standard HTTP client:

import java.net.URI;
import java.net.http.*;
import java.nio.file.*;

public class ScreenshotNeoCall {
  public static void main(String[] args) throws Exception {
    String key = System.getenv("SCREENSHOTNEO_API_KEY");
    String url = "https://stripe.com";
    String requestUrl = "https://api.screenshotneo.com/v1/shot?access_key="
        + java.net.URLEncoder.encode(key, java.nio.charset.StandardCharsets.UTF_8)
        + "&url=" + java.net.URLEncoder.encode(url, java.nio.charset.StandardCharsets.UTF_8);
    HttpResponse<byte[]> r = HttpClient.newHttpClient().send(
        HttpRequest.newBuilder(URI.create(requestUrl)).GET().build(),
        HttpResponse.BodyHandlers.ofByteArray());
    if (r.statusCode() / 100 != 2) throw new IllegalStateException("HTTP " + r.statusCode());
    Files.write(Path.of("shot.webp"), r.body());
  }
}

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for all 63 options, including full-page lazy-image loading, CSS-selector capture, device and retina settings, PDF paper sizes and page ranges, custom CSS and JavaScript, click and wait conditions, request blocking, headers and cookies, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data, and the OpenAPI specification. Every feature is on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots, with yearly billing providing two months free. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Does an agent require memory?

No. Tool selection and iteration define the practical agent loop; memory is an optional capability for preserving context between turns.

Should every task use an autonomous agent?

No. Use a code-defined workflow when stages and rules are known. Use dynamic tool selection when the path cannot be reliably predetermined.

Can the model execute Java methods directly?

No. The model can request a declared tool. Your Java application validates and executes the corresponding method.

When should I use MCP?

Use MCP when tools must be shared across clients or hosted separately from the Java service. Apply the same authentication, authorization, validation, and timeout controls as for local tools.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is LangChain4j faster than Spring AI?

The available official documentation does not provide a controlled comparison. Benchmark both with your model, prompts, tools, and workload if latency matters.

Frequently Asked Questions

What is the smallest useful Java agent?

One model, one narrowly scoped read-only tool, a bounded request/result loop, and a final response are enough to demonstrate agent behavior.

How do I prevent duplicate side effects?

Require authorization, use idempotency keys, persist operation status, and add human approval for irreversible actions.

Can I add retrieval later?

Yes. Start with the tool loop, then add embeddings and a vector store when answers need private-corpus grounding.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.