October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Mountain View desk4 min

Gemini Interactions API in TypeScript: Task-Aware Thinking Routing

Task-aware thinking routing is application logic: classify a request, then pass a supported value through generation_config.thinking_level in the Gemini Interactions API.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To route Gemini requests by task complexity in TypeScript, classify the task in your application and pass a model-supported value through generation_config.thinking_level when calling client.interactions.create(). The Interactions API exposes the per-request thinking control; the documented API does not automatically classify tasks or dispatch them to a thinking level.

What task-aware thinking routing means

Task-aware routing is an application policy: your code decides whether a request needs relatively little, moderate, or substantial reasoning, then selects a thinking level supported by the model you are using. The API parameter sets the level; your application supplies the classification and decision.

Google describes the Interactions API as generally available as of June 2026 and recommends it for new projects. It provides a unified interface for Gemini models and agents, including text, multimodal input, tool orchestration, and agentic workflows. See Google’s Interactions API documentation.

Set the thinking level in TypeScript

Install and use Google’s JavaScript/TypeScript SDK, @google/genai. The configuration property is spelled thinking_level in snake case, including in TypeScript.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import { GoogleGenAI } from "@google/genai";

const client = new GoogleGenAI({});

type TaskClass = "simple" | "standard" | "complex";
type ThinkingLevel = "low" | "medium" | "high";

function chooseThinkingLevel(task: TaskClass): ThinkingLevel {
  if (task === "simple") return "low";
  if (task === "complex") return "high";
  return "medium";
}

async function runTask(task: TaskClass, input: string) {
  const interaction = await client.interactions.create({
    model: "gemini-3.8-flash",
    input,
    generation_config: {
      thinking_level: chooseThinkingLevel(task),
    },
  });

  return interaction.output_text;
}

const answer = await runTask("standard", "Summarize the supplied material.");
console.log(answer);

The mapping in this example illustrates how to make a policy explicit; it is not a universal recommendation that these levels are optimal for every model or workload. Google’s guide documents model-specific defaults and supported values. Check the current guide for the model you actually deploy before sending a level, and handle rejected or unavailable model/configuration combinations. See Google’s thinking documentation.

Design a useful routing policy

Choose task categories based on what the request needs, not simply on its length. A long extraction task may be straightforward, while a short question may require multi-step analysis. Treat thinking level as one routing decision among several: the selected model, latency budget, and consequences of an incomplete answer also matter.

Rank #2
TypeScript Programming Language - Software Engineer & Coder T-Shirt
  • TypeScript implements a superset of syntax for strictly typed development, facilitating deep static analysis and enhanced development environment integration. The compiler translates source into standard script formats, ensuring parity across any runtime.
  • TypeScript is ideal for front-end developers, full-stack engineers, and software architects who build large-scale web applications. It serves those looking to improve code excellence, reduce bugs through static checking, and maintain complex projects more.
  • Lightweight, Classic fit, Double-needle sleeve and bottom hem
  • Simple: Consider a lower level for routine transformations or direct retrieval when the task has little ambiguity.
  • Standard: Use a middle policy category for ordinary requests that need some synthesis but do not warrant the most reasoning effort.
  • Complex: Consider a higher level for tasks involving multiple constraints, substantial synthesis, or difficult reasoning.

These are policy examples, not measured performance claims. The cited documentation establishes the configuration mechanism and model-specific supported levels; it does not establish that a particular level is fastest, cheapest, or best for a given workload. Measure your own requests if those trade-offs determine the route.

Check model compatibility before routing

Thinking-level names and defaults are not portable assumptions across all Gemini models. Maintain the model ID and its allowed values together in configuration, rather than letting a router emit a level without regard to the destination model. When changing models, recheck the documented default and valid values, and test error handling for unsupported combinations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical decision record can include the task class, chosen model, chosen level, and whether the interaction completed. This makes it possible to adjust the policy when real traffic shows that a category is being routed poorly. Avoid claiming a benchmark or cost reduction unless you have measured it under your own conditions.

Account for output-token limits

max_output_tokens counts thinking tokens as well as answer tokens. If reasoning consumes the available ceiling, an interaction may finish with status incomplete and a truncated or empty answer. Google advises lowering thinking_level to reduce cost or latency rather than setting an artificially small output cap when avoiding truncation matters. See Google’s thinking documentation.

Set an output ceiling appropriate to the response you need, then inspect completion status and handle incomplete results explicitly. Do not assume that a non-throwing request produced a complete, usable answer.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose stateful or stateless routing across turns

The Interactions API stores requests by default to support server-side conversation state. For a follow-up turn, pass the earlier interaction’s ID as previous_interaction_id. Set store: false for stateless operation. See Google’s Interactions API documentation.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decide whether a multi-turn policy should keep the same model and thinking level across a conversation or reassess each turn. If you use stateless requests, your application must manage any context needed for continuity. This is a routing design choice, not an automatic behavior supplied by task classification.

Inspect steps without mistaking thoughts for answers

The TypeScript example in Google’s documentation iterates over interaction.steps and checks whether a thought step has a summary. A summary may be absent or empty, so code should treat it as optional. Do not rely on thought summaries being present, or confuse them with the final response; use the interaction’s answer output for the user-facing result. See Google’s Interactions API documentation.

Implementation checklist

  • Classify tasks in your own application; do not expect thinking_level to infer task type.
  • Pass the selected, model-supported value as generation_config.thinking_level to client.interactions.create().
  • Verify the chosen model’s current allowed levels and default before deployment.
  • Choose an output-token ceiling with the fact in mind that it includes thinking tokens, and handle incomplete results.
  • Make an explicit decision about state: use previous_interaction_id for stored continuation or store: false for stateless requests.
  • Evaluate latency, cost, and answer quality on your actual workload rather than assuming a level has universal performance characteristics.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.