To stream Claude output to a browser, configure an API Gateway REST API Lambda proxy integration for response transfer mode STREAM, then have Lambda forward Claude’s server-sent events as they arrive. This guide uses Anthropic’s direct Messages API as the Claude route; it does not use Amazon Bedrock’s AWS event-stream format. The browser receives SSE events, not a guarantee that each network chunk equals one token.
How the real-time Claude chat API fits together
The request travels from the browser to API Gateway, then to a Lambda proxy integration. Lambda makes a streaming request to Anthropic’s Messages API and relays the response body to API Gateway, which streams it to the browser. The client can render text as it arrives rather than waiting for Claude’s complete answer.
As an Amazon Associate I earn from qualifying purchases.
- Browser: sends a chat request and reads the response incrementally.
- API Gateway: accepts the request and invokes Lambda using response streaming.
- Lambda: keeps Anthropic credentials on the server, starts the upstream request, and forwards the SSE response.
- Anthropic Messages API: emits events describing the response lifecycle and text content.
The implementation below relays Anthropic’s SSE bytes rather than parsing and rebuilding its events. That keeps the example small and avoids confusing Anthropic SSE with Bedrock’s legacy AWS event-stream framing. If you choose Bedrock instead, use its matching API, credentials, SDK, and framing; Anthropic documents a newer Bedrock Messages endpoint that uses SSE as well as legacy InvokeModel/Converse integrations that use AWS event-stream encoding.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What to configure in API Gateway and Lambda
Use a REST API with streaming enabled
API Gateway response streaming is supported for REST APIs with proxy integrations, including Lambda proxy integrations. Set the integration’s response transfer mode to STREAM; the default buffered mode does not provide incremental delivery. API Gateway invokes Lambda with InvokeWithResponseStream. AWS documents generative-AI chat and incremental progress as use cases for this feature.
#1 Best Overall
Use an AWS Region that supports Lambda response streaming, and set the function timeout high enough for the longest response you intend to allow. The route should return the SSE content type, and any proxy, CDN, or client in the path must also permit streaming rather than buffer the response.
Return the Lambda streaming proxy format
A streaming Lambda proxy response needs a metadata section and a payload section. In Node.js, use awslambda.streamifyResponse() and awslambda.HttpResponseStream.from() so the runtime formats the metadata correctly. If writing the format manually, valid JSON metadata must be followed by eight null bytes within the first 16 KB. The supported metadata fields include headers, multiValueHeaders, cookies, and statusCode; do not add arbitrary fields.
Lambda example: relay Claude SSE events
This Node.js example expects a JSON request body with a messages array. Configure CLAUDE_MESSAGES_URL with the current Anthropic Messages API endpoint, and set ANTHROPIC_API_KEY, ANTHROPIC_VERSION, and CLAUDE_MODEL as Lambda environment variables. Use Anthropic’s current documentation for the endpoint, required API version header, and model identifier; model IDs and lifecycle information can change.
exports.handler = awslambda.streamifyResponse(async (event, rawResponseStream) => {
let input;
try {
input = JSON.parse(event.body || "{}");
} catch {
const response = awslambda.HttpResponseStream.from(rawResponseStream, {
statusCode: 400,
headers: { "content-type": "application/json; charset=utf-8" }
});
response.end(JSON.stringify({ error: "Request body must be valid JSON." }));
return;
}
if (!Array.isArray(input.messages) || input.messages.length === 0) {
const response = awslambda.HttpResponseStream.from(rawResponseStream, {
statusCode: 400,
headers: { "content-type": "application/json; charset=utf-8" }
});
response.end(JSON.stringify({ error: "messages must be a non-empty array." }));
return;
}
let upstream;
try {
upstream = await fetch(process.env.CLAUDE_MESSAGES_URL, {
method: "POST",
headers: {
"content-type": "application/json",
"x-api-key": process.env.ANTHROPIC_API_KEY,
"anthropic-version": process.env.ANTHROPIC_VERSION
},
body: JSON.stringify({
model: process.env.CLAUDE_MODEL,
max_tokens: 1024,
stream: true,
messages: input.messages
})
});
} catch {
const response = awslambda.HttpResponseStream.from(rawResponseStream, {
statusCode: 502,
headers: { "content-type": "application/json; charset=utf-8" }
});
response.end(JSON.stringify({ error: "Could not connect to the Claude API." }));
return;
}
if (!upstream.ok || !upstream.body) {
const body = upstream.body ? await upstream.text() : "Claude API returned no response body.";
const response = awslambda.HttpResponseStream.from(rawResponseStream, {
statusCode: upstream.status || 502,
headers: { "content-type": "application/json; charset=utf-8" }
});
response.end(body);
return;
}
const response = awslambda.HttpResponseStream.from(rawResponseStream, {
statusCode: 200,
headers: {
"content-type": "text/event-stream; charset=utf-8",
"cache-control": "no-cache",
"x-content-type-options": "nosniff"
}
});
try {
for await (const chunk of upstream.body) {
if (!response.write(chunk)) {
await new Promise(resolve => response.once("drain", resolve));
}
}
response.end();
} catch {
// HTTP headers are already committed; a new HTTP status cannot report this failure.
response.write(`event: errorndata: ${JSON.stringify({ error: "The stream ended unexpectedly." })}nn`);
response.end();
}
});
The example checks for an upstream connection failure and non-success response before beginning the downstream stream, so those cases can use an HTTP error status. Once streaming headers and body data have started, the status cannot be changed. The catch block therefore sends an application-level terminal error event if the relay fails mid-stream. A client should handle that event as well as normal completion events.
Rank #3
Do not trust client-supplied model names or credentials. Keep API keys in server-side configuration, validate the message schema and size, and apply authentication and rate limits appropriate to your application. The example’s fixed max_tokens is a request setting, not a service limit; choose a value based on your product’s requirements and the current model documentation.
How should the browser read the SSE stream?
For a simple same-origin POST endpoint, use fetch() and read its response body as a stream. The browser’s native EventSource API is designed for opening an SSE connection, commonly with GET; it does not provide the same arbitrary POST-body interface. A byte chunk from the network may split an SSE event or contain multiple events, so buffer decoded text and parse complete events at the blank-line boundary rather than treating each read as a token.
const response = await fetch("/chat", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ messages })
});
if (!response.ok || !response.body) {
throw new Error(`Chat request failed: ${response.status}`);
}
const reader = response.body.getReader();
const decoder = new TextDecoder();
let pending = "";
while (true) {
const { value, done } = await reader.read();
pending += decoder.decode(value || new Uint8Array(), { stream: !done });
const events = pending.split("nn");
pending = events.pop() || "";
for (const eventText of events) {
if (!eventText.trim()) continue;
const eventName = eventText.match(/^event:s*(.*)$/m)?.[1] || "message";
const data = eventText.match(/^data:s*(.*)$/m)?.[1] || "";
if (eventName === "error") {
throw new Error(data || "The response stream failed.");
}
// Parse and render the data according to Anthropic's current SSE event schema.
}
if (done) break;
}
The final parsing and rendering step must follow the current Anthropic Messages API event schema. The relay intentionally does not assume that each SSE event is a text delta or that a particular event name means the answer is complete. Render only text-bearing events, handle the API’s completion event, and surface error events distinctly from ordinary content.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Why is API Gateway buffering my response?
First confirm that the deployed REST API integration is configured for STREAM, not the default buffered response mode. Then test the deployed endpoint with a client that displays data as received:
curl -i --no-buffer
-H 'content-type: application/json'
-d '{"messages":[{"role":"user","content":"Say hello in one sentence."}]}'
https://YOUR_DEPLOYED_API/chat
Replace the example host and route with your deployed endpoint. AWS notes that API Gateway’s test invocation buffers a streaming response and returns it after completion, after 35 seconds, or after more than 1 MB has accumulated. It cannot prove that a real client receives incremental chunks. Also check that any intermediary between the client and API Gateway is not buffering.
For diagnosis, AWS documents streaming-specific access-log values including response transfer mode, time to all headers, time to first content, and integration latency. Compare time to first content with total integration time to distinguish a slow upstream first response from a path that delivers the response only after it finishes.
Limits and trade-offs to plan for
| Service or behavior | Documented limit or constraint | What it means for this API |
|---|---|---|
| API Gateway stream duration | Up to 15 minutes | Keep expected Claude response duration within the gateway’s stream window. |
| API Gateway idle timeout | 5 minutes for Regional and private endpoints; 30 seconds for edge-optimized endpoints | A stream with no data for longer than the applicable idle interval can time out. |
| API Gateway payload bandwidth | 2 MB/s for payload beyond the first 10 MB | Large responses are subject to a throughput cap after that initial amount. |
| Lambda streamed response size | Up to 200 MB; first 6 MB uncapped, later data limited to 2 MB/s | Lambda’s size and bandwidth rules are separate from API Gateway’s rules; the stricter applicable constraint governs the full path. |
| Buffered Lambda response size | 6 MB maximum | Streaming supports larger responses than buffered Lambda responses, but it does not remove API Gateway’s own constraints. |
| API Gateway features requiring buffering | Endpoint caching, VTL response transformation, and API Gateway content encoding are unavailable for response streaming | Perform any required transformation in the application or select an architecture that does not depend on those features. |
The service figures above are documented AWS product limits, not benchmark results. Lambda streaming is not available in every AWS Region. Check current regional support and deployed service limits before choosing a Region. If the client disconnects, Lambda execution may continue and consume duration; set a sensible function timeout and account for the possibility that work continues after the caller is gone.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Choose and maintain the Claude route deliberately
This example uses Anthropic’s direct Messages API: the Lambda function authenticates with an Anthropic API key and relays SSE. Amazon Bedrock is a different configuration, not a drop-in change to the credential header or event parser. Anthropic distinguishes legacy Bedrock InvokeModel/Converse integrations, which use AWS event-stream encoding, from its newer Bedrock Messages endpoint using SSE. Follow the documentation for the specific route you deploy.
Before deploying or updating the function, verify the endpoint, authentication method, required headers, model identifier, and model availability for your account and region. Anthropic’s model deprecation documentation is the relevant place to check lifecycle changes; do not assume a model identifier remains valid indefinitely.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




