What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The reliable pattern is a browser chat UI connected to your own server endpoint. The browser sends messages to your backend; the backend authenticates the user, validates and limits the request, keeps provider credentials private, calls an LLM, and streams the answer back. Build that boundary first, then add rendering, tools, retrieval, safety controls, observability, and a deliberate retention policy.
What you are building
An LLM interface is more than a text box. It is an application with a trust boundary:
- Browser: displays messages, accepts input, shows streaming state, and handles retry and cancellation.
- Application server: authenticates requests, checks origin and quotas, validates message shape and size, applies system instructions, calls the model, and records only approved telemetry.
- Model service: generates text and, when enabled, may call tools or use retrieved content.
Do not put a provider key in JavaScript shipped to the browser. Anyone can inspect it, reuse it, and spend your quota. The browser should know only your application endpoint and a user session token or cookie.
Decide the assistant’s scope before writing code
Define the job and boundaries
Write down the tasks the assistant may perform, the topics it must refuse or hand off, and what counts as a consequential action. A support assistant might answer from approved documentation but never change an account. A booking assistant may propose a reservation but require explicit confirmation before submitting it.
#1 Best Overall
Choose the conversation contract
Decide whether conversations are anonymous or account-bound, whether users can start multiple threads, the maximum message and context sizes, and how cancellation works. Define a stable request such as {"messages":[{"role":"user","content":"..."}]}, then reject unknown roles, oversized content, and malformed JSON on the server.
Choose an API or SDK surface
Common options include a provider SDK, an OpenAI-compatible API, a framework AI SDK, or a gateway that normalizes several providers. Vercel documents AI SDK, OpenAI-compatible Chat Completions and Responses, Anthropic Messages, and OpenResponses surfaces. Their support for streaming, tool calls, and structured outputs differs by surface and model, so verify the exact combination you deploy.
| Decision | Prefer a normalized SDK when | Prefer a provider-native API when |
|---|---|---|
| Provider portability | You expect fallback routing or multiple vendors. | You need a provider-specific capability immediately. |
| Feature access | Your required features are supported consistently. | You need the newest native tools or response controls. |
| Operations | You want one instrumentation and error shape. | Your team already operates the provider’s SDK confidently. |
There is no source-backed universal winner for latency, price, or answer quality. Measure your own prompts, traffic patterns, and failure budget.
Implement a minimal server endpoint
The following Express-style example shows the responsibilities, not a requirement to use Express or a particular vendor SDK. Keep the model call behind an environment variable and return a stream.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
import express from "express";
import crypto from "node:crypto";
const app = express();
app.use(express.json({ limit: "64kb" }));
app.post("/api/chat", async (req, res) => {
const user = await authenticate(req); // Your session or token check
if (!user) return res.status(401).json({ error: "unauthorized" });
const messages = req.body?.messages;
if (!Array.isArray(messages) || messages.length > 40) {
return res.status(400).json({ error: "invalid messages" });
}
for (const m of messages) {
if (!["user", "assistant"].includes(m.role) ||
typeof m.content !== "string" || m.content.length > 12000) {
return res.status(400).json({ error: "invalid message" });
}
}
if (!(await withinQuota(user.id))) {
return res.status(429).json({ error: "rate limit exceeded" });
}
const controller = new AbortController();
req.on("close", () => controller.abort());
res.setHeader("Content-Type", "text/event-stream; charset=utf-8");
res.setHeader("Cache-Control", "no-cache, no-transform");
res.setHeader("Connection", "keep-alive");
try {
const upstream = await fetch(process.env.LLM_ENDPOINT, {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.LLM_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: process.env.LLM_MODEL,
stream: true,
messages: [
{ role: "system", content: "Follow the product policy. Never claim an action was completed unless the server confirmed it." },
...messages
]
}),
signal: controller.signal
});
if (!upstream.ok || !upstream.body) {
return res.status(502).json({ error: "model unavailable" });
}
for await (const chunk of upstream.body) res.write(chunk);
res.end();
} catch (error) {
if (!controller.signal.aborted) {
res.write(`event: error\ndata: {"error":"generation failed"}\n\n`);
res.end();
}
}
});
app.listen(process.env.PORT || 3000);
Use your provider’s documented stream format. Some return server-sent events, while others expose an SDK async iterator. Never forward raw provider errors, keys, internal prompts, or stack traces to the browser.
Connect the browser and stream progressively
The client should append each received text delta to the current assistant message, keep the submit control disabled while a request is active, and expose cancel, retry, and error states. With a framework hook such as Vercel’s useChat, the hook can manage message state and streaming; with plain browser JavaScript, use fetch and read the response body.
async function sendMessage(text) {
addMessage({ role: "user", content: text });
const response = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages: [...messages, { role: "user", content: text }] })
});
if (!response.ok || !response.body) throw new Error("Request failed");
const reader = response.body.getReader();
const decoder = new TextDecoder();
let answer = "";
addMessage({ role: "assistant", content: "" });
while (true) {
const { value, done } = await reader.read();
if (done) break;
answer += decoder.decode(value, { stream: true });
updateLastAssistantMessage(answer);
}
}
Parse the actual event framing returned by your API rather than assuming every chunk is plain text. Flush proxies and disable response buffering where your hosting platform requires it.
Render model output as untrusted data
Markdown is convenient but is not automatically safe. A rendered image URL, link, or HTML fragment can trigger a browser request that discloses information or reaches an unexpected host. Sanitize HTML with a maintained allowlist, strip dangerous URL schemes, restrict remote images where practical, and consider rendering a constrained subset instead of arbitrary HTML. Test the production renderer, not only the model response.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Rank #3
Handle prompt injection and tools
Prompt injection is untrusted text attempting to override your instructions. It can come from a user, a retrieved document, a web page, or tool output. Treat every one of those inputs as data, not policy.
Layered controls
- Keep policy instructions separate from user and retrieved content, and state prohibited actions explicitly.
- Use structured outputs for classifications and routing decisions; validate the schema before acting.
- Give tools the least privilege possible. Separate read-only tools from write tools.
- Require a user confirmation for purchases, account changes, messages, deletions, or other consequential operations.
- Screen tool output before returning it to the model, especially when it contains web or document text.
- Limit sensitive fields in prompts and component properties, redact personal data from logs, and monitor anomalous requests.
- Evaluate adversarial prompts continuously; guardrails reduce risk but do not make an agent infallible.
Authentication, quotas, and abuse controls
Authenticate at your endpoint, enforce authorization for each conversation, and apply per-user and per-IP limits. Validate content type, length, number of turns, and uploaded files server-side. Add request timeouts, cancellation, concurrency limits, and a budget ceiling. Rate limiting is an application responsibility even when a provider offers its own limits.
Privacy and retention decisions
Choose what you retain, why, and for how long before launch. Publish the policy, provide deletion where applicable, and avoid putting secrets or unnecessary personal information into prompts or logs. Separate content logs from operational metrics and redact identifiers.
Provider terms are feature- and account-specific. Anthropic’s current API documentation says standard retained data is not used for model training without express permission; conversation content is not retained by default except for specified covered-model cases requiring 30-day retention; and zero-data-retention is an organization-level arrangement that must be enabled separately. Verify the current policy, API feature, and contract instead of generalizing these statements to another provider.
Rank #4
Production checklist
- Secrets exist only in server-side configuration and are rotated.
- Authentication, authorization, validation, quotas, timeouts, and cancellation are tested.
- Streaming works through your production proxy and reconnect behavior is defined.
- Markdown/HTML, links, images, code blocks, and error messages are safely rendered.
- Tool scopes, confirmation prompts, injection tests, and audit events are reviewed.
- Retention, deletion, redaction, provider terms, and incident contacts are documented.
- Metrics distinguish user cancellations, provider failures, rate limits, timeouts, and safety blocks.
Troubleshooting common failures
The browser says 401 or 403
Check session cookies, CSRF/origin rules, and server-side authorization for the requested conversation. Do not solve this by exposing the provider key.
The response arrives all at once
Inspect proxy buffering, response headers, and the provider’s stream mode. Confirm your client parses the provider’s event format and that your hosting runtime supports streaming.
Requests time out
Set an explicit server timeout, abort upstream work when the client disconnects, reduce context size, and return a retryable error. Do not automatically duplicate a non-idempotent tool action.
Answers reveal hidden instructions or data
Reduce prompt context, remove secrets, tighten retrieval filters, validate tool output, and add tests for extraction and indirect-injection attempts. Review logs for the first untrusted source that entered the context.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Usage costs spike
Cap input and output tokens, limit history, cache eligible requests, enforce quotas, and alert on spend per user and route. Recheck model and API-surface pricing before changing providers.
Or skip the browser setup
If your interface needs screenshots of pages, you can call ScreenshotNeo directly rather than operating a browser service. Its API accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, with the result identified by response headers. It also provides an MCP server for AI agents, including Claude and Cursor, with take_screenshot, get_page_info, and capture_pdf tools.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the complete options and response behavior in the ScreenshotNeo API documentation. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots. Sign up for ScreenshotNeo free.
Frequently Asked Questions
How do I add an AI chatbot to my website?
Create a browser chat component that posts messages to an authenticated server endpoint, call the model from that server, and stream validated output back to the component.
Recommended Free Tools
How do I stream LLM responses to a web UI?
Enable streaming in the model request, return the provider’s stream through your endpoint with buffering disabled, parse its event format in the browser, and append deltas to the active assistant message.
Should chat history be stored in the browser or database?
Use browser storage only for explicitly local, low-sensitivity drafts. Store account conversations server-side only when a product need justifies it, with a documented retention and deletion policy.
The Bottom Line
Keep the model behind your server, stream through a controlled endpoint, render output safely, treat every external text source as untrusted, and make retention and tool permissions explicit before launch.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

