October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
AI

LM Studio: Run Local LLMs on Your Computer

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes—LM Studio lets you download model weights, load them into your computer’s memory, and chat with a large language model locally. It runs on macOS, Windows, and Linux, can work fully offline after model files are available, and can expose local or network APIs for scripts and other applications. Your practical limit is hardware: system RAM, GPU/VRAM, model quantization, and context length determine which models run comfortably.

What LM Studio does

LM Studio is a desktop application for discovering, downloading, loading, and chatting with local large language models. Its main areas are:

  • Discover: search for model files and download them.
  • Model loader: place a downloaded model into memory with the runtime settings you choose.
  • Chat: converse with the loaded model and work with local documents.
  • Developer: run a localhost or local-network server using compatible APIs.
  • MCP connections: connect configured Model Context Protocol servers and tools.
  • Management: organize local models, prompts, and configurations.

LM Studio is not a hosted chatbot. Inference happens on your computer when the model is loaded, so the model’s memory requirements and your computer’s available resources matter more than an account tier.

System requirements by operating system

Check the requirements for your exact machine before downloading large model files. The figures below are recommendations or minimums as stated in LM Studio’s 2026 requirements documentation; they are not guarantees of a particular token-per-second speed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
GMKtec AI Mini PC Ryzen Al Max+ 395 (up to 5.1GHz) Mini Gaming Computers
  • EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
Platform Documented support and requirements Practical implication
macOS Apple Silicon M1, M2, M3, or M4; macOS 14.0 or newer; 16GB or more RAM recommended. Intel Macs are not currently listed as supported. Unified memory is shared by the operating system, model, and graphics workload. A 16GB machine is a sensible starting point, while larger models or long contexts need more headroom.
Windows x64 and ARM, including Snapdragon X Elite; AVX2 required on x64; at least 16GB RAM recommended; at least 4GB dedicated VRAM recommended. Verify that an x64 processor exposes AVX2 and that a discrete GPU has enough dedicated VRAM for the model configuration you intend to use.
Linux x64 and ARM64; distributed as an AppImage; Ubuntu 20.04 or newer listed as required. Make the AppImage executable and check your distribution’s graphics drivers before troubleshooting model acceleration.

RAM recommendations are not a model-size calculator. A quantized model’s file size, runtime overhead, KV cache, context window, and operating-system usage all consume memory. If loading fails or the system starts swapping, choose a smaller quantization, shorten the context, or close other applications.

Install LM Studio and download a model

  1. Install the current build. Download the version for macOS, Windows, or Linux. On Linux, use the supplied AppImage and grant execute permission if your desktop does not do so automatically.
  2. Open Discover. Search by model family or task. Read the model card and select a weight format offered for your platform.
  3. Choose a file format. LM Studio’s getting-started material describes GGUF and safetensors files. Select a quantization that fits available memory rather than automatically choosing the largest file.
  4. Download the weights. Model files can be several gigabytes or more, so use a stable connection and verify that you have sufficient disk space for the download and any additional files.
  5. Open the model loader. Select the downloaded model and configure the available GPU offload, context, and other runtime parameters.
  6. Load the model. Loading allocates memory for model weights and other parameters. Wait for the load to complete before starting a chat.
  7. Start Chat. Open the Chat tab, select the loaded model, and send a short test prompt before attempting a long document or tool workflow.

Choosing a model that will actually run

Match the model to memory

Start with the model file’s stated size, then leave room for runtime overhead and the context cache. A computer with 16GB of RAM should not be treated as having 16GB available to the model: the operating system and LM Studio itself need memory too. On Windows, dedicated VRAM can reduce pressure on system RAM, but the documented recommendation is at least 4GB, not a promise that every model will fit.

Quantization and quality

Quantized files reduce memory use at the cost of some numerical precision. Smaller quantizations generally make loading easier; larger ones may preserve more quality but can trigger swapping or an out-of-memory failure. Compare files from the same model family at the same context length when evaluating quality.

Context length is a separate cost

A model that loads successfully can still run out of memory when you increase its context window or attach lengthy documents. Begin with a conservative context, test a representative prompt, and increase it only when your machine remains responsive.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can LM Studio run offline?

Yes. LM Studio’s documentation states: “Offline Operation LM Studio can operate entirely offline, just make sure to get some model files first.” In practice, downloading model files normally requires internet access, unless you transfer or sideload them yourself. After the files are present, inference and document work can remain on-device.

Offline does not mean that every possible integration is offline. Starting a network server, calling a remote MCP tool, downloading a model, or using an external service creates a network data path. Review each tool and connection separately if data locality is important.

Rank #2
AMD Ryzen™ AI Halo - Personal AI Desktop Computer - Developer Platform - Linux OS
  • Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
  • 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
  • AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
  • Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
  • Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.

Use LM Studio as a local API server

Open the Developer tab to start a server on localhost or, where appropriate, your local network. LM Studio documents native REST, OpenAI-compatible, Anthropic-compatible, Python, and TypeScript interfaces. The v1 REST API, released with LM Studio 0.4.0, adds stateful chats, MCP through the API, authentication configuration, and model download, load, and unload endpoints.

Choose the right exposure

  • Localhost: best default for a script running on the same computer.
  • Local network: useful for another trusted device, but it expands who can reach the service. Configure authentication and firewall rules before enabling it.
  • OpenAI-compatible clients: point the client’s base URL at the address and port shown in Developer, then use the model identifier LM Studio reports.

Do not assume an OpenAI-compatible endpoint has identical model behavior or every hosted API feature. Confirm the request format, authentication setting, streaming behavior, and tool support in the interface you are using.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

MCP support and tool connections

LM Studio supports configured MCP connections, and its v1 API can expose MCP functionality programmatically. MCP lets a model interact with tools or resources supplied by an MCP server, but the tool server determines what data leaves the machine and what actions are possible.

Safer MCP practice

  • Connect only servers you understand and trust.
  • Give tools the minimum filesystem, network, or account access they require.
  • Keep sensitive work offline by disabling remote tools for that session.
  • Test a connection with a harmless read-only request before allowing writes or external actions.
  • When using a network-accessible LM Studio server, require authentication and restrict inbound access.

Common problems and fixes

The model will not load

Likely cause: insufficient RAM or VRAM, an overly large context, or a file/runtime mismatch. Fix: close memory-heavy applications, select a smaller quantization, reduce context length, or reduce GPU offload. Confirm that the downloaded file is complete.

The computer becomes slow or starts swapping

Likely cause: the model plus context exceeds available memory. Fix: stop the model, lower context, choose a smaller file, and leave operating-system headroom. Swapping can make a technically successful load unusably slow.

GPU acceleration is unavailable

Likely cause: unsupported hardware, missing drivers, or a platform-specific runtime setting. Fix: update the relevant graphics driver, verify that your hardware meets the platform requirements, and retry with a CPU-oriented configuration. A CPU fallback may work but can be slower.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
GMKtec EVO-X2 AI Mini PC AMD Ryzen Al Max+ 395 Up to 5.1GHz, 16C/32T
  • EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 64GB pool, which is perfect for running LLMs such as Deepseek 32B, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 4% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.

The Windows build reports an instruction error

Likely cause: an x64 processor without AVX2. Fix: verify AVX2 support; the documented Windows requirement is AVX2 on x64. ARM Windows follows the separate ARM support path.

The API client cannot connect

Likely cause: the Developer server is stopped, the client is using the wrong port or base path, a firewall blocks the connection, or authentication is enabled. Fix: start the server, copy its displayed address exactly, test localhost first, then inspect firewall and authentication settings before exposing it to the network.

Chat quality is poor

Likely cause: the selected model, quantization, prompt, or context does not suit the task. Fix: test another model from Discover, compare quantizations at the same context, and use a short, explicit prompt. Do not infer a universal model winner without controlling hardware, quantization, context, and workload.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and cost decisions

Performance

Speed depends on model architecture and size, quantization, context length, CPU, GPU, memory bandwidth, and how much work is offloaded. Benchmark only with the same model file, prompt, context, and hardware if you need a meaningful comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliability

Keep model files on fast local storage, maintain free disk space, and avoid running at the edge of available memory. For API use, add client timeouts appropriate to model generation and handle server restarts. A local service is available only while the computer, LM Studio, and its Developer server are running.

Cost

LM Studio’s local workflow avoids per-request hosted-model charges, but you supply the computer, storage, electricity, and downloads. The cheapest reliable setup is not always the smallest model: repeated swapping, failed loads, or unusable latency can cost more time than moving to a machine with additional RAM or GPU capacity.

Rank #4
MINISFORUM MS-S1 Max Mini Workstation AMD Ryzen AI Max+ 395(16C/32T) 64GB LPDDR5 2TB SSD Mini PC, HDMI+2X USB4+2X USB4 V2 Video Output, 2x10G RJ45 Port, WiFi7, BT5.4, Radeon 8060S Graphics Computer
  • 【Leading AI Mini Workstation】MINISFORUM AI MS-S1 Max Workstation comes with AMD Ryzen AI Max+ 395 processor, which uses AMD's latest generation Zen 5 architecture. It has 16 Cores and 32 Threads, the boost clock is up to 5.1GHz. The overall processor performance is up to 126 TOPS, and the NPU performance reaches up to 50 TOPS. AMD Ryzen AI enables improved productivity, advanced collaboration, and improved efficiency.
  • 【AMD Radeon 8060S Graphics 】The MS-S1 Max Mini PC equipped with AMD Radeon 8060S Graphics which built on the new generation of RDNA 3.5 architecture AMD graphics, it brings ultra-high frame rate experiences and advanced content creation features anywhere and delivers staggering performance. It can handle all your computing and multimedia tasks efficiently.
  • 【Five 8K Video Output】This MS-S1 Max Workstation comes with five video outputs, 1x HDMI (8K@60Hz), 2x USB4(40Gbps,Alt DP2.0,PD out 15W) and 2x USB4 V2(80Gbps,Alt DP2.0,PD out 15W) Outputs, which support multiple monitors display at the same time and provide a larger and wider filed of view and improve your work efficiency. It is used in fields that require high-performance computing and graphics processing, including digital signage and securities trading, as well as work that uses CAD, such as engineering design, scientific calculations, animation production, and post-production for movies and television
  • 【 Fast and Stable Wire & Wireless Speed】It comes with Two 10G Lan Ports for wired connection and and Wi-Fi 7 / BT5.4 for wireless connection, which increased the network speed greatly and expand its functions and improved performance of computer to a large extent and allows you to use more networks such as software routers (OpenWRT / DD-WRT / Tomato etc.), firewalls, NAT, network isolation etc.
  • 【Large Storage & Flexible Expandability】This Workstation equipped with 64GB LPDDR5-8000MHz + 2TB M.2 2280 PCIe4.0 SSD. There is another PCIe4.0 SSD slot available for up to 8TB, these SSD slots are compatible with RAID0 and RAID1, you can store movies, videos, photos, important files easily. What’s more, it also comes with 1x standard PCIex16 slot(PCIe4.0x4) inside.

Or skip the browser setup

If your documentation or workflow also needs a clean screenshot of a web page, ScreenshotNeo returns a PNG, JPEG, WebP, or PDF from one request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the complete parameter reference in the ScreenshotNeo documentation. A one-call cURL example is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Equivalent Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Equivalent Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

FAQ

Can I run LM Studio on an Intel Mac?

Intel Macs are not currently included in LM Studio’s documented macOS requirements; the listed Mac path is Apple Silicon M1 through M4 with macOS 14.0 or newer.

Do I need a dedicated GPU?

No universal GPU requirement is stated for every platform. CPU execution may be possible, while Windows documentation recommends at least 4GB of dedicated VRAM. Your model choice and desired speed determine whether a GPU is worthwhile.

Can another application use my loaded model?

Yes. Start the Developer server and use the documented REST, OpenAI-compatible, Anthropic-compatible, Python, or TypeScript interface, observing the server’s address, model identifier, and authentication settings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Are MCP tools automatically safe because the model is local?

No. The model can be local while an MCP tool accesses remote services or sensitive files. Review every server’s permissions and network behavior before connecting it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.