Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsYou can use a local coding model in VS Code by installing a model-provider extension and selecting its model in the Chat view. For Ollama, the current recommended route is Ollama’s official VS Code extension; VS Code’s older built-in Ollama provider is deprecated. Local chat can work without a Copilot plan and offline after setup, but it does not replace Copilot-backed inline suggestions, semantic search, or embedding features.
Connect Ollama to VS Code chat
This setup uses Ollama to run a model on your computer and its official extension to make that model available as a VS Code chat provider. The steps follow Microsoft’s VS Code language-model documentation and the Ollama extension recommendation in the VS Code 1.127 release notes.
- Install Ollama and download a model. Follow Ollama’s installation instructions, then download a compatible model using Ollama. The Foundry Toolkit guide documents the command pattern
ollama pull <model-name>; use the model name you intend to run. - Open the language-model management screen. In VS Code, open the Chat view’s model picker and choose Manage Language Models. You can also run Chat: Manage Language Models from the Command Palette.
- Install the provider. Choose Install Model Providers, or search the Extensions view for
@tag:language-models. Install the official Ollama extension published by Ollama and complete its setup flow. - Select the model and test it. Return to the Chat view’s model picker, select your local model, and try a small coding question or task. If the model is missing or cannot respond, confirm that Ollama has the model and follow the extension’s setup guidance.
Do not follow older instructions that rely on VS Code’s built-in Ollama provider as the default. Microsoft’s 1.127 release notes say that provider is deprecated and recommend the official Ollama extension instead.
Choose between the Ollama extension and Foundry Toolkit
These are different workflows, not two required steps in one setup. Use the Ollama extension when your main goal is to select an Ollama model as a provider in VS Code chat. Microsoft’s Foundry Toolkit documentation describes a broader model-discovery and experimentation workflow that can include Ollama and other supported local sources.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
| Option | Best fit | Setup detail | Relevant limitation |
|---|---|---|---|
| Official Ollama VS Code extension | Using an Ollama model through VS Code’s chat model picker | Install Ollama, download a model, then install the Ollama provider extension and select the model in chat. | Model and workflow capabilities depend on the model and provider. |
| Foundry Toolkit for VS Code | Exploring models, using a catalog or playground, or working in AI app-development flows | Download the model in Ollama first. In Foundry Toolkit, choose Add Ollama Model, acknowledge the third-party provider, and select an installed model. A custom Ollama endpoint is also supported in that workflow. | The documented Ollama integration does not support attachments. |
Foundry Toolkit supports Ollama as well as other local sources such as Foundry Local and ONNX, alongside hosted sources. Its ability to add an Ollama model should not be confused with the direct provider-extension setup for VS Code chat.
What a local model does—and does not—replace
Microsoft says VS Code BYOK (bring your own key/provider) models can be used for chat without a GitHub account or Copilot plan, and local-model chat can work offline once the model and provider are set up. You can also direct some utility tasks, including title or commit-message generation, to local models through the chat.utilityModel and chat.utilitySmallModel settings. See Microsoft’s language-model documentation for provider and setting details.
Rank #2
- Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
- 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
- AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
- Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
- Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.
BYOK is not a replacement for every Copilot capability. Microsoft identifies inline suggestions, semantic search, and embedding-dependent features as requiring GitHub Copilot services rather than BYOK models. Availability of features such as tool calling, vision, and thinking varies by model and may also vary by harness. Check the requirements of the specific workflow you want to use in Microsoft’s language-model guide and model capability overview.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshoot setup and choose a suitable model
- Ollama is missing from the provider list: verify that the official Ollama extension is installed and follow its current setup flow. The older built-in provider is deprecated.
- Foundry Toolkit shows no Ollama models: download a model in Ollama first; the Toolkit integration lists models already installed there.
- Chat works offline, but another feature does not: local chat does not bring GitHub-service features such as inline suggestions, semantic search, or embeddings offline.
- An agent workflow cannot use tools: confirm that both the chosen model and provider expose the tool-calling capabilities the workflow needs. Support is not universal.
Choose a model based on coding performance for your tasks, required tool or vision support, context needs, and the computer resources it requires. Those requirements vary by model and runtime; the cited Microsoft setup documentation does not establish universal memory, storage, or GPU minimums.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Best Value
Rank #4
- 🚨 Your Productivity AI Companion: Built for designers, editors, creators and studios, IT13 Max blends cloud AI inspiration with local NPU acceleration while keeping files private. For stable 24/7 workflows, it features quiet cooling, solid construction, original-grade SSD flash and rigorous testing. Backed by a 3-year warranty, it is a reliable Productivity AI Companion
- ➊ 3-Year Warranty + Precision Engineering for Long-Term Reliability & Business Use: From design to components, GEEKOM maintains highest quality standards. Each unit undergoes rigorous reliability testing for stable, long-term operation. Backed by a 3-year official warranty – peace of mind for home and business. Stable, durable, reliable. More than performance – a trusted partner (𝙂𝙚𝙩 𝘽𝙧𝙖𝙣𝙙-𝘿𝙞𝙧𝙚𝙘𝙩 𝙎𝙪𝙥𝙥𝙤𝙧𝙩: 𝙂𝙀𝙀𝙆𝙊𝙈 𝙊𝙛𝙛𝙞𝙘𝙞𝙖𝙡 𝙒𝙚𝙗𝙨𝙞𝙩𝙚)
- ➋ Intel Core Ultra 9 185H (TDP 65W) 2–3× AI Power for Developers & Engineers:2× faster graphics, 2–3× higher AI power, 20–30% faster video editing than i9. Run LLMs, computer vision, and ML workloads locally – no cloud latency, no privacy concerns. From AI inference to model training, this mini PC handles it all. For scientists, engineers, developers, and creatives – a ready-to-deploy productivity machine for intensive workloads
- ➌ Why pay more for less? 16GB DDR5 (higher bandwidth, better stability)+1TB SSD. Outperforms traditional desktops at a lower cost. Run office apps, edit 4K video in DaVinci Resolve (Linux or Windows), or handle heavy creative workloads – smooth and responsive. Desktop power, mini PC convenience. Smaller, more efficient, space-saving
- ➍ Silent Operation with IceBlast 3.0 for Hospitals, Schools & Shared Environments: Tired of loud fans disrupting patient care or classrooms? IT13 MAX with IceBlast 3.0 delivers 65W sustained performance while whisper-quiet – 40% quieter than typical mini PCs. Deploy in hospital nurse stations, school computer labs, or work late without waking family. High-performance computing – without the noise
Rank #3
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




