AI Models & LLMs

Best Tools to Run LLMs Locally

Also known as: tools to run llms locally · local AI apps

You can run a large language model (LLM) on your own laptop or desktop, with no internet connection and no monthly bill, and your prompts and files never leave the machine. You need two things: an open model small enough for your memory, and a tool that downloads it, loads it onto your graphics card or Apple chip, and gives you a chat window or an API.

This page ranks those tools. It does not rank the models themselves: for that, see our best local LLMs ranking. We compared 10 apps, runtimes and front ends on ease of use, speed and hardware support, model support, features, licence and privacy, and how actively each one is maintained. Versions, GitHub stars and prices were checked on 25 September 2026. If you are buying hardware, see best GPUs for AI and best AI laptops.

Quick answer

Ollama is the best tool to run LLMs locally for most people in September 2026. It is free and MIT-licensed, runs on Mac, Windows and Linux, starts a model with one command, and almost every other AI app can connect to it. The best Ollama alternatives are LM Studio for a polished point-and-click app, Jan for a fully open-source desktop app, AnythingLLM for chatting with your documents, llama.cpp for maximum control and vLLM for serving a model to a whole team.

Top picks at a glance

Scoreboard

Scores are out of 10. The overall score is the weighted average of the criteria below.

#ToolOverallEase of useSpeed & hardware supportModel supportFeatures & integrationsLicence & privacyMaintenance & communityPrice fromBest for
1Ollama
Ollama Inc.
9.19.38.89.09.28.89.6Free (local); cloud plans from $20/month
Free tier
Most people, and any app that needs a local model behind it
2LM Studio
Element Labs
8.89.69.09.08.87.09.0Free (local); Bionic+ $20/month for cloud models
Free tier
Beginners who want a polished point-and-click app
3llama.cpp
ggml-org (part of Hugging Face since February 2026)
8.66.89.49.58.09.89.6Free (open source)
Free tier
Maximum control, the newest models first, and embedding in products
4Jan
Menlo Research
8.58.68.38.38.59.57.8Free (open source)
Free tier
An open-source desktop app that mixes local and cloud models
5AnythingLLM
Mintplex Labs
8.58.87.27.89.29.38.8Free (desktop and self-hosted)
Free tier
Chatting privately with your own documents
6vLLM
vLLM project (open source)
8.35.59.69.08.59.89.5Free (open source)
Free tier
Serving a model to a team, an app or an internal API
7LocalAI
LocalAI (open-source project led by Ettore Di Giacinto)
8.36.88.08.89.09.88.5Free (open source)
Free tier
One private server that replaces several cloud AI APIs
8Open WebUI
Open WebUI Inc.
8.27.57.58.59.67.89.3Free (self-hosted)
Free tier
A private, ChatGPT-style chat site for a family or team
9MLX-LM
Apple (ml-explore)
7.96.28.58.57.59.88.0Free (open source)
Free tier
Mac users who want new models first or to fine-tune locally
10GPT4All
Nomic AI
7.18.56.55.57.59.53.0Free (open source)
Free tier
Existing users on older hardware

Expert reviews

#1 · Most people, and any app that needs a local model behind it

Ollama

by Ollama Inc. · Freemium · Free (local); cloud plans from $20/month
9.1/10

Ollama is the default way to run open models on your own machine, and the tool most other apps plug into. Install it on macOS, Windows or Linux, type ollama run gemma4, and it downloads the model and starts a chat. The desktop app adds a chat window with file drag-and-drop, and a local API on port 11434 accepts OpenAI- and Anthropic-style requests, so coding tools, note apps and front ends such as Open WebUI connect in seconds. The ollama launch command sets up Claude Code, Codex or OpenCode against a local model.

On Macs, Ollama moved to Apple's MLX framework in 2026. In Ollama's own test on an M5 Max, output speed rose from 58 to 112 tokens per second. NVIDIA cards (compute capability 5.0 and up) run through CUDA, and AMD cards through ROCm or Vulkan.

Is Ollama free? Yes for local use: running models on your own hardware is unlimited and the code is MIT-licensed. The paid plans (Pro $20/month, Max $100/month) only cover its optional cloud models.

Pick it if you want the simplest, best-supported local runtime. Skip it if you prefer browsing models in a full graphical app (LM Studio) or need to serve many users at once (vLLM).

Score breakdown

Ease of use9.3
Speed & hardware support8.8
Model support9.0
Features & integrations9.2
Licence & privacy8.8
Maintenance & community9.6

Key facts

Pricing
Free (local); cloud plans from $20/month (Running models on your own hardware is free and unlimited. Optional cloud models: Free plan with starter credits; Pro $20/month ($60 of usage); Max $100/month ($300 of usage); Team $500/month; Enterprise custom. Cloud usage is billed per token, e.g. gpt-oss:120b at $0.15 / $0.60 per 1M tokens.)
Free option
Yes
Platforms
macOS, Windows, Linux, Docker, API
Licence
MIT
GitHub stars
About 181.7k (ollama/ollama, 25 Sep 2026)
Latest release
v0.34.4 (23 Sep 2026)
Engines
MLX on Apple silicon (since 2026); GGUF models through llama.cpp

What we like

  • One command to download and run a model
  • Local OpenAI- and Anthropic-style API that most apps support
  • Fast MLX engine on Apple silicon
  • MIT licence and a huge community (181k GitHub stars)

Watch out for

  • Fewer tuning options than llama.cpp or LM Studio
  • Not built to serve many users at once
  • Cloud plans blur the line between local and hosted
#2 · Beginners who want a polished point-and-click app

LM Studio

by Element Labs · Freemium · Free (local); Bionic+ $20/month for cloud models
8.8/10

LM Studio is the easiest local AI app for people who do not want a terminal. It looks like a normal desktop program: search for a model, pick a size, click download and start chatting. It runs GGUF models through llama.cpp and, on Macs, Apple's MLX format. Developers get a local server with OpenAI- and Anthropic-compatible endpoints, a headless daemon (llmster) for servers, the lms command-line tool, SDKs for JavaScript and Python, and MCP tool support.

In July 2026 the company added LM Studio Bionic, a separate agent app for coding and file work that can use local models or optional paid cloud models. The original LM Studio app is still available and free.

The licence is the main thing to check. LM Studio is closed source. It has been free for work since July 2025, and the current terms allow personal and internal business use, but you may not modify it, reverse engineer it or resell it as a hosted service. Its privacy policy says local chats never leave your device and the app has no telemetry.

Pick it if you want a polished app with built-in model search. Skip it if you need open-source software you can audit or ship inside your own product.

Score breakdown

Ease of use9.6
Speed & hardware support9.0
Model support9.0
Features & integrations8.8
Licence & privacy7.0
Maintenance & community9.0

Key facts

Pricing
Free (local); Bionic+ $20/month for cloud models (The app and local models are free, including at work since July 2025. Bionic+ ($20/month) and Pro ($100/month) add US-hosted cloud models with zero data retention. Enterprise plans with SSO and model controls are by quote.)
Free option
Yes
Platforms
macOS, Windows, Linux, API
Licence
Closed source; free for personal and internal business use
Latest versions
LM Studio 0.4.25; Bionic agent app 1.1.6 (23 Sep 2026)
Engines
llama.cpp (GGUF) and Apple MLX
Requirements
Apple silicon Mac on macOS 14+, or x64/ARM Windows and Linux; 16GB RAM recommended

What we like

  • Friendliest app: search, download and chat without a terminal
  • Runs both GGUF (llama.cpp) and MLX models
  • Free for work use since July 2025
  • No telemetry; local chats stay on your device

Watch out for

  • Closed source; no modifying or reselling as a service
  • Intel Macs are not supported
  • Bionic and cloud plans make the product line confusing
#3 · Maximum control, the newest models first, and embedding in products

llama.cpp

by ggml-org (part of Hugging Face since February 2026) · Open source · Free (open source)
8.6/10

llama.cpp is the engine under much of local AI. LM Studio, Jan, LocalAI and GPT4All run models with it, Ollama uses it for GGUF models, and GGUF, its file format, is the standard way local models are shared. That is why new open models usually work here first. Using it directly gives you the most control over quantization, context length, GPU offload and speculative decoding.

It has also become easier to start with. The project's new home, llama.app, offers a one-line installer and a llama serve command that starts an OpenAI-compatible server with a built-in web chat. Version 0.5.0, released on 23 September 2026, added faster CUDA and Metal paths and new server options, and fresh builds ship several times a day.

It runs on almost anything, from Apple silicon and NVIDIA, AMD and Intel graphics cards to plain CPUs. The MIT licence lets you ship it inside your own product. Its team (ggml.ai) joined Hugging Face in February 2026 and says the project stays fully open source; NVIDIA has since agreed to buy Hugging Face.

Pick it if you want full control, the newest models first, or an engine to build on. Skip it if you want a point-and-click app.

Score breakdown

Ease of use6.8
Speed & hardware support9.4
Model support9.5
Features & integrations8.0
Licence & privacy9.8
Maintenance & community9.6

Key facts

Pricing
Free (open source) (MIT licence: free to use, modify and ship inside your own products.)
Free option
Yes
Platforms
macOS, Windows, Linux, Docker, API
Licence
MIT
GitHub stars
About 129.5k (ggml-org/llama.cpp, 25 Sep 2026)
Latest release
v0.5.0 (23 Sep 2026), plus several builds a day
Model format
GGUF, the standard file format for local models

What we like

  • New models usually work here first
  • Runs on almost any GPU or CPU
  • MIT licence: free to embed in products
  • Fine control over quantization, offload and context

Watch out for

  • Command line first; settings take time to learn
  • No built-in document chat (RAG)
  • Ownership now tied to Hugging Face and its pending NVIDIA deal
#4 · An open-source desktop app that mixes local and cloud models

Jan

by Menlo Research · Open source · Free (open source)
8.5/10

Jan is the best fully open-source alternative to LM Studio. It is a desktop app for macOS, Windows and Linux with a clean chat window, a model hub, projects, file upload and web search, and it runs models locally through llama.cpp or, on Macs, MLX. You can also add cloud models from OpenAI, Anthropic, Gemini, Groq or OpenRouter with your own API keys and switch between local and cloud in the same app.

For developers there is a local API server, a command-line tool and MCP support for connecting tools, plus integrations with Claude Code. Jan's team also ships its own small models, such as Jan-v3-4B and Jan-Code-4B, built for agent tasks.

The code is Apache 2.0, and Jan says it collects no data until you pick tracking settings at first launch. The team says the app has passed 6.7 million downloads.

The main weakness is pace. The latest stable release, v0.8.4, came out on 23 July 2026, two months before we checked, while Ollama and llama.cpp ship every week or every day.

Pick it if you want an open-source, private app that mixes local and cloud models. Skip it if you need the newest model architectures on day one.

Score breakdown

Ease of use8.6
Speed & hardware support8.3
Model support8.3
Features & integrations8.5
Licence & privacy9.5
Maintenance & community7.8

Key facts

Pricing
Free (open source) (Apache 2.0. The desktop app is free; cloud models from OpenAI, Anthropic, Gemini, Groq, OpenRouter and others use your own API keys.)
Free option
Yes
Platforms
macOS, Windows, Linux, API
Licence
Apache 2.0
GitHub stars
About 44.6k (janhq/jan, 25 Sep 2026)
Latest release
v0.8.4 (23 Jul 2026)
Engines
llama.cpp and MLX

What we like

  • Apache 2.0 open source
  • Local (llama.cpp, MLX) and cloud models in one app
  • MCP support and a local API server
  • Usage analytics only if you opt in

Watch out for

  • Slower release pace than Ollama or llama.cpp
  • Fewer third-party integrations than Ollama
  • Some features, such as Cowork, are still in preview
#5 · Chatting privately with your own documents

AnythingLLM

by Mintplex Labs · Freemium · Free (desktop and self-hosted)
8.5/10

AnythingLLM is the best choice if your main goal is chatting with your own documents. Drop PDFs, Word files or web pages into a workspace and it builds a private search index (RAG) on your machine, so answers can draw on your files. Around that sit AI agents with tools such as web search, file editing and scheduled jobs, a meeting assistant that transcribes calls on your computer, and an Android app.

The desktop app is one download for macOS, Windows and Linux. It ships with a built-in model engine, based on Ollama's open-source engine, and can recommend a model for your hardware, so you do not need anything else installed. If you already run Ollama, LM Studio or LocalAI, or want a cloud model, you can connect those instead.

It is MIT-licensed and free on desktop or self-hosted with Docker, which also adds multi-user workspaces. Mintplex Labs sells a hosted cloud version from $50 a month.

Raw speed depends on the engine underneath, and the docs say the built-in engine is not a full Ollama replacement, so power users often pair it with Ollama or LM Studio.

Pick it if you want private document chat and agents with little setup. Skip it if you only need a fast model runner.

Score breakdown

Ease of use8.8
Speed & hardware support7.2
Model support7.8
Features & integrations9.2
Licence & privacy9.3
Maintenance & community8.8

Key facts

Pricing
Free (desktop and self-hosted) (MIT licence. The desktop app and Docker self-hosting are free. Hosted cloud: Basic $50/month, Pro $99/month, Enterprise by quote.)
Free option
Yes
Platforms
macOS, Windows, Linux, Android, Docker
Licence
MIT
GitHub stars
About 66.5k (Mintplex-Labs/anything-llm, 25 Sep 2026)
Latest release
v1.16.2 (22 Sep 2026)
Built-in engine
Based on Ollama's open-source engine; can also connect to Ollama, LM Studio, LocalAI or cloud APIs

What we like

  • Strong built-in document chat (RAG)
  • One-file desktop app with a built-in model engine
  • Agents, scheduled jobs and a local meeting assistant
  • MIT licence; free self-hosting with multi-user workspaces

Watch out for

  • Speed depends on the engine underneath
  • Built-in engine lacks some Ollama features
  • Busy interface if you only want a simple chat
#6 · Serving a model to a team, an app or an internal API

vLLM

by vLLM project (open source) · Open source · Free (open source)
8.3/10

vLLM is what you use when a local model has to serve many people at once. Instead of a desktop app, it is a Python server that loads a model onto your GPUs and exposes an OpenAI-compatible API. Its PagedAttention memory system raised throughput 2 to 4 times over earlier serving systems at the same latency in the original research paper, and it batches many requests together so one GPU can serve a whole team.

It loads a wide range of Hugging Face models and supports tool calling, structured outputs, speculative decoding and multi-GPU setups. If you would rather not run the servers yourself, the hosts on our best AI inference providers list sell the same kind of service by the token.

It is not a beginner tool. vLLM officially runs on Linux (Windows users need WSL), wants an NVIDIA GPU with compute capability 7.5 or newer (or AMD ROCm, Intel, or Apple silicon through the vLLM-Metal plugin), and you configure it from the command line. For one person on a laptop, Ollama or LM Studio is much quicker to set up.

Pick it if you are hosting a model for a team, an app or an internal API. Skip it if you just want to chat with a model on your own computer.

Score breakdown

Ease of use5.5
Speed & hardware support9.6
Model support9.0
Features & integrations8.5
Licence & privacy9.8
Maintenance & community9.5

Key facts

Pricing
Free (open source) (Apache 2.0. You pay only for your own hardware or rented GPUs.)
Free option
Yes
Platforms
Linux, Docker, API
Licence
Apache 2.0
GitHub stars
About 92.7k (vllm-project/vllm, 25 Sep 2026)
Latest release
v0.30.0 (22 Sep 2026)
Requirements
Linux (Windows via WSL); NVIDIA compute capability 7.5+, AMD ROCm, Intel, or Apple silicon via vLLM-Metal

What we like

  • Highest throughput when many users share one GPU
  • OpenAI-compatible server with tool calling and structured outputs
  • Apache 2.0 and very active (92.7k GitHub stars)
  • Scales across several GPUs

Watch out for

  • Linux only (Windows through WSL)
  • Needs a recent NVIDIA GPU, or extra setup for other hardware
  • No graphical app
#7 · One private server that replaces several cloud AI APIs

LocalAI

by LocalAI (open-source project led by Ettore Di Giacinto) · Open source · Free (open source)
8.3/10

LocalAI is a self-hosted, drop-in replacement for cloud AI APIs. It runs as one server that answers OpenAI-, Anthropic-, Ollama- and ElevenLabs-style requests, so existing apps usually need only a new URL. Behind that API it can switch engines per model: llama.cpp, vLLM or MLX for text, Whisper and Parakeet for speech, diffusers for images, plus text-to-speech, voice cloning and vision models.

That breadth is its selling point. One install can cover chat, transcription, image generation and voice for a small team or a home lab, and its model gallery lists more than 1,200 models. The project says every feature has a CPU path, so it works without a GPU, and it can spread work across several machines.

It is MIT-licensed and actively released (v4.10.0 on 17 September 2026). The trade-off is complexity: you set models up through config files and choose backends yourself, and it is less polished than Ollama or LM Studio for one person who only wants a chat window.

Pick it if you want one private server that replaces several cloud APIs. Skip it if you only need a local chat app.

Score breakdown

Ease of use6.8
Speed & hardware support8.0
Model support8.8
Features & integrations9.0
Licence & privacy9.8
Maintenance & community8.5

Key facts

Pricing
Free (open source) (MIT licence. Runs on CPU with no GPU required; a GPU makes it faster.)
Free option
Yes
Platforms
Linux, macOS, API
Licence
MIT
GitHub stars
About 49.3k (mudler/LocalAI, 25 Sep 2026)
Latest release
v4.10.0 (17 Sep 2026)
API compatibility
OpenAI, Anthropic, Ollama and ElevenLabs

What we like

  • Speaks OpenAI, Anthropic, Ollama and ElevenLabs APIs
  • Text, speech, image and video models in one server
  • Runs on CPU without a GPU
  • MIT licence

Watch out for

  • More setup than Ollama or LM Studio
  • Less polished for single-user chat
  • Speed varies with the backend you pick
#8 · A private, ChatGPT-style chat site for a family or team

Open WebUI

by Open WebUI Inc. · Free · Free (self-hosted)
8.2/10

Open WebUI is not a model runner. It is the best self-hosted chat interface to put in front of one: a ChatGPT-style web app that connects to Ollama, vLLM, any OpenAI-compatible server or cloud APIs. Run one Docker command and your household or team gets a shared, private chat site with accounts, permissions, document search (RAG), web search, voice, tools and plugins.

It has grown into an agent platform too. Its Open Terminal feature gives a model a real terminal and file system inside the chat, and it connects to autonomous agents. It has about 153,000 GitHub stars, and releases arrive often (v0.11.4 on 21 September 2026). A desktop app now exists for people who do not want Docker.

Check the licence before a big rollout. Since v0.6.6 Open WebUI uses its own licence: it is free and BSD-style, but deployments with more than 50 users in any 30 days must keep the Open WebUI name and logo unless they buy an enterprise licence. Contributors also sign a contributor licence agreement.

Pick it if you want a private ChatGPT-style site for a team on top of Ollama or vLLM. Skip it if you want a single desktop app with no server to run.

Score breakdown

Ease of use7.5
Speed & hardware support7.5
Model support8.5
Features & integrations9.6
Licence & privacy7.8
Maintenance & community9.3

Key facts

Pricing
Free (self-hosted) (Open WebUI License (BSD-3 style with a branding clause). Free to self-host; deployments with more than 50 users in any 30 days must keep the Open WebUI branding unless they buy an enterprise licence.)
Free option
Yes
Platforms
Web (self-hosted), Docker, Python, Desktop app
Licence
Open WebUI License: BSD-3 style plus a branding clause (since v0.6.6)
GitHub stars
About 153.1k (open-webui/open-webui, 25 Sep 2026)
Latest release
v0.11.4 (21 Sep 2026)
Connects to
Ollama, vLLM, any OpenAI-compatible API, Anthropic and more

What we like

  • Polished multi-user chat site that you host yourself
  • Works with Ollama, vLLM and any OpenAI-compatible API
  • RAG, web search, voice, tools and plugins built in
  • Very active project (153k GitHub stars)

Watch out for

  • Needs a separate model runner
  • Branding clause above 50 users
  • Full version needs Docker or Python setup
#9 · Mac users who want new models first or to fine-tune locally

MLX-LM

by Apple (ml-explore) · Open source · Free (open source)
7.9/10

MLX is Apple's own machine learning framework for Apple silicon, and MLX-LM is its Python package for running and fine-tuning language models with it. Install it with pip install mlx-lm, then generate text, chat, or start a local server (mlx_lm.server, with an API similar to OpenAI's) from the terminal. Thousands of ready-converted models sit in the mlx-community group on Hugging Face.

MLX is built around the Mac's unified memory, and Ollama, LM Studio and Jan all use the MLX framework on Apple silicon. Using MLX-LM directly gives you features the apps do not: LoRA and full fine-tuning on your Mac, quantizing and uploading your own models, and spreading one model across several Macs.

The downsides are reach and polish. It is built for Apple silicon (the core MLX framework also has Linux CUDA and CPU builds), there is no app, and its own docs say the built-in server is not meant for production. The last PyPI release was 0.31.3 in April 2026, though the code on GitHub was still being updated in September.

Pick it if you have an Apple silicon Mac and want the newest MLX features or to fine-tune locally. Skip it if you are on Windows or want a point-and-click app.

Score breakdown

Ease of use6.2
Speed & hardware support8.5
Model support8.5
Features & integrations7.5
Licence & privacy9.8
Maintenance & community8.0

Key facts

Pricing
Free (open source) (MIT licence. Install with pip install mlx-lm.)
Free option
Yes
Platforms
macOS (Apple silicon), Python, API
Licence
MIT
Maker
Apple's machine learning research team (ml-explore on GitHub)
Latest release
0.31.3 on PyPI (22 Apr 2026)
GitHub stars
About 7.1k (mlx-lm); 28.5k for the MLX framework

What we like

  • Apple's own framework, built for Apple silicon
  • Fine-tune models with LoRA on a Mac
  • Thousands of converted models on Hugging Face
  • MIT licence; simple local server included

Watch out for

  • Apple silicon only for most users
  • Command line and Python only
  • Built-in server not meant for production
#10 · Existing users on older hardware

GPT4All

by Nomic AI · Open source · Free (open source)
7.1/10

GPT4All was one of the first easy desktop apps for local AI, and it still installs and runs on Windows, macOS and Linux with a simple chat window. Its LocalDocs feature lets you chat with a folder of files on your computer, it has a local API server, and its MIT licence is as open as it gets. It also runs on old hardware: Nomic lists an Intel Core i3 2nd generation or AMD Bulldozer CPU as the minimum, and there is a build for Snapdragon Windows laptops.

The problem is that it has stopped moving. The last release, v3.10.0, came out on 25 February 2025, and the GitHub repository has had no new code since May 2025, even though users still open issues. Local AI changes every month, and a runner that has not been updated for 19 months is likely to miss newer model families and the faster engines that Ollama, LM Studio and Jan now use.

We keep it on the list because many older guides still recommend it. For now, it is not our pick for new users.

Pick it if it already works for you on an older machine. Skip it if you are starting fresh; choose Jan or LM Studio instead.

Score breakdown

Ease of use8.5
Speed & hardware support6.5
Model support5.5
Features & integrations7.5
Licence & privacy9.5
Maintenance & community3.0

Key facts

Pricing
Free (open source) (MIT licence. Free desktop app for Windows, macOS and Linux.)
Free option
Yes
Platforms
Windows, macOS, Linux
Licence
MIT
Latest release
v3.10.0 (25 Feb 2025)
Last code change
May 2025 (nomic-ai/gpt4all on GitHub)
GitHub stars
About 77.4k

What we like

  • Simple desktop app that runs on older CPUs
  • LocalDocs for chatting with your own files
  • MIT licence

Watch out for

  • No release since February 2025
  • May not support newer model families
  • No MLX engine on Macs

How we scored these tools

Each tool is scored 0–10 on the criteria below, using public evidence: independent benchmarks, vendor documentation and pricing pages, aggregate user ratings and reputable reviews. The overall score is the weighted average. Nobody pays to be listed. Read our full methodology.

CriterionWeightWhat we look at
Ease of use25%How quickly a non-expert can install it, find a model that fits their machine and start chatting or calling it.
Speed & hardware support20%Speed on common hardware, and support for NVIDIA, AMD and Intel GPUs, Apple silicon (Metal and MLX) and plain CPUs.
Model support15%Which model formats it loads (GGUF, MLX, Hugging Face weights) and how quickly new open models work.
Features & integrations15%Local API server, document chat (RAG), tools and MCP, multi-user support and how many other apps connect to it.
Licence & privacy15%Open-source licence, terms for work use, telemetry, and whether everything can run offline.
Maintenance & community10%Release pace, GitHub activity, backing and community size, checked on 25 September 2026.

What hardware do you need to run an LLM locally?

Memory decides what you can run. The whole model has to fit in your graphics card's video memory (VRAM) or, on an Apple silicon Mac, in its unified memory, with a few gigabytes spare for the conversation. The download sizes below are Ollama's default builds (mostly 4-bit), checked on 25 September 2026.

Your machine Usable memory What runs well Example download size
Laptop with 8GB RAM 8GB Small models (2B-4B) Gemma 4 E2B: 4.3GB (QAT build)
Laptop or Mac with 16GB 16GB Models up to about 12B-20B Gemma 4 12B: 7.6GB; gpt-oss-20b: 14GB
PC with a 24GB GPU, or a 32GB Mac 24-32GB Strong mid-size models Qwen3.8 27B: 18GB; Gemma 4 31B: 20GB
Mac Studio, DGX Spark or a 96-128GB workstation 96-128GB The biggest local models gpt-oss-120b: 65GB

Rule of thumb: take the download size and add 2-4GB for the context (the conversation the model keeps in memory). Very long contexts need more. Higher-precision builds are much bigger: Qwen3.8 27B grows from 18GB at the default setting to 30GB in 8-bit and 56GB at full 16-bit precision.

Minimums by tool: LM Studio needs an Apple silicon Mac on macOS 14 or newer (Intel Macs are not supported), or a Windows or Linux PC; on x64 Windows the CPU must support AVX2, and it recommends 16GB of RAM and 4GB of VRAM. Ollama needs macOS 14 or Windows 10 22H2 or newer, and uses NVIDIA cards with compute capability 5.0 and up, plus AMD cards through ROCm or Vulkan. vLLM needs Linux (or WSL on Windows) and an NVIDIA GPU with compute capability 7.5 or newer. GPT4All runs on CPUs as old as an Intel Core i3 2nd generation.

Buying hardware? See our best GPUs for AI and best AI laptops rankings. To pick a model that fits, see best local LLMs and best small language models.

Is Ollama free? Ollama pricing explained

Yes. Running models on your own computer with Ollama is free and unlimited, and the software is MIT-licensed. You only pay if you use Ollama's cloud models, which run on Ollama's servers and suit models too big for your machine. Since 31 August 2026 those plans use per-token pricing with a monthly usage allowance:

Plan Price Included cloud usage Concurrent cloud requests
Free $0 Starter credits for a small set of models 1
Pro $20/month or $200/year $60 a month 3
Max $100/month $300 a month 10
Team $500/month $1,000 a month, shared, unlimited users 10
Enterprise Custom Custom Custom

Cloud usage is billed at published rates, for example $0.15 / $0.60 per million input / output tokens for gpt-oss:120b and $1.40 / $4.40 for GLM-5.3. Ollama says cloud prompts are never logged or trained on, and are processed in the US and Europe (plus Singapore for some Qwen models).

LM Studio works the same way: the app and local models are free, and Bionic+ ($20/month) or Pro ($100/month) add cloud models. Every other tool on this page is free, apart from AnythingLLM's optional hosted cloud (from $50/month) and Open WebUI's enterprise licence for large rebranded deployments.

Licences and work use: what you are allowed to do

All ten tools are free to use at work, but the terms differ if you want to change them, rebrand them or build them into a product.

Tool Licence Watch out for
Ollama MIT Cloud models send prompts to Ollama's servers
LM Studio Proprietary, free Personal and internal business use only; no modifying, reverse engineering or reselling as a hosted service
llama.cpp MIT Nothing unusual
Jan Apache 2.0 Nothing unusual
AnythingLLM MIT Hosted cloud is a separate paid service
vLLM Apache 2.0 Nothing unusual
LocalAI MIT Nothing unusual
Open WebUI Open WebUI License (BSD-3 plus branding clause) Keep the branding above 50 users in 30 days, or buy an enterprise licence
MLX-LM MIT Nothing unusual
GPT4All MIT No updates since 2025

Remember that each model has its own licence too. OpenAI's gpt-oss models use Apache 2.0, for example, while other model families have their own terms. Check the model card before you use a model in a product.

On privacy, all ten can run fully offline once a model is downloaded. LM Studio's privacy policy says local chats never leave your device and the app has no telemetry, and Jan only collects usage analytics if you opt in.

Ollama alternatives: which one should you pick?

These tools do different jobs, and many people combine two of them. Pick by what you want to do:

  • You want an app with menus, not a terminal: LM Studio. Choose Jan instead if the app must be open source.
  • You want to chat with your PDFs and notes: AnythingLLM, with its built-in engine or on top of Ollama.
  • You want a shared ChatGPT-style site for a family or team: Open WebUI in front of Ollama or vLLM.
  • You want the most control, or the newest model on day one: llama.cpp. On a Mac, MLX-LM.
  • You are serving an app or many users: vLLM. If you also need speech, images and voice from one server, LocalAI.
  • You just want it to work, and other apps to find it: stay with Ollama.

A common setup is an engine (llama.cpp, MLX or vLLM) wrapped by a runner (Ollama, LM Studio or LocalAI) with a front end on top (Open WebUI, AnythingLLM or Jan). If your hardware is too small for the model you need, a hosted API may be simpler: see our best AI inference providers ranking.

At a glance: platforms, APIs and activity

Tool Type Platforms Local API GitHub stars (25 Sep 2026) Latest release
Ollama Runner and app Mac, Windows, Linux OpenAI- and Anthropic-style 181.7k v0.34.4 (23 Sep 2026)
LM Studio App and runner Mac (Apple silicon), Windows, Linux OpenAI- and Anthropic-compatible Closed source (lms CLI: 5.3k) 0.4.25
llama.cpp Engine and server Mac, Windows, Linux OpenAI-compatible 129.5k v0.5.0 (23 Sep 2026)
Jan Desktop app Mac, Windows, Linux Local API server 44.6k v0.8.4 (23 Jul 2026)
AnythingLLM App with RAG and agents Mac, Windows, Linux, Android, Docker Developer API 66.5k v1.16.2 (22 Sep 2026)
vLLM Serving engine Linux (Windows via WSL) OpenAI-compatible 92.7k v0.30.0 (22 Sep 2026)
LocalAI API server Linux, Mac OpenAI, Anthropic, Ollama, ElevenLabs 49.3k v4.10.0 (17 Sep 2026)
Open WebUI Web front end Any, via Docker, Python or desktop app Connects to other runners 153.1k v0.11.4 (21 Sep 2026)
MLX-LM Python package Mac (Apple silicon) Similar to OpenAI's 7.1k 0.31.3 (22 Apr 2026)
GPT4All Desktop app Windows, Mac, Linux Local API server 77.4k v3.10.0 (25 Feb 2025)

GitHub stars measure interest, not quality, but release dates tell you whether a tool will support next month's models.

Expert tips
  1. Check the download size before you pull a model: it should be at least 2-4GB smaller than your VRAM, and on a Mac you also need to leave room for macOS and your open apps. If it does not fit, pick a smaller model or a lower-bit build rather than letting it spill onto the CPU.
  2. On an Apple silicon Mac, prefer MLX builds. Ollama's own test on an M5 Max showed output speed nearly doubling (58 to 112 tokens per second) after it moved to MLX, and LM Studio and Jan offer MLX models too.
  3. Raise the context length only when you need it. Ollama and LM Studio let you set it per model, and doubling it can add gigabytes of memory use and slow every reply.
  4. Run Ollama or LM Studio's server once and point everything at it: Open WebUI, AnythingLLM, your code editor and coding agents can all share one loaded model instead of each loading its own copy.
  5. If you roll Open WebUI out to more than 50 people, keep its branding or budget for an enterprise licence; the licence requires it above that size.

Jargon explained

Local LLM
A language model that runs on your own computer instead of a company's servers, so it works offline and your prompts stay private.
VRAM
The memory on a graphics card. The model has to fit in it (or in an Apple silicon Mac's shared memory) to run quickly.
GGUF
The standard file format for local models, created by the llama.cpp project. Most tools on this page can load GGUF files.
MLX
Apple's machine learning framework for Apple silicon Macs. MLX builds of a model usually run faster on a Mac than GGUF builds.
Quantization
Storing a model's numbers with fewer bits (for example 4-bit instead of 16-bit) so it is smaller and faster, with a small loss in quality.
RAG (retrieval-augmented generation)
Letting a model search your own documents and use what it finds in its answer. AnythingLLM, Open WebUI and GPT4All have it built in.

Frequently asked questions

What is the best tool to run LLMs locally?

For most people, Ollama. It is free, MIT-licensed, works on Mac, Windows and Linux, starts a model with one command, and almost every other AI app can connect to it. If you prefer a point-and-click app, use LM Studio.

What are the best Ollama alternatives?

LM Studio (easiest app), Jan (open-source desktop app), AnythingLLM (chat with your documents), llama.cpp (most control), vLLM (serving many users) and LocalAI (one server for text, speech and images). Open WebUI is not a replacement but a chat interface you put in front of Ollama or vLLM.

Is Ollama free?

Yes. Running models locally with Ollama is free and unlimited, and the code is MIT-licensed. Ollama only charges for its optional cloud models: Pro is $20/month (with $60 of usage), Max $100/month ($300 of usage) and Team $500/month.

Is LM Studio free for commercial use?

Yes, for use inside your business. LM Studio has been free for work since 8 July 2025, and its terms allow personal and internal business use. It is closed source, so you may not modify it, reverse engineer it or resell it as a hosted service. Paid Bionic+ and Pro plans add cloud models.

Ollama vs LM Studio: which is better?

Ollama is better as a background service that other apps and coding tools connect to, and it is open source. LM Studio is better as a desktop app for browsing, downloading and chatting with models without a terminal. Both use llama.cpp and Apple's MLX under the hood, so the choice is mostly about how you like to work.

How much RAM do I need to run an LLM locally?

About the model's download size plus 2-4GB. With 8GB you can run small 2B-4B models; 16GB handles models such as Gemma 4 12B (7.6GB) or gpt-oss-20b (14GB); 24-32GB of VRAM or Mac memory fits Qwen3.8 27B (18GB); and gpt-oss-120b (65GB) needs a 96-128GB machine. See our best local LLMs ranking for models that fit.

Can I run a local LLM without a GPU?

Yes, but it is slower. llama.cpp, Ollama, LocalAI and GPT4All all run on a normal CPU, and LocalAI says every feature has a CPU path. Small models are usable on a CPU; bigger ones are much faster on a GPU or an Apple silicon Mac.

Is running an LLM locally private?

Yes, if you use a local model. Once the model is downloaded, all ten tools here can run offline, so your prompts never leave the machine. Be careful with optional cloud features: Ollama's cloud models, LM Studio's Bionic cloud and any cloud API you add to Jan or AnythingLLM send prompts to a server.

Sources

Every fact on this page comes from public information. Vendor figures are labelled as vendor claims.

  1. Ollama pricing (Ollama)
  2. Ollama's transparent pricing (Ollama)
  3. Ollama is now powered by MLX on Apple silicon (preview) (Ollama)
  4. Ollama blog (funding, releases) (Ollama)
  5. Ollama's new app (Ollama)
  6. Ollama on Windows: system requirements (Ollama)
  7. Ollama on macOS: system requirements (Ollama)
  8. Ollama hardware support (Ollama)
  9. Ollama OpenAI compatibility (Ollama)
  10. Ollama gpt-oss model tags (Ollama)
  11. Ollama qwen3.8 model tags (Ollama)
  12. Ollama gemma4 model tags (Ollama)
  13. ollama/ollama on GitHub (GitHub)
  14. LM Studio pricing (LM Studio)
  15. LM Studio is free for use at work (LM Studio)
  16. LM Studio desktop app terms of service (LM Studio)
  17. LM Studio desktop app privacy policy (LM Studio)
  18. LM Studio system requirements (LM Studio)
  19. LM Studio OpenAI compatibility endpoints (LM Studio)
  20. Introducing LM Studio Bionic (LM Studio)
  21. LM Studio download page (LM Studio)
  22. llama.app: official home for llama.cpp (ggml-org)
  23. llama.cpp releases (GitHub)
  24. GGML and llama.cpp join Hugging Face (Hugging Face)
  25. NVIDIA to acquire Hugging Face (NVIDIA)
  26. Jan homepage (Menlo Research)
  27. Jan desktop docs (Menlo Research)
  28. Jan's privacy approach (Menlo Research)
  29. janhq/jan releases (GitHub)
  30. AnythingLLM homepage (Mintplex Labs)
  31. AnythingLLM cloud pricing (Mintplex Labs)
  32. AnythingLLM default (built-in) LLM (Mintplex Labs)
  33. vLLM installation and hardware support (vLLM)
  34. vLLM GPU installation requirements (vLLM)
  35. Efficient Memory Management for LLM Serving with PagedAttention (arXiv)
  36. vllm-project/vllm on GitHub (GitHub)
  37. LocalAI homepage (LocalAI)
  38. mudler/LocalAI releases (GitHub)
  39. Open WebUI documentation (Open WebUI)
  40. Open WebUI license (Open WebUI)
  41. open-webui/open-webui on GitHub (GitHub)
  42. MLX LM README (GitHub)
  43. MLX LM HTTP server docs (GitHub)
  44. mlx-lm on PyPI (PyPI)
  45. MLX framework on GitHub (GitHub)
  46. nomic-ai/gpt4all on GitHub (GitHub)
  47. GPT4All API server docs (Nomic AI)