You do not need a powerful GPU to use DeepSeek online. Its website, mobile assistant and hosted API run models on remote servers. Hardware requirements become a different question when you download a model and run it on your own PC or Mac.
For a first local setup, start with DeepSeek-R1 8B on a computer with 16 GB of RAM. If you have 8 GB, try the smaller 1.5B model first. These are planning recommendations for the quantized R1 models below, not requirements for the full R1 model, DeepSeek-V4.1-Flash or V4 Pro. A dedicated graphics card can help, but CPU-only use is possible when the model fits in system memory.
Jump to the RAM and GPU table, full R1 and V4 models, or the short local setup.
Which DeepSeek are you trying to run?
For official DeepSeek Chat, your device runs the browser while DeepSeek runs the model. The same distinction applies to the official mobile assistant. You need internet access and a compatible browser or app, not enough local memory to hold the model. Developers using the hosted API also avoid local model hardware, although they still need the usual account, key and service access.
DeepSeek Harness is a separate desktop product. Installing a desktop client does not mean you have installed the full model for offline inference. Its default official-model service requires a DeepSeek account and prepaid Open Platform funds, as explained in the Harness terms. Our download guide covers the supported desktop and mobile versions.
The table below is for downloaded models running through Ollama. Before choosing one, check your computer’s RAM, free disk space and, if present, the dedicated GPU’s VRAM. A PC with 16 GB of RAM and an 8 GB graphics card does not have a single 24 GB memory pool.
Choose a model that fits your memory
Start with the memory you actually have. On a Windows or Linux PC with a dedicated graphics card, system RAM and the card’s VRAM are separate. Integrated graphics can use system RAM instead. On an Apple silicon Mac, the CPU and GPU share unified memory, which also has to support macOS and your other apps.
The table gives starting recommendations with spare memory, not the lowest workable capacities. These are our planning estimates, not official minimum requirements or hardware test results. They assume a Q4 model, one conversation and a modest context of around 4,096 tokens.
| Ollama model | Download | RAM for CPU use or Mac unified memory | Dedicated VRAM to aim for |
|---|---|---|---|
deepseek-r1:1.5b | 1.1 GB | 8 GB | A GPU is optional for an initial trial |
deepseek-r1:8b | 5.2 GB | 16 GB | 8 GB |
deepseek-r1:14b | 9 GB | 24–32 GB | 16 GB |
deepseek-r1:32b | 20 GB | 48–64 GB | 24–32 GB; 24 GB leaves less room for context |
deepseek-r1:70b | 43 GB | 96 GB or more | 64–80 GB or more |
Download sizes come from Ollama’s model listings, checked on October 4, 2026. Read the memory columns as alternative paths: CPU or Apple silicon on the left, a supported dedicated GPU on the right. A GPU-equipped PC still needs system RAM, but it does not also need to meet the CPU-only recommendation.
On a phone, swipe the table sideways to see all four columns.
Q4 stores model weights at roughly four-bit precision to reduce their memory footprint. Quantization can also change answer quality; the smaller file is not an identical full-precision copy. The estimates above do not apply to every download with the same parameter count. For example, Ollama lists the 8B Q4 package at 5.2 GB and its Q8 version at 8.9 GB. Check the format before downloading.
A larger model is worth trying if the smaller one struggles with your task and you have spare memory. Size alone does not tell you whether its answers will be useful for your work.
How RAM, VRAM and context affect the choice
The download size covers the model files. Running the model also takes memory for the conversation and the software itself. That is why a 5.2 GB download should not be treated as a promise that everything will fit on a 6 GB graphics card.
The same distinction matters for larger models. The 32B Q4 package is about 20 GB, so its weights cannot fit entirely in 16 GB of VRAM. Ollama may put part of the model in system RAM instead. This is called offloading; it makes more configurations possible, but a model that loads this way may respond more slowly than you want.
For memory planning, a 24 GB RTX 4090 is a candidate for the listed R1 32B Q4 package at a modest context. A 32 GB RTX 5090 gives that package more memory headroom. Check Ollama’s hardware support list for your card and operating system, especially with AMD hardware.
On Apple silicon, use the unified-memory figures; on Intel Macs, use the CPU RAM figures. Ollama supports GPU acceleration on Apple silicon, while Intel Macs run on the CPU. Ollama’s Mac requirements explain the supported systems.
Context is the amount of text the model can keep available while answering. A longer document or conversation needs more of it, which increases memory use. Start with a short conversation, then raise the context when the task calls for it. The maximum shown on a model’s download page is not a promise that your computer can run that entire context. See Ollama’s context settings.
What about full R1, DeepSeek-V4.1-Flash and V4 Pro?
The compact R1 models above are distilled models: smaller models trained using outputs from a larger one. The current 8B Ollama package is DeepSeek-R1-0528-Qwen3-8B. It is a separate model from full R1 and from the newer models served through DeepSeek’s hosted products.
For scale, Ollama’s full 671B R1 Q4 package is about 404 GB before runtime memory. That is a different hardware project from running an 8B model on a laptop. The 8 GB and 16 GB starting recommendations above do not apply to it.
DeepSeek-V4.1-Flash is not a small 8B download. Its official model card lists a 552B-parameter backbone and a separate 196B-parameter Engram component. The often-quoted 8B/16B figures describe parameters activated during different stages of processing. They do not mean that the rest of the model disappears from the storage and memory plan.
Flash’s reference inference instructions describe weight conversion and parallel execution, using eight parallel ranks in the example. They do not establish a universal minimum GPU-memory figure. Treat that example as a deployment reference, not a promise that any eight GPUs will work.
For DeepSeek-V4-Pro-0813, the official model card gives a vLLM serving example on a single node with four GB300 GPUs. That is one documented configuration, not a statement that four GPUs are the minimum or that four consumer cards are equivalent. The card links to the serving engines’ recipes for other configurations.
Before planning a full-model deployment, match the exact checkpoint to a supported inference engine, weight format and hardware configuration. Leave capacity for the context, concurrent requests and runtime as well as the weights. A guide for an older V4 preview or a different quantization may describe a different requirement.
For a personal PC, the smaller R1 models are the practical starting route in this guide. If you want access to current hosted models without buying server hardware, use DeepSeek Chat or follow our API setup guide. The model overview explains how the model families relate.
Start a local model with Ollama
Download Ollama from its official site. On Windows, you need Windows 10 22H2 or newer; on a Mac, you need macOS Sonoma 14 or newer. Linux users can follow the official installation instructions.
Check free disk space before downloading. Allow for the model, Ollama and future updates. On Windows, the Ollama installation itself requires at least 4 GB, separate from the model files listed above.
Once Ollama is installed and running, open PowerShell on Windows or Terminal on macOS or Linux and enter:
ollama run deepseek-r1:8b
This downloads the model if needed and opens a chat. Replace the model name with your choice from the table; for an 8 GB computer, use deepseek-r1:1.5b. You do not need a DeepSeek API key. Try a short task of your own before moving to a larger model.
Prefer choosing a model through a graphical interface? Follow the separate DeepSeek setup guide for LM Studio.
If the model will not run—or runs too slowly
Not enough memory: close other memory-heavy applications. Ollama may use a larger context than the 4K assumed in the table. If the model loads but memory is tight, enter /set parameter num_ctx 4096 inside the Ollama chat, before your next prompt. This is an Ollama chat command, not a PowerShell or system-terminal command. Reducing the context frees working memory without shrinking the model’s weights. If the model cannot load at all, start with a smaller one.
Slow responses: with the model still loaded, open another terminal window and run ollama ps. As the Ollama FAQ explains, under PROCESSOR, 100% GPU means the model is loaded on the GPU; a CPU/GPU split means it uses both. A smaller model may fit entirely on the GPU and respond faster. If it shows 100% CPU when you expected GPU use, check hardware support and drivers.
The command is not found: close and reopen the terminal after installation. If Ollama cannot connect, make sure the application or service is running.
Questions before you download
Can I run DeepSeek offline?
Yes, for a downloaded model running locally with all required files already installed. The R1 tags in this guide use local inference. A cloud model or a tool that searches the web still needs a connection. Ollama also documents a local-only mode that disables its cloud models and web search. DeepSeek’s website and hosted API are online services.
Will a bigger SSD let me run a larger model?
It gives you room to store larger downloads, but disk space does not replace the working memory in the table. Ollama lets you choose a different model-storage location through OLLAMA_MODELS, as described in its storage instructions. A model kept on a larger drive still needs sufficient RAM or VRAM when you run it.
