Qwen includes models you can run on your own hardware for answering questions about documents, editing code, reading scanned pages, and handling classification tasks. Each puts different demands on the model and the machine.
Start with one task, a few representative inputs and the hardware you intend to use. Then choose a checkpoint you can test under those conditions. The examples below cover selected branches of the Qwen family, with model cards you can use to check the details.
Choose the model around its inputs
For text work, the original Qwen3 release includes dense models from 0.6B to 32B parameters, alongside mixture-of-experts models. Its 4B and 8B checkpoints are candidates for an initial local experiment when memory is limited. The release also introduced thinking and non-thinking modes, so include the selected mode in any comparison. Qwen3 release notes.
For software work, consider a coding checkpoint. Qwen3-Coder-30B-A3B-Instruct is designed for coding and tool use, and its model card specifies a non-thinking model. A practical trial might ask it to fix a bug exposed by a failing test in a small repository, with the same files and tools available on every run. Qwen3-Coder model card.
Images require a model that accepts them. Qwen3-VL-8B-Instruct supports image-and-text input; it is a candidate for experiments with screenshots, diagrams and scanned documents. Evaluate the actual page images, including rotated text and poor scans, rather than judging it on a clean example alone. Qwen3-VL model card.
Naming conventions change between generations. Qwen3.5-9B includes a vision encoder without a “VL” suffix. Read the exact model card before deciding what an unfamiliar checkpoint can accept. Qwen3.5-9B model card.
For document search, Qwen3-Embedding and Qwen3-Reranker provide separate models in 0.6B, 4B and 8B sizes. An embedding model represents queries and passages as vectors; a reranker scores retrieved passages against the query. In a document assistant, test retrieval separately from the model that writes the answer. Otherwise, a missing passage can look like a generation failure. Qwen3 embedding and reranking models.
Read both numbers in an MoE model
A mixture-of-experts model routes each token through a subset of its expert networks. Dense models have no such expert routing. Qwen3-30B-A3B has 30.5 billion total parameters and 3.3 billion active parameters, according to its model card. The active count describes the computation selected for a token; the complete set of weights still needs storage. Qwen3-30B-A3B model card.
The Qwen3-Coder-Next model card lists 80 billion total parameters and 3 billion active. Plan for the full set of weights when choosing storage and memory. A runtime may distribute or offload weights, but that changes the deployment and needs measuring on the intended machine. Qwen3-Coder-Next model card.
Avoid selecting an MoE model by the smaller number alone. Record total parameters, the downloaded file size and observed memory consumption alongside response time.

Budget for the whole session
Quantisation reduces the precision used to store model weights, reducing their memory footprint. Qwen documents several quantisation options for local inference with llama.cpp. Qwen local-inference guide.
As a rough calculation, eight billion parameters at four bits each occupy four billion bytes before metadata and other overhead. That is a weight-storage estimate. You need additional memory to run the model.
The runtime also stores state while processing the conversation. Longer inputs can increase that requirement. Ollama explicitly warns that increasing context length increases memory use, and provides controls for setting it. Start with enough context for your test documents and increase it when the task requires it. Ollama context-length documentation.
Keep the quantisation fixed during an initial comparison. Changing the checkpoint, precision and context length together makes it difficult to explain a change in quality or speed. If a smaller download is necessary, rerun the same task set after changing precision.
Keep the first installation reproducible
Ollama provides a local route to trying Qwen3-8B after installing the runtime:
ollama run qwen3:8b
The command downloads the model if needed and starts an interactive session. Use this example for a text-model trial on hardware with enough available memory. Check the model entry before downloading. Qwen3-8B in Ollama.
For a repeatable experiment, save the exact model identifier and file digest, runtime version, quantisation, prompt template, context limit and generation settings. Include thinking mode where the checkpoint supports it. A colleague should be able to recreate the configuration without guessing which “Qwen” you meant.
Check compatibility before choosing a different family member. A runtime that loads one Qwen architecture may need an update for another, and an application must support the input types you plan to send.
Define where local processing ends
Ollama says it does not receive prompts or data from local runs. It also offers cloud models and web search, and documents a local-only setting that disables those features. If keeping inference on the machine is a requirement, configure that explicitly. Ollama privacy and local-only settings.
Then review the application around the model. A locally hosted model can still be connected to remote search, document storage, embeddings or monitoring. Trace a sample document through the complete workflow and inspect what leaves the machine. Decide which prompts and outputs to retain, who can access them, and when to delete them.
For a coding agent, include its tools in that review. The ability to generate a shell command and permission to execute it are separate configuration decisions.
Evaluate the work you expect it to do
Build a small test set before tuning prompts. For document extraction, include missing fields, ambiguous dates and tables split across pages. For coding, use changes with executable tests. For questions over internal documents, include questions whose answers are absent from the supplied material.
Define an acceptable result for each example. Track correctness, required output format, time to first output, total completion time and peak memory. Repeat difficult cases so that one successful answer does not determine the choice.
Review the mistakes with the person who would use the output. Record which require human correction and which should stop the workflow. Keep those examples for the next checkpoint or runtime upgrade, and run them again before changing the deployed configuration.