     [Blog](https://scrapfly.io/blog)   /  [ai](https://scrapfly.io/blog/tag/ai)   /  [How to Run a Local LLM with Ollama](https://scrapfly.io/blog/posts/guide-to-local-llm)   # How to Run a Local LLM with Ollama

 by [Mostafa Gouda](https://scrapfly.io/blog/author/mostafa) Oct 01, 2026 16 min read [\#ai](https://scrapfly.io/blog/tag/ai) 

 [  ](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fguide-to-local-llm "Share on LinkedIn") [  ](https://x.com/intent/tweet?url=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fguide-to-local-llm&text=How%20to%20Run%20a%20Local%20LLM%20with%20Ollama "Share on X") [  ](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fguide-to-local-llm "Share on Facebook")    

 

 

Summarize this article with

 [  ](https://chat.openai.com/?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fguide-to-local-llm) [  ](https://claude.ai/new?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fguide-to-local-llm) [  ](https://x.com/i/grok?text=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fguide-to-local-llm) [  ](https://www.perplexity.ai/search/new?q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fguide-to-local-llm) [  ](https://www.google.com/search?udm=50&aep=11&q=Summarize%20this%20article%20and%20explain%20how%20Scrapfly%20helps%20me%20scrape%20any%20website%20at%20scale%20and%20bypass%20anti-bot%20systems%20for%20my%20use%20case%3A%20https%3A%2F%2Fscrapfly.io%2Fblog%2Fposts%2Fguide-to-local-llm) 



   

We downloaded `qwen3.5:9b` in Q4\_K\_M quantization as a 6.6 GB file, loaded it with 8,192-token context, and ran it to extract a title and price from one product page in the public domain via Ollama. But before you do this again, what matters isn't which model has topped the leaderboard. It is rather whether the quantized model size and context you give it fit within the available Ollama memory on your machine.



## Key Takeaways

- `qwen3.5:9b` at Q4\_K\_M is the tested starting point in this guide, not a claim that one model is best for every local LLM task.
- Model weights and context length use separate memory budgets. Start with the context your task needs, then inspect the loaded process with `ollama ps`.
- A GPU is optional. Ollama can use system RAM, but CPU use or partial offload normally reduces response speed.
- Ollama keeps the model inference local in this workflow. Scrapfly remains a cloud collection service, so the target URL and scrape configuration leave the machine.
- The demonstrated Python flow returned the expected product title and price for one public demo page. It does not guarantee results for other pages or schemas.

**Get web scraping tips in your inbox**Trusted by 100K+ developers and 30K+ enterprises. Unsubscribe anytime.







## What Is a Local LLM, and What Actually Stays Local?

A local LLM is a set of model weights loaded and executed on hardware you control, rather than sent to a hosted API for inference. When you run Ollama, it starts a [local server](https://docs.ollama.com/quickstart.md), normally at `http://localhost:11434`, and every inference request goes through that address instead of leaving your machine.

Running a model locally hands you a specific set of controls

- Which model tag you pull
- Which quantization you load
- How much context you allocate to it
- When you decide to update it.

None of that depends on a provider's release schedule or rate limit.



With Ollama, prompts and model inference stay on the local machine, while a hosted API sends the prompt across the network for processing on a provider’s infrastructure.Local inference is not the same claim as a fully local application. A local LLM covers the step where the model reads tokens and generates a response. Files on disk, third-party APIs, telemetry a library sends home and any web content you fetch from the internet are separate concerns with their own data paths.

A short example makes the boundary concrete. Text you type directly into a prompt and pass to Ollama stays on the local inference path the whole way through. A URL fetched through a cloud scraping or search API does not. The request and the fetched content cross the network before your local model ever sees them.

Retrieval-augmented generation adds a further retrieval and indexing layer on top of local inference, covered separately in

[How to Power-Up LLMs with Web Scraping and RAGIn depth look at how to use LLM and web scraping for RAG applications using either LlamaIndex or LangChain.](https://scrapfly.io/blog/posts/how-to-use-web-scaping-for-rag-applications)



## How Much Memory Does an Ollama Local LLM Need?

Memory fit depends on

- The quantized weight file you load
- The context you allocate
- Runtime overhead, whether any layers offload to CPU
- How much you run in parallel

Parameter count alone tells you very little on its own.

The architecture shown in this tutorial is the [qwen3.5:9b listing](https://ollama.com/library/qwen3.5:9b), which is a 9.7 billion parameters model deployed at Q4\_K\_M quantization with a file size of 6.6 GB, using an 8,192-token context for a one-page extraction task.

This listing promises a much larger maximum context of 262,144 tokens. However this tutorial does not suggest using such an extremely large value as the default setting. Large context values are reserved for tasks requiring them.

The model memory refers to the VRAM or unified memory that Ollama can allocate to the model process, not the total RAM of the machine. Even CPU-only machines can run `qwen3.5:9b`. In this case Ollama adapts to the system RAM with a lower response speed than the GPU or unified memory solution.

For any value you figure out for the model always keep some headroom for the operating system and the Python process running LangChain. Consider this memory usage as just one element in a larger budget.

### How Do Model Weights and Q4\_K\_M Quantization Affect Memory?

Parameter count is a measure of the scale of the model architecture. Quantization refers to the precision of storage of individual weights after a model has been compressed for local deployment. These two concepts relate to different aspects and only when they come together you can know the storage requirements of the model.

Example illustration of the conservative lower bound: storing 10 billion weights using 4 bits per weight gives you about 5 GB in size before any file headers and runtime overheads. This math is the lower bound, but definitely not the size you will get from Ollama’s reporting itself.

Q4\_K\_M is not an exact representation of four-bit quantization. It is a mixed precision scheme and the actual storage size would be slightly higher as in `qwen3.5:9b` at Q4\_K\_M which is 6.6 GB. Use the reported download size in Ollama for fit planning.

### How Does Ollama Context Length Affect Memory?

Context is a separate runtime budget from the weight file. It covers the input tokens you send and the tokens the model generates in response, and Ollama allocates memory for that budget in addition to the loaded weights.

[Ollama's context length documentation](https://docs.ollama.com/context-length.md) states the general rule plainly, a larger allocated context requires more memory, independent of which model you run. A model that loads comfortably at an 8K context can fail to load, or fail to leave room for anything else, at a much larger allocation on the same hardware.

Let the task set the number instead of the model's maximum. The single-page extraction demonstrated later in this guide used an 8,192-token context, which comfortably covers one page of Markdown and a short JSON response. Raise it only when measured input length or output needs justify the change not by default.

### What Does `ollama ps` Show About Context and Offload?

There are two commands that describe what is really going on after your model was loaded

1. `ollama show qwen3.5:9b` describes the metadata of the loaded model, such as parameters number, quantization type and templates.
2. `ollama ps` shows the context size that was allocated by Ollama for the process and its distribution between GPU and CPU.

Splitting model into CPU/GPU where part of the model is loaded to system RAM and other is loaded into GPU, can be exactly what allows your model to fit in hardware which is not able to accommodate all layers of the model in VRAM. However, the price should be paid for the above fitting of the model.

The response speed is usually reduced when any part of the model is calculated by CPU but not GPU. The exact timings obtained on the machine that is used for the above guide are omitted on purpose.

| Available model memory | Practical starting point | Context starting point | Required qualifier |
|---|---|---|---|
| 4 to 8 GiB | 1B to 4B model at Q4 | 4K to 8K | Starting guidance, not a guarantee |
| 8 to 16 GiB | 7B to 10B model at Q4 | 8K | The demonstrated 9B path belongs here by fit planning, not by a benchmark across machines |
| 16 to 32 GiB | 14B to 20B model at Q4, or a smaller model with more context | 8K to 32K | Architecture and KV-cache settings change the result |
| 32 GiB or more | 30B-class Q4 model, subject to architecture and offload | Start at 8K | Raise context only for a measured need |

These are planning ranges derived from quantized-weight size and memory headroom not benchmarked guarantees.



## How Do You Install Ollama and Run `qwen3.5:9b`?

Start from the [official Ollama download page](https://ollama.com/download), which covers Linux, macOS, and Windows installers. On Linux, the installer does not always start the background service automatically, so run `ollama serve` first if a later command reports that it cannot reach the server.

Confirm the install before pulling anything

bash```bash
ollama --version
```



Pull the exact tag used in this guide

bash```bash
ollama pull qwen3.5:9b
```



That downloads the 6.6 GB Q4\_K\_M artifact described in the memory section above. Once the pull finishes, start the model

bash```bash
ollama run qwen3.5:9b
```



This opens an interactive prompt against the local model. Leave that session running, and check two things about it from a second terminal. First, model metadata

bash```bash
ollama show qwen3.5:9b
```



This confirms the model tag, parameter count, and quantization that actually loaded, not just the one you requested. Second, the live process

bash```bash
ollama ps
```



`ollama ps` reports the context size and the CPU/GPU split Ollama chose for the running process, which is the number to trust over any assumption you made before loading the model.

The workflow documented here ran on Ollama 0.32.5 with `qwen3.5:9b` at Q4\_K\_M. Later compatible releases of Ollama or the model tag could work the same way, but treat that as your own verification step.

This section only covers the native Ollama installer and CLI. Running the model through Docker, LM Studio, llama.cpp directly, or Ollama's own hosted Cloud product follows a different setup path with different memory and network characteristics which is out of scope here.

None of the commands above prove anything about output quality either. A chat response that reads well in the interactive prompt is not evidence of correct extraction, which is what the next two sections actually check.



## How Do Scrapfly and Ollama Split Cloud Collection and Local Inference?

Ollama runs inference entirely on your machine. Scrapfly is a separate cloud collection layer. It takes a target URL and scrape configuration, fetches the page in its cloud, and returns Markdown, a format built for clean [handoff to LangChain](https://scrapfly.io/docs/integration/langchain). That split doesn't change once you wire both together in Python.

Three options describe how the content gets to your local model in the context of the integration, regardless of any legal or policy considerations

- Content that you have locally moves to Ollama directly. No collection in the cloud happens in this case.
- Public URL is provided to Scrapfly, which collects the page. Then Scrapfly returns the content in Markdown, which then goes to your local Ollama model.
- Hosted inference service that you use instead of Ollama receives your prompt and any content you provide within its own data boundary and not Ollama’s.

Apply your organization's data policy before you send a sensitive URL or sensitive local content to any external collection service including Scrapfly. Scrapfly is a collection layer in this workflow. It is not a local runtime, and it is not the model provider. The `qwen3.5:9b` weights and the inference step stay entirely with Ollama.

| Workflow step | Where it runs | Data that crosses the boundary | Output |
|---|---|---|---|
| Ollama inference | Reader machine at `localhost` | None to a hosted model provider in this workflow | Model response |
| Scrapfly collection | Scrapfly cloud | Target URL and scrape configuration | Page Markdown |
| Python and LangChain composition | Reader machine | Receives the Scrapfly result, then sends it to local Ollama | Parsed JSON |



Scrapfly handles cloud-based page collection, while Python, LangChain, and Ollama remain on the reader’s machine for local composition and inference.Scrapfly's collection step is not run offline. Treat the Markdown it returns as content that left your machine and came back not as content that was always local. The next section runs this exact split end to end from a public product page to a parsed JSON result.



Scrapfly

#### Scale your web scraping effortlessly

Scrapfly handles proxies, browsers, and anti-bot bypass — so you can focus on data.

[Try Free →](https://scrapfly.io/register)## How Do You Extract Web Data with Scrapfly, LangChain, and Ollama?

Create a project-local virtual environment before installing anything, so these exact package versions do not collide with other projects on the same machine:

bash```bash
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate
pip install langchain-ollama==1.1.0 scrapfly-sdk==0.12.0
```



The run documented here used Python 3.13.11 and Ollama 0.32.5, with `qwen3.5:9b` already pulled from the previous section. `langchain-ollama` is the package behind the [`ChatOllama` integration](https://docs.langchain.com/oss/python/integrations/chat/ollama.md) used below. Set your Scrapfly API key as an environment variable rather than hardcoding it in the script, output, or a screenshot:

bash```bash
export SCRAPFLY_API_KEY="your-scrapfly-key"
```



The example fetches <https://web-scraping.dev/product/1> which is a public demo product page through [Scrapfly's SDK directly](https://scrapfly.io/docs/scrape-api/getting-started) rather than through a LangChain document loader. Calling the SDK directly keeps the collection step and its response shape visible, instead of hiding it behind a loader abstraction



python```python
import os

from langchain_core.output_parsers import JsonOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_ollama import ChatOllama
from scrapfly import ScrapeConfig, ScrapflyClient

MODEL = "qwen3.5:9b"
TARGET_URL = "https://web-scraping.dev/product/1"

# 1. Collect the page as Markdown with Scrapfly (runs in Scrapfly's cloud)
scrapfly = ScrapflyClient(key=os.environ["SCRAPFLY_API_KEY"])
result = scrapfly.scrape(ScrapeConfig(url=TARGET_URL, format="markdown"))
markdown = result.content

# 2. Prompt the local model for two fields only
prompt = ChatPromptTemplate.from_messages([
    (
        "system",
        "Extract the product title and price from the provided markdown. "
        "Return JSON with exactly two string fields: title and price.",
    ),
    ("human", "{markdown}"),
])

# 3. Local inference through Ollama, nothing leaves this machine at this step
llm = ChatOllama(
    model=MODEL,
    temperature=0,
    reasoning=False,
    format="json",
    num_ctx=8192,
)

# 4. Compose the chain: prompt -> local model -> JSON parser
chain = prompt | llm | JsonOutputParser()
print(chain.invoke({"markdown": markdown}))
```



Running this script against the public demo page above produced

json```json
{"title": "Box of Chocolate Candy", "price": "$9.99"}
```



That output demonstrates one target page, one package set, one model at one context size and a two-field schema. It is not a benchmark and it does not guarantee the same title-and-price result for a different page, a different schema or a different context size. Treat it as evidence that the flow above works end to end not as a claim about extraction accuracy in general.

This example is Python because the LangChain integration demonstrated here is Python. Scrapfly itself is not Python-only. Official SDKs also cover TypeScript, Go, and Rust if your application runs on a different stack.

For a broader LangChain workflow beyond this single extraction, see

[LangChain Web Scraping: Build AI Agents &amp; RAG ApplicationsLearn to integrate LangChain with Scrapfly for web scraping. Build AI agents and RAG applications that extract, process, and understand web data at scale.](https://scrapfly.io/blog/posts/langchain-web-scraping-complete-guide-scrapfly)



## When Should You Use Ollama, Hosted Inference, or Scrapfly Extraction API?

Use Ollama when the workload fits your memory budget and you want inference on hardware you control. Use hosted inference when the task needs more capability or burst capacity than your machine can provide. Use Scrapfly's Extraction API to hand off both collection and extraction to a managed service instead of running a local model.

Collection and inference are separate decisions. You can collect through Scrapfly and still infer locally, exactly as the previous section demonstrated, or send the same content to a hosted model instead. Content already on disk skips collection entirely.

| Route | Best fit | Data boundary | Operational tradeoff |
|---|---|---|---|
| Ollama local inference | Stable tasks that fit local memory | Prompt and response stay on the local inference path | Reader operates model files, memory, and runtime |
| Hosted inference | Capability or burst demand beyond the local machine | Prompt and supplied content go to the hosted service | Less local operation, external service dependency |
| Scrapfly Extraction API | Managed structured extraction from collected content | Collection and extraction run in Scrapfly cloud | No local model operation, separate cloud boundary |

Scrapfly's [Extraction API](https://scrapfly.io/docs/extraction-api/automatic-ai) covers a set of prebuilt extraction models for common structured objects. Rather than reproducing its request shape and output here, this guide links to its documentation directly, since the API's behavior is best read from the current docs, not copied into a second, possibly stale, example.

Whichever route you choose, evaluate it against your own prompts and schemas rather than inferring quality from parameter count or a vendor's benchmark table.

For workflows that chain multiple tool calls or make decisions on top of local inference, rather than running a single extraction, see

[Guide to Understanding and Developing LLM AgentsExplore how LLM agents transform AI, from text generators into dynamic decision-makers with tools like LangChain for automation, analysis &amp; more!](https://scrapfly.io/blog/posts/practical-guide-to-llm-agents)



## FAQ

Do Ollama Local Models Need an API Key?No. Local Ollama models do not need an Ollama API key. Requests go straight to the local server at `http://localhost:11434` with no authentication required. The Scrapfly collection step in this guide is a separate cloud service and uses its own project key, set as the `SCRAPFLY_API_KEY` environment variable, because that request leaves your machine and Ollama's local server never sees it.







Is LangChain Required to Run Ollama Locally?No. Ollama can be called directly from its command-line interface or its local HTTP API without any additional framework. LangChain is useful in this guide for composing the prompt template, the model call, and the JSON output parser into one chain, which is a convenience for the Python example, not a requirement for running `qwen3.5:9b` through Ollama on its own.









## Which Local LLM Setup Should You Use?

Choosing a local LLM setup comes down to the same sequence this guide walked through. Pick a quantized model whose actual downloaded file fits comfortably inside the memory Ollama has available, start with the context your task actually needs rather than the model's maximum, inspect the loaded process with `ollama show` and `ollama ps`, and evaluate the result against your own prompts before trusting it.

`qwen3.5:9b` at Q4\_K\_M with an 8,192-token context is this article's demonstrated setup, not a universal recommendation for every machine or task.

When your task needs current public web content rather than text you already have, use [Scrapfly's Web Scraping API](https://scrapfly.io/products/web-scraping-api) to collect Markdown, and keep the inference step that follows in local Ollama. If your source text is already local, or cannot leave your environment at all, skip cloud collection and go straight to the model.



### Web Scraping API

Scrape any website with our powerful API. Anti-bot bypass, JavaScript rendering, and rotating proxies built-in.



[Try Web Scraping API](https://scrapfly.io/docs/scrape-api/getting-started)



 

   [  Add as a preferred source ](https://google.com/preferences/source?q=scrapfly.io) Table of Contents















 

  Table of Contents- [Key Takeaways](#key-takeaways)
- [What Is a Local LLM, and What Actually Stays Local?](#what-is-a-local-llm-and-what-actually-stays-local)
- [How Much Memory Does an Ollama Local LLM Need?](#how-much-memory-does-an-ollama-local-llm-need)
- [How Do Model Weights and Q4\_K\_M Quantization Affect Memory?](#how-do-model-weights-and-q4-k-m-quantization-affect-memory)
- [How Does Ollama Context Length Affect Memory?](#how-does-ollama-context-length-affect-memory)
- [What Does ollama ps Show About Context and Offload?](#what-does-ollama-ps-show-about-context-and-offload)
- [How Do You Install Ollama and Run qwen3.5:9b?](#how-do-you-install-ollama-and-run-qwen3-5-9b)
- [How Do Scrapfly and Ollama Split Cloud Collection and Local Inference?](#how-do-scrapfly-and-ollama-split-cloud-collection-and-local-inference)
- [How Do You Extract Web Data with Scrapfly, LangChain, and Ollama?](#how-do-you-extract-web-data-with-scrapfly-langchain-and-ollama)
- [When Should You Use Ollama, Hosted Inference, or Scrapfly Extraction API?](#when-should-you-use-ollama-hosted-inference-or-scrapfly-extraction-api)
- [FAQ](#faq)
- [Which Local LLM Setup Should You Use?](#which-local-llm-setup-should-you-use)
 
    Join the Newsletter  Get monthly web scraping insights 

 

  



Scale Your Web Scraping

Anti-bot bypass, browser rendering, and rotating proxies, all in one API. Start with 1,000 free credits.

  No credit card required  1,000 free API credits  Anti-bot bypass included 

 [Start Free](https://scrapfly.io/register) [View Docs](https://scrapfly.io/docs/onboarding) 

 Not ready? Get our newsletter instead. 

 

 ## Related Articles

 [     

 ai 

### Top LangChain Alternatives in 2026: Which One Should You Choose?

Explore the best LangChain alternatives in 2026 for building powerful AI applications. Compare features, performance, an...

 

 ](https://scrapfly.io/blog/posts/top-langchain-alternatives) [  

 ai 

### Guide to Understanding and Developing LLM Agents

Explore how LLM agents transform AI, from text generators into dynamic decision-makers with tools like LangChain for aut...

 

 ](https://scrapfly.io/blog/posts/practical-guide-to-llm-agents) [  

 ai 

### What Is MCP? Understanding the Model Context Protocol

What is MCP? Learn how the Model Context Protocol powers tools like Copilot Studio by giving AI models access to real-ti...

 

 ](https://scrapfly.io/blog/posts/what-is-mcp-understanding-the-model-context-protocol) 

  ## Related Questions

- [ Q How to select any element using wildcard in XPath? ](https://scrapfly.io/blog/answers/how-to-select-elements-of-any-name-using-wildcards-in-xpath)
- [ Q How to use cURL in Python? ](https://scrapfly.io/blog/answers/how-to-use-curl-in-python)
 
  



   



 Scale your web scraping effortlessly, **1,000 free credits** [Start Free](https://scrapfly.io/register)