Skip to content

docs: show how much memory a larger Ollama context needs - #1442

Open
159753a52 wants to merge 1 commit into
open-webui:mainfrom
159753a52:docs/ollama-context-memory
Open

159753a52 wants to merge 1 commit into
open-webui:mainfrom
159753a52:docs/ollama-context-memory

Conversation

@159753a52

Copy link
Copy Markdown

The context length section ends with "A larger context uses more VRAM and RAM, so size it to what your hardware can hold", but doesn't say how much. This adds a short subsection so people can size num_ctx / OLLAMA_CONTEXT_LENGTH before hitting a CPU offload:

  • Ollama reserves the KV cache for the whole context at load time.
  • The f16 per-token formula, with one worked row for Llama 3.1 8B (32 layers, 8 KV heads, head dim 128 from its config.json: 128 KiB/token, 4 GiB at 32768, 16 GiB at 131072). Weight size (4.92 GB) is from the llama3.1:8b manifest on registry.ollama.ai.
  • OLLAMA_NUM_PARALLEL multiplies it, OLLAMA_KV_CACHE_TYPE shrinks it (linking Ollama's FAQ rather than repeating it), sliding-window / MLA models need less.
  • How to confirm with ollama ps (CONTEXT and PROCESSOR columns).

Disclosure: the last sentence links ModelVRAM, a free VRAM calculator that I built (no sign-up, runs in the browser, MIT core at https://github.com/159753a52/llm-vram-calculator). If you'd rather not link third-party tools, I'm happy to drop that sentence; the rest stands on its own.

Checked with npx docusaurus-mdx-checker -c docs/getting-started/quick-start/connect-a-provider (all 9 files compile).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant