diff --git a/docs/getting-started/quick-start/connect-a-provider/starting-with-ollama.mdx b/docs/getting-started/quick-start/connect-a-provider/starting-with-ollama.mdx index 3087da5320..2a8e7ad10a 100644 --- a/docs/getting-started/quick-start/connect-a-provider/starting-with-ollama.mdx +++ b/docs/getting-started/quick-start/connect-a-provider/starting-with-ollama.mdx @@ -75,6 +75,10 @@ Here’s an example of how it works: This method is perfect if you want to skip navigating through the Admin Settings menu and get right to using your models. +### Downloading GGUF Models from Hugging Face + +Ollama can also download GGUF models straight from Hugging Face. In **Pull a model from Ollama.com** or the model selector, enter the model as `hf.co/{username}/{repository}:{quantization}`, for example `hf.co/ggml-org/gemma-3-270m-it-GGUF:Q8_0`. The quantization matches the one in the GGUF file's name, such as `Q4_K_M` or `Q8_0`. The model is then listed under that full name. See Hugging Face's [Ollama guide](https://huggingface.co/docs/hub/en/ollama) for which repositories work. + --- ## Updating and Deleting Models @@ -98,6 +102,27 @@ Deleting removes the model from the Ollama server itself and frees its disk spac --- +## Creating a Model + +**Create a model**, also in the **Manage** dialog, makes a new model on that Ollama server from one it already has, with its own system prompt or parameters. Nothing new is downloaded, so it takes a few seconds. + +1. Enter a name for the new model, for example `pirate-gemma`. +2. In the box below it, describe the model as JSON, using the fields of Ollama's [create API](https://docs.ollama.com/api/create). `from` is the model to build on, and must already be on that server. For example: + + ```json + {"from": "gemma3:270m", "system": "You answer in one short sentence.", "parameters": {"temperature": 0.3, "num_ctx": 8192}} + ``` + +3. Click **Create Model**, the button next to the name. It stays disabled until both the name and the JSON are filled in. + +When Ollama finishes, Open WebUI shows **success**, and the new model appears in the model selector as `pirate-gemma:latest`. If the JSON can't be read, the error is shown and nothing is sent. If the `from` model isn't on that server, Ollama tries to download it from Ollama.com instead and fails with **pull model manifest: file does not exist** when no such model exists there. + +:::note A system prompt set here is a fallback +Ollama uses the model's own system prompt only when a chat sends none. A system prompt from the model's settings in **Workspace** > **Models**, from a user's settings, or added by a filter replaces it. To give a model a system prompt that always applies in Open WebUI, set it in **Workspace** > **Models** instead. +::: + +--- + ## Unloading Loaded Models Open WebUI shows a green "Loaded" indicator next to any Ollama model that is currently kept warm by the runtime, and admins see an **Eject** button on the model row to unload it without restarting the server. Behind the scenes, Open WebUI calls `POST /api/models/unload` (admin-only), which forwards a `keep_alive=0` generate call to every Ollama node serving that model.