Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -75,6 +75,10 @@ Here’s an example of how it works:

This method is perfect if you want to skip navigating through the Admin Settings menu and get right to using your models.

### Downloading GGUF Models from Hugging Face

Ollama can also download GGUF models straight from Hugging Face. In **Pull a model from Ollama.com** or the model selector, enter the model as `hf.co/{username}/{repository}:{quantization}`, for example `hf.co/ggml-org/gemma-3-270m-it-GGUF:Q8_0`. The quantization matches the one in the GGUF file's name, such as `Q4_K_M` or `Q8_0`. The model is then listed under that full name. See Hugging Face's [Ollama guide](https://huggingface.co/docs/hub/en/ollama) for which repositories work.

---

## Updating and Deleting Models
Expand All @@ -98,6 +102,27 @@ Deleting removes the model from the Ollama server itself and frees its disk spac

---

## Creating a Model

**Create a model**, also in the **Manage** dialog, makes a new model on that Ollama server from one it already has, with its own system prompt or parameters. Nothing new is downloaded, so it takes a few seconds.

1. Enter a name for the new model, for example `pirate-gemma`.
2. In the box below it, describe the model as JSON, using the fields of Ollama's [create API](https://docs.ollama.com/api/create). `from` is the model to build on, and must already be on that server. For example:

```json
{"from": "gemma3:270m", "system": "You answer in one short sentence.", "parameters": {"temperature": 0.3, "num_ctx": 8192}}
```

3. Click **Create Model**, the button next to the name. It stays disabled until both the name and the JSON are filled in.

When Ollama finishes, Open WebUI shows **success**, and the new model appears in the model selector as `pirate-gemma:latest`. If the JSON can't be read, the error is shown and nothing is sent. If the `from` model isn't on that server, Ollama tries to download it from Ollama.com instead and fails with **pull model manifest: file does not exist** when no such model exists there.

:::note A system prompt set here is a fallback
Ollama uses the model's own system prompt only when a chat sends none. A system prompt from the model's settings in **Workspace** > **Models**, from a user's settings, or added by a filter replaces it. To give a model a system prompt that always applies in Open WebUI, set it in **Workspace** > **Models** instead.
:::

---

## Unloading Loaded Models

Open WebUI shows a green "Loaded" indicator next to any Ollama model that is currently kept warm by the runtime, and admins see an **Eject** button on the model row to unload it without restarting the server. Behind the scenes, Open WebUI calls `POST /api/models/unload` (admin-only), which forwards a `keep_alive=0` generate call to every Ollama node serving that model.
Expand Down
Loading