A self-hosted personal knowledge base with natural-language Q&A — save notes, articles, and links, organize them with tags and collections, and (in later phases) query your own knowledge using local and cloud AI.
- Save notes, articles, and links — full CRUD with pagination and sorting
- Organize with tags (many-to-many) and filter notes by tag (AND-intersection)
- Group notes into collections
- Keyword full-text search over titles and content via MySQL FULLTEXT
- Multi-user with JWT auth (access + refresh tokens) and full per-user data isolation
- Auto-generate summaries from a local LLM (Ollama), persisted on the note (Phase 5)
- Suggest tags for a note via local LLM — suggest-only, never auto-attached (Phase 5)
- Ask questions in natural language, grounded in your own notes with source citations (RAG via Google Gemini by default, with local Ollama or Anthropic Claude as alternatives) (Phase 6, in progress)
Status: Phases 1–5 are complete — repo foundation, async MySQL + Notes CRUD API, JWT auth with per-user data isolation, tags/collections/full-text search, and local AI via Ollama. Summarization and tag suggestion are both verified end-to-end against a real local model. Phase 6 (RAG pipeline) is in progress: embedding pipeline, cosine similarity search over note chunks, and grounded natural-language Q&A that cites its sources — or declines when no note is relevant. See Project Status.
| Layer | Technology |
|---|---|
| Runtime | Python 3.12 |
| API framework | FastAPI 0.115 + Uvicorn |
| Database | MySQL 8.4 (InnoDB, utf8mb4, FULLTEXT) |
| ORM / migrations | SQLAlchemy 2.x async (asyncmy) + Alembic |
| Auth | JWT (PyJWT) + Argon2 password hashing (pwdlib) |
| Local AI | Ollama (llama3.2:3b for summarization, nomic-embed-text for embeddings) |
| Cloud AI | Google Gemini (RAG Q&A, default) — Ollama (zero-egress) or Anthropic Claude also selectable via RAG_PROVIDER |
| Infrastructure | Docker + Docker Compose (dev/live + prod) |
| Package manager | uv |
| Linting / typing | ruff + mypy |
| CI/CD | GitHub Actions |
Prerequisites: Docker Desktop running on Windows (or Docker + Docker Compose on Linux/macOS).
git clone https://github.com/Zauwx/second-brain.git
cd second-brain
cp .env.example .envEdit .env and set real values. The ones that matter for the current stack (api + MySQL + auth):
MYSQL_ROOT_PASSWORD,MYSQL_USER,MYSQL_PASSWORD,MYSQL_DATABASE— database credentialsDATABASE_URL— must match the MySQL user/password/database above (host ismysql, the compose service name)JWT_SECRET_KEY— generate one withpython -c "import secrets; print(secrets.token_hex(32))"
Ollama variables are needed from Phase 5 onward; the RAG tuning variables and whichever provider's key RAG_PROVIDER selects (GOOGLE_API_KEY for the default gemini, or ANTHROPIC_API_KEY for RAG_PROVIDER=anthropic) from Phase 6. A missing key for the selected provider does not block startup — only the cloud RAG endpoint degrades.
docker compose up -d --buildThis starts three containers on an internal network: api (FastAPI, exposed on port 8000), mysql (MySQL 8.4, internal only), and ollama (local LLM runtime, internal only). The API waits for MySQL to pass its healthcheck.
docker compose exec api alembic upgrade headRequired on a fresh volume — the schema is not created automatically. Skipping this leaves every write returning 500 with Table 'secondbrain.users' doesn't exist.
docker compose exec ollama ollama pull llama3.2:3b~2 GB, needed once per ollama_data volume. Only the AI endpoints depend on it; notes, tags, search, and auth work without it.
GET /health tells you whether this step is done. Its ollama field reports ok (server reachable and the configured chat model pulled), model_missing (server up, model not pulled — AI endpoints will fail), or unreachable (server down). The response stays 200 with status: "ok" in every case: local AI is optional and the rest of the app is unaffected.
Open http://localhost:8000/docs for the Swagger UI. From there you can:
POST /auth/registerthenPOST /auth/loginto obtain a JWT access token- Authorize in Swagger with the token
- Create notes (
POST /notes/), attach tags, and filter withGET /notes?tag=python&tag=docker - Generate a summary with
POST /ai/summarize(persisted on the note) or get tag ideas withPOST /ai/suggest-tags(suggest-only — never auto-attached)
A 200 OK from GET /health confirms the API container is up. If Ollama is stopped, the AI endpoints return a clean 503 while all note operations keep working.
To wipe all data and start fresh: docker compose down -v.
The application follows a domain-per-folder layout inside app/:
app/
main.py # FastAPI app factory, lifespan hooks, router registration
core/ # Settings (pydantic-settings), JWT/security helpers, shared dependencies
database.py # Async SQLAlchemy engine + session factory
auth/ # JWT register/login/refresh/logout, User model, per-user isolation
notes/ # Note CRUD — router, service, repository, schemas, ORM model
tags/ # Tag many-to-many — same layered structure
collections/ # Collection grouping — same layered structure
search/ # Full-text search; semantic search (Phase 6)
ai/ # Ollama local-LLM provider seam, summarize + suggest-tags (Phase 5);
# switchable cloud RAG client — Gemini default, Ollama or Anthropic (Phase 6)
Each domain uses the same layered structure: router (HTTP) → service (business logic) → repository (data access) → model/schemas (ORM + Pydantic).
Services run as Docker containers connected on an internal network:
- api — FastAPI application (Uvicorn in dev, Gunicorn in prod), exposed on port 8000
- mysql — MySQL 8.4, internal network only (not exposed to host)
- ollama — Local LLM runtime, internal network only (added in Phase 5)
Cloud LLM calls (Google Gemini by default, or Anthropic when RAG_PROVIDER=anthropic) go out over HTTPS directly from the API container. The API key is injected at runtime via .env — it never enters the Docker image. RAG_PROVIDER=ollama keeps this call on the internal network instead, with zero cloud egress.
This project deliberately mixes local and cloud AI ("IA hybride" — see Constraints), and it's worth being explicit about where that line sits, since this app stores personal notes.
- Chunking, embedding, and retrieval are 100% local. Splitting a note into chunks, generating its embedding vector, and finding similar chunks for a question all run through Ollama, inside the internal Docker network. None of that leaves the machine.
- Only the final answer-generation step calls a cloud LLM, and by default that's Google Gemini. When
POST /ai/askfinds relevant chunks, it sends the text of those retrieved chunks (not your whole note corpus, not other users' notes) to whichever providerRAG_PROVIDERselects, to generate a grounded, cited answer.RAG_PROVIDER=geminiis the default;RAG_PROVIDER=anthropicis retained as a paid alternative;RAG_PROVIDER=ollamakeeps this step local too (see below). - On Gemini's free tier, that retrieved note text is used to improve Google's products. Google's own pricing page states free-tier prompt and response content is "used to improve our products"; paid tiers guarantee it is not. Under the default configuration, this means the text of every note chunk sent to answer a question becomes Google training data — not a hypothetical, the actual behavior of the free key most people will use first. This is a conscious tradeoff for a zero-cost headline feature, not a compliance claim. If it's unacceptable for your notes, set
RAG_PROVIDER=ollamaand answer generation never leaves the machine, orRAG_PROVIDER=anthropicto use Anthropic's paid API instead. GET /notes/{id}/relatedmakes no cloud call at all. It's pure cosine similarity over already-computed local embeddings — no LLM, no network egress.- The rest of the app never talks to the cloud. Notes CRUD, tags, collections, full-text search, auth, and the local-LLM features (
/ai/summarize,/ai/suggest-tags) run entirely on MySQL + Ollama.
If you plan to store genuinely sensitive content in your own instance, know exactly where the boundary sits — RagService's answer-generation call in app/ai/rag/service.py — and that leaving the selected provider's API key unset simply disables /ai/ask (a clean 503) while everything else, including local summarization and tagging, keeps working.
| Phase | Scope | Status |
|---|---|---|
| 1 — Repo Foundation | Git hygiene, project scaffold, Docker base, FastAPI skeleton, GET /health |
✅ Complete (2026-06-24) |
| 2 — Database + API Skeleton | Async MySQL, Alembic migrations, Note CRUD, OpenAPI docs, pagination, tests | ✅ Complete (2026-06-24) |
| 3 — Auth + Per-User Data Isolation | JWT auth with refresh tokens, per-user query isolation, cross-user access tests | ✅ Complete (2026-06-25) |
| 4 — Tags, Collections, Full-Text Search | Many-to-many tags, collections, MySQL FULLTEXT search, REST surface polish | ✅ Complete (2026-06-29) |
| 5 — Local AI (Ollama) | Ollama in Docker, LLM provider abstraction, auto-summarization, tag suggestion | ✅ Complete (2026-07-20) |
| 6 — RAG Pipeline | Embedding pipeline, note chunks, cosine similarity, natural-language Q&A via Claude | 🚧 In progress |
| 7 — CI/CD Hardening + Portfolio Readiness | GitHub Actions lint/test/build/release, prod Compose, versioned images, secrets audit | ⏳ Planned |
# Install dependencies (requires uv)
uv sync
# Run linter
uv run ruff check app/
# Run type checker
uv run mypy app/
# Run tests
uv run pytest
# Start the stack with rebuild (dev)
docker compose up -d --buildContributions are welcome! Please read the Contributing Guide for development setup, coding conventions, and the pull request process. By participating you agree to the Code of Conduct.
Found a security issue? Please report it privately — see the Security Policy.
MIT — see LICENSE.
