TickerFlow is a Python backend that turns local market CSVs into validated, queryable API data.
Problem: make local OHLCV data quality decisions visible before data is queryable.
- Workflow: CSV → canonical UTC frame → validation report → partitioned Parquet → half-open queries → hourly/daily time bars → FastAPI demo.
- Hard decision: quarantine invalid rows without silent repair. Quality policy and the invalid-OHLC regression.
- Validation: dirty fixture report tests and ingestion/storage/query regressions.
- Quick run:
uv sync --extra dev && uv run pytest; then follow the complete seed → query demonstration. - Limits: local filesystem and one writer; identical-key repeat ingestion is idempotent, but does not establish concurrent-write safety. Writes directly replace partitions; no atomic crash recovery or multi-partition transactions are promised. Trades/quotes and tick/volume/dollar bars remain planned.
Market-data pipelines need deterministic ingestion, explicit schemas, visible quality decisions, durable storage, and stable query boundaries. TickerFlow provides that local workflow for financial time series, from CSV normalization through Parquet storage and feature-ready time bars.
- Ingest local OHLCV CSV files with explicit schemas and configuration.
- Normalize timestamps to UTC and report data-quality issues.
- Store valid OHLCV rows as partitioned Parquet.
- Query datasets by symbol and half-open date range through Python and FastAPI.
- Build hourly and daily time bars with explicit interval boundaries.
- Explore the local workflow through the
/demopage.
- Trade and quote ingestion.
- Tick, volume, and dollar bars.
- Reproducible performance benchmarks on larger datasets.
- Python 3.12+
- Polars for DataFrame transformations.
- DuckDB for local analytical queries.
- PyArrow/Parquet for storage.
- Pydantic and FastAPI for backend contracts.
- pytest, ruff, and mypy for quality.
- Schemas before pipelines.
- Deterministic local fixtures before live data.
- Validation reports before silent cleaning.
- Small vertical slices over broad unsupported features.
- Bootstrap Python package, CI, linting, typing, and tests.
- Define canonical schemas for OHLCV and trades.
- Implement local CSV ingestion with validation reports.
- Store normalized data as partitioned Parquet.
- Add query service and FastAPI endpoint.
- Implement time bars (done); volume bars remain planned.
- Add benchmarks and quality-report examples.
TickerFlow currently implements the first OHLCV backend slice:
- Load local OHLCV CSV files with typed ingestion configuration.
- Normalize timestamps to timezone-aware UTC at microsecond precision.
- Validate dirty rows and return a structured quality report instead of silently dropping data.
- Store valid rows as partitioned local Parquet under
ohlcv/symbol=<SYMBOL>/date=<YYYY-MM-DD>/data.parquet. - Query OHLCV rows by symbol and half-open UTC date range
[start, end). - Discover available local datasets and symbols from Parquet partitions.
- Build hourly or daily time bars with explicit half-open boundaries.
- Expose
/health,/datasets,/symbols,/ohlcv, and/bars/timethrough FastAPI. - Provide a browser market-data demo at
/demo.
Input CSV fixtures use these columns:
timestamp,symbol,open,high,low,close,volume
Canonical rows use:
timestamp_utc: datetime[us, UTC]
symbol: uppercase string
open: float
high: float
low: float
close: float
volume: float
source: string
Prices are unadjusted fixture values in arbitrary currency units. OHLC prices must be finite, positive, and consistent with low/high bounds. Volume is a finite non-negative numeric quantity. Corporate actions and live data sources are intentionally out of scope.
uv sync --extra dev
uv run ruff format .
uv run ruff check .
uv run mypy src tests
uv run pytestRun the API locally:
uv run uvicorn tickerflow.api.main:app --reloadOpen the market-data demo UI:
open http://127.0.0.1:8000/demoExample query after writing Parquet data into the configured data directory:
curl "http://127.0.0.1:8000/ohlcv?symbol=AAPL&start=2024-01-02T00:00:00Z&end=2024-01-04T00:00:00Z"Catalog endpoints:
curl "http://127.0.0.1:8000/datasets"
curl "http://127.0.0.1:8000/symbols?dataset=ohlcv"Time-bar endpoint:
curl "http://127.0.0.1:8000/bars/time?symbol=AAPL&start=2024-01-02T14:00:00Z&end=2024-01-02T16:00:00Z&interval=1h"By default, the API reads from .tickerflow. Set TICKERFLOW_DATA_DIR to point at another local Parquet root.
The deterministic demo script lives in docs/DEMO.md.
The prepared homelab manifest uses sudo homelab project deploy tickerflow,
running the repository quality gate before activating an immutable release.
Uvicorn is a runtime dependency, retained by production-only
uv sync --frozen --no-extra dev --no-editable. The API binds loopback behind
Caddy; /health checks readiness and the existing root redirects to /demo.
Set TICKERFLOW_DATA_DIR to writable state outside the release (the VPS
manifest uses /var/lib/homelab/tickerflow/data). Demo input remains synthetic.
Public DNS and privileged host setup require separate publication; this README
does not claim that the prepared VPS demo is already live.
The main /demo uses bounded, committed cases through GET /demo/cases and
POST /demo/run. Each run keeps original string cells, normalized values and
actual row-level validation decisions, writes accepted rows to a request-local
Parquet store, queries with DuckDB and traces query rows into hourly/daily bars.
Temporary files are cleaned after the request; the shared persistent dataset is
not altered. The legacy /demo/seed and data/query APIs remain compatible.
Accepted + quarantined counts reconcile to the same input. Issue counts can overlap. Duplicate pairs are both quarantined. Invalid original values remain inspectable even when normalization produces null/non-finite values; response JSON uses safe nulls for non-finite normalized cells. Cases are synthetic, unadjusted OHLCV, UTC microseconds, currency-unit prices and synthetic volume.
Bars aggregate only the selected query's observations, use half-open UTC
intervals and do not fill gaps. Clipped edge intervals are marked. Exports contain
actual report/query/bar data and fixture hashes. Repeating a run demonstrates
canonical result reproducibility, not byte-identical Parquet files. No upload or
arbitrary SQL endpoint is exposed. The production wheel includes fixtures/assets.
Set DEMO_BUILD_REVISION to the actual build SHA or accept the honest unknown.
Run browser checks with node tests/browser/demo.cjs, Playwright available and
TICKERFLOW_DEMO_URL pointing at a local server. Hand-checked case expectations
live in tests/fixtures/demo_expected.json.
