Skip to content

Repository files navigation

RoboSystems Content Machine

License: MIT

Automated equity-research content pipeline. Turns a company's SEC filings into a narrated 16:9 video, a purpose-built 9:16 short, a written brief that publishes as-is, and an audio edition of that brief - one analysis, every surface.

  • Campaign-Driven - reusable campaign templates define the editorial angle, analytical framework, and output specs; apply them to any ticker.
  • Authored in Claude Code - one session reads the filings via RoboSystems MCP tools and writes the brief, the video script, the short, and the social copy. No hand-off to another app.
  • Rendered locally - the deck is HTML built from the script, shot frame-by-frame in headless Chrome, and muxed with ffmpeg. No cloud render service, no per-render cost.
  • No hand-authored slides - you write numbers into script.json; the renderer draws every slide.

🎙️ Voiceover & music run on ElevenLabs. Setting this up? Signing up through our referral link costs you nothing extra and directly supports the project. Affiliate link.

Quick Start

git clone https://github.com/RoboFinSystems/robosystems-content-machine.git
cd robosystems-content-machine

# Scaffold a project from a campaign
just campaign TICKER campaign_name

# Or scaffold from the base template (no campaign)
just new TICKER

The first just command auto-creates .env from .env.example. Fill in your API keys (see Setup).

How It Works

Three stages: scaffold, author, render. Authoring happens in a Claude Code session against a written contract; everything on either side of it is a just recipe.

1. Scaffold a Project

Every project starts from the base template/ (folder structure + the authoring contract + assets). Projects are company-centric - sources accumulate over time, each run produces a new set of outputs.

just new TICKER                      # base template
just campaign TICKER campaign_name   # with a campaign overlay
just campaigns                       # list available campaigns
just recover TICKER campaign_name    # re-cover for a new quarter (archives prior outputs)

Campaigns

Campaigns add an editorial layer for thematic coverage across many companies - the voice, analytical framework, target tickers, and shared reference data.

campaigns/
  my_campaign/
    CAMPAIGN_BRIEF.md           # Editorial strategy and analytical framework
    AUTHORING_INSTRUCTIONS.md   # Authoring instructions (overrides base)
    tickers.md                  # Target companies and production calendar
    sources/                    # Third-party research and reference data (gitignored)
    overrides/                  # File replacements (custom assets/instructions)

The base template is applied first, then the campaign overlays its instructions, brief, and shared sources on top.

2. Author (Claude Code)

Point a Claude Code session at the repo. The scaffolded project carries its own contract: AUTHORING_INSTRUCTIONS.md (the editorial brief) and PRODUCTION_CONTRACT.md (the schema, slide kinds, layout capacity limits, and spoken-form TTS rules). The session reads those plus everything in sources/, verifies every number against the SEC graph over MCP, and writes:

  • Narrative brief (reports/{TICKER}_brief.md) - the written analysis, authored first. Ships verbatim as a native X Article.
  • Video script (scripts/{TICKER}_script.json) - the source of truth: ordered segments carrying narration plus the exact numbers each slide draws.
  • Short script (scripts/{TICKER}_short_script.json) - 5-6 beats for the vertical cut, targeting ~45s.
  • Social copy (social/) - X post, YouTube description, and publish.json (the titles and links the publish step reads).

The repo ships the research-lane skills that drive this - /scout, /collect, /author, /review, /status, /batch, /refresh - as .claude/commands/*.md. They are conveniences around the contract, not a dependency: the contract is the spec, and a session that reads AUTHORING_INSTRUCTIONS.md + PRODUCTION_CONTRACT.md produces the same artifacts without them.

just validate TICKER    # gate the authored output against the contract before rendering

validate is schema-level. It catches missing fields, capacity overruns and duplicate refs; it cannot see that a chart is visually wrong, so check the rendered frames too.

3. Render

just webdeck-pipeline TICKER         # long-form 16:9 -> videos/{TICKER}_final.mp4
just webdeck-short-pipeline TICKER   # 9:16 short    -> videos/{TICKER}_short.mp4
Step Command What it does
Everything just webdeck-pipeline TICKER Runs the five steps below end to end
Validate just validate TICKER Checks the authored output against the production contract
Voiceover just voiceover TICKER Sends narration to ElevenLabs TTS (idempotent; --force to regen)
Build just webdeck TICKER Builds the animated HTML deck from script.json + VO durations
Render just webdeck-render TICKER Renders the deck to frames via headless Chrome (puppeteer-core)
Mux just webdeck-mux TICKER Muxes narration, and narration + ducked music, with ffmpeg

The 9:16 short runs the same engine at 1080x1920. It is purpose-built vertical, not a crop: its own beat kinds (hook / stat / cards / points / cta), burned-in kinetic captions, a progress bar and a $TICKER chip. One asset serves both X native video and YouTube Shorts.

Rendering is entirely local - puppeteer-core drives headless Chrome, ffmpeg does the mux. There is no cloud render service and no per-render cost, so a re-render costs only wall clock. The mux writes videos/{TICKER}_timestamps.txt with the authoritative YouTube chapter times.

Two pre-render inspection recipes save a full render when the layout is in question:

just webdeck-stills TICKER "3,17,42"        # single frames at given seconds
just webdeck-short-stills TICKER "1,12,30"

4. Thumbnails, publish, post

just thumbnails TICKER   # YouTube thumbnail, generated from the brief via OpenAI
just narrate TICKER      # the audio edition: an ElevenLabs read of the brief (auto-runs on publish)
just publish TICKER      # upload deliverables to the S3 artifact store + reindex the catalog
just postpack TICKER     # assemble the per-platform publish pack (paste-ready copy + S3 links)

Posting is per-surface, each asset in its best format:

just yt-upload TICKER   && just yt-publish TICKER         # long-form (uploads private, then flips)
just yt-short TICKER    && just yt-short-publish TICKER   # the Short (auto-links the long-form)
just x-article TICKER                                     # the brief as a native X Article (draft)
just x-article TICKER --publish
just x-short TICKER                                       # the 9:16 as X native video
just x-post TICKER                                        # the text post
just sync-youtube                                         # capture published URLs into the catalog
just analytics [tickers] · just insights                  # per-post rollup · channel-level reach

just yt-auth and just x-auth do the one-time OAuth for the YouTube and X APIs.

Brief-only coverage

Not every name needs a video. A ticker can ship as a written brief alone - published to /research and posted as an X Article - by skipping the script and render entirely:

just validate-brief TICKER   # validate a brief-only project (skips every script/render check)
just publish-brief TICKER    # validate + publish the brief straight to /research
just x-article TICKER        # then post it as a native X Article

Useful for covering more filings than you have render time for, and for names where the analysis is the whole product. A brief-only name still ships audio: publish-brief narrates it like any other, so the page has a written report and a spoken one even with no video.

The audio edition

Every report ships a "Listen to this report" narration - a single-voice ElevenLabs read of reports/{TICKER}_brief.md, written to reports/{TICKER}_narration.mp3 and played from a card on the /research page. It is the same affordance every blog post has, and it took the slot the Q&A podcast held before that format was retired (2026-07-21; its assets were deleted 2026-09-05).

just narrate TICKER             # generate it on its own
just narrate TICKER --force     # re-read it (re-bills TTS)
just narrate TICKER --dry-run   # print the cleaned text, call no API, bill nothing
just publish TICKER --no-audio  # publish without generating one

just publish narrates any report that has no audio yet, so the feature stays consistent across the catalog; a TTS failure logs a warning and still publishes the report. The brief is the script - nothing extra is authored. Tables are stripped (they read terribly aloud), cashtags lose the $ that would otherwise be read as "dollar S T X", and [PROMO_CODE] resolves exactly as it does for the published text. Narration shares one engine with the blog (tools/narrate_common.py), so both lanes use the same brand voice and the same chunk-and-concat path.

Publishing (S3 artifact store)

just publish {TICKER} uploads the final deliverables (long-form, short, thumbnail, brief, narration, social copy) to s3://$AWS_S3_BUCKET/content/{TICKER}/ and prints public URLs (served via $AWS_CDN_DOMAIN_URL when set, else https://$AWS_S3_BUCKET.s3.amazonaws.com/content/{TICKER}/…)

  • a durable artifact store, separate from posting to YouTube / X. The bucket policy grants public read on the content/* + blog/* prefixes only (no user data - the store is public by design); everything else in the bucket stays private. The bucket + CloudFront CDN are managed by cloudformation/content.yaml (just infra-deploy - see Infrastructure below).

Blog pipeline

A lighter sibling of the research pipeline for markdown essays. A post is one file - blog/<slug>/post.md (YAML frontmatter + body), authored and git-versioned in this repo. Narration, cover image, and social copy are all optional and additive; a post with just post.md publishes cleanly.

One catalog feeds two sites. The frontmatter site field (robosystems, the default, or roboledger) says which app renders the post: robosystems-app lists the graph and platform essays, roboledger-app lists the buyer posts. Moving a post is one frontmatter line plus a reindex; canonicalUrl must sit on the same domain (publish refuses a mismatch).

just blog-new <slug> [site] # scaffold blog/<slug>/post.md from the template (site: robosystems | roboledger)
just blog-publish <slug>    # auto-narrate (default-on) + upload blog/<slug>/* to S3 + reindex
just blog-narrate <slug>    # (re)generate narration on its own; --force to redo
just blog-social <slug>     # optional: paste-ready distribution pack (uses <slug>_x_post.txt if present)
just blog-x-article <slug>  # publish the essay as a native X Article
just blog-reindex           # rebuild blog/index.json (the catalog the app's /blog routes read)

Every post ships with a "Listen to this story" narration - blog-publish auto-narrates any post that has no audio yet (pass --no-audio to skip), so the feature stays consistent across the whole catalog. Narration reuses the same ElevenLabs path as the research voiceover (one brand voice; body stripped of code/tables, chunked for TTS, concatenated with ffmpeg). blog-publish also writes a self-describing meta.json and refreshes blog/index.json - a versioned contract (version: 1) with absolute CDN asset URLs, the same consumption shape the /research catalog uses. The app consumes it via SSG/ISR; publishing or editing a post no longer needs an app redeploy.

Shared Media Libraries

The short pulls from reusable, mood/tag-tagged libraries that compound across every ticker:

just broll          # show the b-roll library + coverage by category
just broll-sync     # register new clips dropped into assets/broll/
just music-sync     # register new tracks dropped into assets/music/
just music "<prompt>"   # generate a music bed via the ElevenLabs Music API

Clips and tracks are selected by theme - a broll_theme / music_mood (tags) or an explicit list. Manifests are tracked; the heavy .mp4/.mp3 binaries are gitignored (local-only).

Product demo capture

A second, separate capture track records the live product rather than a deck: a headless browser drives the real UI, a drawn cursor travels and clicks, and the page responds. It shares this repo's audio and mux stages but none of its deck code. See renderer/README.md for just render-setup, render-capture, demo-probe, and demo-pipeline.

Batch Operations

just projects              # List all projects
just play PROJECT          # Play the final video
just durations PROJECT     # Show media durations via ffprobe
just clean PROJECT         # Remove generated assets (keeps source files)

Setup

Required Tools

  • uv - Python package manager
  • just - command runner
  • ffmpeg / ffprobe - media processing (render mux, shorts, probing)
  • Node.js - the webdeck renderer (puppeteer-core) and renderer/
  • AWS CLI - S3 uploads for publishing + the CloudFront CDN

API Keys

Configure in .env after first run:

Service Keys Purpose
ElevenLabs ELEVEN_LABS_API_KEY, ELEVEN_LABS_VOICE_ID Voiceover + the Music API
OpenAI OPENAI_API_KEY Thumbnail generation (just thumbnails)
AWS AWS_PROFILE, AWS_REGION, AWS_S3_BUCKET, AWS_CDN_DOMAIN_URL, AWS_CLOUDFRONT_DISTRIBUTION_ID Asset uploads + CloudFront CDN
YouTube YT_CLIENT_ID, YT_CLIENT_SECRET, YT_REFRESH_TOKEN, YT_CHANNEL_ID Uploads + analytics (just yt-auth writes the refresh token into .env)
X X_CONSUMER_KEY, X_SECRET_KEY, X_ACCESS_TOKEN, X_ACCESS_SECRET, X_HANDLE Articles, posts, native video (just x-auth verifies)

The ElevenLabs link above is a referral link.

Filing data (RoboSystems MCP)

Authoring verifies numbers against SEC XBRL filings through the RoboSystems MCP server. Configure it in your Claude Code session; the sec graph is the read-only shared repository the research lane queries. SEC_RAW_BUCKET (optional) points /collect at a store of raw filing archives; without it, fetch filings from EDGAR by hand.

Infrastructure

The content bucket + CloudFront CDN are defined in cloudformation/content.yaml and deployed locally via the AWS CLI (no GitHub Actions), mirroring the platform repo's just bootstrap flow. Config comes from .env (AWS_PROFILE, AWS_S3_BUCKET, optional AWS_CDN_DOMAIN_URL; AWS_ROUTE53_HOSTED_ZONE_ID is auto-resolved from the CDN domain).

just infra-validate    # validate the template
just infra-deploy      # create the bucket + CDN stack (+ wait, + print outputs)
just reindex           # rebuild content/index.json (CDN urls)
just infra-outputs     # show bucket / CDN url / distribution id

Resources

Support

Acknowledgements

Backed by an ElevenLabs Grant - the credits power the voiceover and music generation behind every video this pipeline produces.

ElevenLabs Grants

Using ElevenLabs yourself? Our referral link costs you nothing extra and supports the project.

License

This project is licensed under the MIT License - see the LICENSE file for details.

MIT © 2026 RFS LLC

About

Automated video content pipeline for financial analysis. Combines AI-generated content with production automation to turn SEC filings into narrated videos, podcasts, and social posts. Uses Claude Cowork + RoboSystems MCP for research, Claude Design, ElevenLabs, and Shotstack for production.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Used by

Contributors

Languages