This repository is deprecated. All of its content and history has been moved to googleapis/google-cloud-node.
-
Updated
Jul 13, 2023
This repository is deprecated. All of its content and history has been moved to googleapis/google-cloud-node.
Use the Moondream 2 model to detect faces and their gaze directions in videos.
A powerful video summarization tool that utilizes Moondream alongside multiple AI models to provide comprehensive video understanding through audio transcription, intelligent frame selection, visual description, and content summarization.
Context based video seek and search
Uses Video Intelligence to analyse and edit a video based on a given sentence.
Foundational framework for mission-critical surveillance, autonomous video intelligence, and situational awareness.
Media Info Comparison
This tool uses Moondream 2B, a powerful yet lightweight vision-language model, to detect and redact objects from videos. Moondream can recognize a wide variety of objects, people, text, and more with high accuracy while being much smaller than most vision models.
Local-first video intelligence orchestrator using Intel OpenVINO, Tauri, gRPC, SQLite, and agentic AI routing.
Real-time AI agent for querying live courtroom video with sub-500ms latency. Multimodal search combining video intelligence, speech-to-text, and hybrid search. Built with Stream, Twelve Labs, Deepgram, and Gemini Live API.
Learn public speaking from talks you admire — ask how they do it, watch captioned clips cut from the video. Built with VideoDB.
Universal video intelligence: turn any video + a prompt into a timestamped timeline, evidence-grounded findings, and structured reports — ready for AI agents.
Multimodal video dossiers for agents: transcripts, frames, OCR, evidence, and RAG-ready knowledge.
VideoMind AI - AI Video Intelligence OS: collection → ASR → AI analysis → reports. Cross-platform desktop app.
UnReel is an AI-powered Video Intelligence engine built to decode the context of any short-form content. It is designed for users who encounter language barriers, missed situational context, or struggle to find resources mentioned in a video via a dedicated video analysis pipeline.
BRI — empathetic video intelligence with production Streamlit, FastAPI MCP, SQLite durability, and multimodal ML tooling
The missing middle between raw video and reasoning models. Turn video into structured intelligence. Citable, queryable, AI-ready. CLI + MCP server. No Docker. No GPU.
Multimodal video-to-music recommender. Google Video Intelligence labels + real Spotify audio features (with librosa fallback) + Keras emotion classifier.
This Multimodal AI Agent is a Streamlit application that uses Gemini 2.0 Flash to analyze video content alongside real-time web research. It enables users to upload videos and receive comprehensive, data-driven answers by synthesizing visual insights with live information from the internet.
Video intelligence pipeline — download, transcribe (Whisper), detect scenes, capture keyframes, analyze frames. (Companion repo — architecture and interface only. Source lives in the private engine.)
Add a description, image, and links to the video-intelligence topic page so that developers can more easily learn about it.
To associate your repository with the video-intelligence topic, visit your repo's landing page and select "manage topics."