Pixelle Video
Backed by Alibaba's AIDC team and carrying over 27,700 GitHub stars, Pixelle-Video turns a single text prompt into a publish-ready short video in approximately three minutes — handling scriptwriting, image generation, voice narration, music selection, subtitle overlay, and final MP4 export in one automated pipeline. The engine supports multiple LLM backends for script generation including GPT-4, Qwen, DeepSeek, and local Ollama deployments, while image and video creation routes through either self-hosted ComfyUI workflows, cloud-based RunningHub pipelines, or direct API connections to DashScope Wan, OpenAI, Seedream, Seedance, and Kling AI. Text-to-speech synthesis uses Edge-TTS, Index-TTS, and other mainstream engines with multi-language voice profiles. Five distinct pipelines cover Quick Create, Standard, Digital Human Avatar broadcasting, Image-to-Video transformation, and Motion Transfer from reference video. The Streamlit web UI on port 8501 provides a visual workflow builder with template selection across portrait (1080x1920), landscape (1920x1080), and square formats, while the FastAPI server on port 8000 exposes a REST API with endpoints for async video generation, task polling, content scripting, TTS and image generation, template listing, and health checks. History persistence tracks all completed generations. HTML-based visual templates support static, image-overlay, and AI-video styles with customizable prompt prefixes. The modular architecture lets operators swap any atomic capability — image model, video model, TTS engine, or VLM — by editing a workflow JSON file without touching Python code. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. Apache 2.0 licensed.
OpenMontage
Reaching #1 on GitHub Trending with over 48,000 stars, OpenMontage is the first open-source agentic video production system — transforming AI coding assistants like Claude Code, Cursor, Copilot, Windsurf, and Codex into complete video studios that handle research, scripting, scene planning, asset generation, editing, and final rendering through natural language prompts. Twelve production pipelines cover animated explainers, cinematic trailers, documentary montages, talking heads, screen demos, podcast repurposing, character animation, localization and dubbing, avatar spokesperson videos, hybrid productions, clip factory batch processing, and animation workflows. Over 100 registered Python tools connect to 60+ providers including Kling, Runway Gen-4, Google Veo 3.1, FLUX, Google Imagen 4, ElevenLabs, and Suno AI for cloud generation, plus Piper TTS, WAN 2.1, Hunyuan, and CogVideo for fully local GPU rendering — while free footage from Archive.org, NASA, Wikimedia Commons, Pexels, and Unsplash powers the documentary montage pipeline's CLIP-indexed retrieval system for real-motion video without paid generation APIs. Two composition engines — Remotion for React-based programmatic video and HyperFrames for HTML/GSAP motion graphics — render final output with spring animations, word-level captions, kinetic typography, and SVG character rigs. A seven-dimension scored provider selector, pre-compose validation gates, post-render ffprobe self-review, slideshow risk scoring, configurable budget caps with per-action approval thresholds, and the Backlot live web dashboard for visual production monitoring enforce production-grade quality at every stage. Running on a dedicated VPS on RepoCloud with guaranteed CPU, RAM, and SSD, full root SSH access, and a browser serial console. AGPL-3.0 licensed.