yt-dlp(yt-dlp)
A feature-rich command-line audio/video downloader with support for thousands of video platforms.
Hand-picked open-source AI video tools, video editors, audio stem splitters, and media utilities on GitHub.
1-Click install commands for yt-dlp, Whisper AI, UVR5, ComfyUI & more.
A feature-rich command-line audio/video downloader with support for thousands of video platforms.
Stable Diffusion browser interface for AI image generation, inpainting, face swapping, and controlnet.
The most powerful visual node-based GUI for Stable Diffusion & AI Video generation.
Image generation software automating prompt weighting & quality tuning like Midjourney.
Leading creative engine for Visual Artists, Animators, and Studios powering AI workflows.
Open-source initiative dedicated to efficiently building OpenAI's Sora video generation pipeline.
State-of-the-art 12 billion parameter image generation model created by original Stable Diffusion founders.
A pipeline-level solution for real-time interactive image & video generation running at 100+ fps.
Tool to remove images background automatically using ONNX U2-Net models.
Free and open-source inpainting tool powered by SOTA AI models for erasing objects and text from images/video frames.
Robust Speech Recognition & Automatic Subtitle Generation Model trained on 680,000 hours of audio.
Bark is a transformer-based text-to-audio model created by Suno. Generates speech, laughter, and sound effects.
Deep learning toolkit for Text-to-Speech synthesis with 3-second voice cloning in 17 languages.
Easy-to-use Voice Conversion framework based on VITS for AI Voice cloning and song cover creation.
Code for the Music Source Separation model Demucs from Meta AI. Separates Drums, Bass, Vocals and Instruments.
Audacity is an easy-to-use, multi-track audio editor and recorder for Windows, macOS, Linux.
State-of-the-art GUI application for music source separation powered by MDX-Net & Demucs AI models.
A fast, local neural text-to-speech system that sounds great and runs on Raspberry Pi & low-end PCs.
Shotcut is a free, open source, cross-platform video editor supporting 4K multi-track editing.
Kdenlive is a non-linear video editor built on MLT Framework with support for multi-track editing.
The Swiss army knife of lossless video/audio editing. Cut & trim videos instantly without re-encoding.
A TypeScript library and editor for creating programmatic motion graphics videos.
Command line application for automatically analyzing and cutting silent sections of video and audio.
Open source API and interchange format created by Pixar for reading & writing video edit timelines between Premiere, DaVinci & Final Cut.
Official Mirror of Blender 3D creation suite - supports 3D pipeline, modeling, rigging, animation, and rendering.
Complete color management solution focused on motion picture production with emphasis on visual effects.
Open-source node-based compositing software for visual effects (VFX) and motion graphics.
3D Reconstruction Software based on AliceVision framework. Turn photos of real objects into textured 3D meshes.
Real-time radiance field rendering for 3D scene reconstruction from video/photos with 1080p 60fps rendering.
Free and open source software for video recording and live streaming.
Complete solution to record, convert and stream audio and video.
Free and Open Source AI Image Upscaler for Linux, MacOS, and Windows.
Command-line program to download image galleries and albums from hundreds of image hosting sites.
Powerful video encoding program for Windows supporting x264, x265, AV1, NVENC and VapourSynth filters.
LivePortrait animates still portrait images using motion drivers from a video source via PyTorch.
Python library for video editing that enables cutting, concatenation, title insertion, video compositing, and custom video creation through scripts.
An open-source few-shot voice cloning and text-to-speech tool that requires as little as one minute of audio data to train a workable voice model.
It implements an end-to-end deep learning pipeline to synthesize a target voice using just a 5-second audio sample.
An open-source AI voice studio that allows users to clone voices, dictate text, and generate speech locally.
Instant voice cloning model by MIT and MyShell that replicates voice tone and style from a short audio clip.
An industrial-level controllable and efficient zero-shot text-to-speech system for high-fidelity voice generation and cloning.
It performs real-time face swapping and one-click video deepfakes using only a single source image.
DeepFaceLive enables real-time face swapping on PC video feeds for live streaming and video calls.
An open-source Python tool that accurately syncs mouth movements in a video to any given audio speech file.
Olive is a free, open-source non-linear video editor built in C++ aimed at providing a serious alternative to professional software.
OpenShot is an open-source, cross-platform non-linear video editing software built with Python and C++.
Palmier Pro is an open-source macOS video editor built natively in Swift with integrated AI capabilities.
CogVideoX is an open-source text-to-video and image-to-video generation model by THUDM that produces short video clips from textual prompts.
Duix Avatar is a truly open-source AI toolkit that generates talking head digital human videos entirely offline using Python.
An open-source screen recorder built with web technologies that lets you capture your desktop into GIF, MP4, WebM, or APNG formats.
Screenity is a free, open-source Chrome extension that lets you record your browser tab, desktop, or camera with built-in annotation, drawing, and audio tools.
An open-source markerless motion capture system that uses standard camera footage to track 3D human movement and export skeletal data.
TripoSR is an open-source AI model that generates a textured 3D mesh object from a single 2D image in less than a second.
Gyroflow is an open-source video stabilization software that uses motion data from camera gyroscopes to smooth out shaky footage without cropping heavily.
It removes image backgrounds directly inside the web browser using client-side WebAssembly and machine learning without sending data to external servers.
SmartSub is a desktop application that automates video transcription, subtitle translation, AI dubbing, and voice cloning in one interface.
It allows autonomous coding agents to execute and automate complex video editing workflows through code and scripts.
An open-source local suite for voice cloning, video dubbing, dictation, and audiobook generation.
Miso TTS is an 8 billion parameter text-to-speech model capable of generating highly emotive spoken audio.
A modern desktop video editor built with Tauri, React, and TypeScript that aims to replicate popular premium features for free.
An open-source AI voice cloning, dubbing, dictation, and text-to-speech studio built for local deployment.
A 3B-active-parameter native unified multimodal model developed by ByteDance for image and video understanding, generation, and editing.
Bernini is an open-source unified framework combining a Multimodal Large Language Model semantic planner with a Diffusion Transformer renderer for video generation and editing.
It is a native macOS application built with Swift that allows users to download and trim online videos directly within the same interface.
A browser-based client-side video and audio processing interface powered by WebAssembly and FFmpeg.
Rescript is an open-source, transcript-based video and audio editor that runs entirely in your web browser.
Performs real-time 3D full-body reconstruction and 70-joint skeleton tracking from a single camera using ONNX and ggml in pure C++.
Zero-shot expressive voice cloning and speech generation tool that creates realistic emotional delivery from a 10-second reference clip.
Scal3R performs scalable test-time training for large-scale 3D reconstruction and depth estimation from visual inputs.
An open-source post-training framework using on-policy self-distillation to speed up few-step autoregressive video generation models.
An open-source Python tool that automates highlight detection, translation, subtitle generation, and voiceovers to turn long YouTube videos into viral short-form clips.
An open-source Android application that performs real-time OCR and translates foreign text directly over games, visual novels, and manga without requiring root access.
An open-source Python tool that uses Google Gemini and FFmpeg to automatically assemble raw footage into a polished ad based on a creative brief.
A local-first browser-based AI video editor powered by WebGPU that handles music generation, audio repair, voiceovers, captions, and talking avatars.
MockingBird is an open-source Python tool that clones a human voice using just a 5-second audio sample to generate arbitrary real-time speech from text.
VoxCPM is an open-source, tokenizer-free text-to-speech model capable of multilingual speech generation, creative voice design, and realistic voice cloning.
An open-source, offline speech-to-text desktop application built with Rust and Tauri that converts audio and video files into text locally.
ScreenToGif is an open-source Windows application that records a selected screen area, webcam, or sketchboard, and lets you edit and export it as a GIF or video.
Recordly is an open-source desktop screen recorder that automatically adds smooth zoom, pan, and cursor animations to your screen captures.
It converts e-books into spoken audiobooks in over 1,158 languages using various text-to-speech engines and custom voice cloning.
NVIDIA NeMo is an open-source conversational AI framework for building, training, and fine-tuning automated speech recognition and text-to-speech models.
Sherpa-ONNX runs speech recognition, text-to-speech, speaker diarization, and audio separation locally using ONNX Runtime without needing an internet connection.
An end-to-end AI platform that generates complete vertical short dramas from a single text prompt, handling everything from scriptwriting to final video assembly.
Qwen3-TTS is an open-source text-to-speech series by Alibaba Cloud that enables stable, expressive speech generation, free-form voice design, and vivid voice cloning.
PaddleSpeech is an open-source Python toolkit providing state-of-the-art models for automatic speech recognition, text-to-speech, speaker verification, and speech translation.
A comprehensive Gradio WebUI bundling Edge-TTS, Kokoro, F5-TTS voice cloning, Whisper transcription, Demucs vocal isolation, and YouTube downloading into one interface.
It provides a Python interface and CLI to access Microsoft Edge's high-quality online text-to-speech voices directly without needing an Edge browser or API key.
Moonshine is a C++ and Python-based speech recognition and synthesis toolkit optimized for extremely low-latency voice agents and transcription.
WhisperLiveKit provides real-time speech-to-text transcription using OpenAI's Whisper models via WebRTC and local processing.
Try adjusting your filters or search query.