Перейти к содержимому

I Made a Complete Local AI Film with MiniMax H3 + ComfyUI

Smart Vision

0:00 / 0:00

I Made a Complete Local AI Film with MiniMax H3 + ComfyUI

689 просмотров · 10 ч назад
Smart Vision
5,14 тыс. подписчиков
689 просмотров · 10 ч назад
In this video, I use MiniMax H3 and ComfyUI to turn two character images and two voice references into a complete 38-second AI short film, entirely locally on my own PC. Instead of generating a collection of disconnected AI clips, this workflow combines Short Clips and Long Video Chaining to create a more complete cinematic scene with recurring characters, recognizable voices, dialogue, and story continuity. The main two-character dialogue is generated as a 23-second Long Clip using four Generation Windows and Motion Context. The close-up reveal and ending are generated separately as Short Clips, making it easier to regenerate important shots without rebuilding the entire scene. This workflow can also run on a GPU with just 8GB of VRAM, although lower resolutions or more aggressive memory offloading may be required if your system only has 32GB of RAM. In this video, you’ll learn: • How to use two character references in MiniMax H3 • How to assign a different Ref Audio voice to each character • How to control dialogue order, expressions, and character actions • How Short Clips and Long Clips serve different purposes • How Motion Context connects multiple Generation Windows • How to maintain character, voice, and motion continuity • How to combine multiple generated shots into a complete AI film • What can cause blur, artifacts, or detail loss in longer generations • How this workflow performs on low-VRAM hardware --- The project includes three workflows: 1. Short Dialogue Workflow 2. Long Dialogue Workflow with Motion Context 3. Complete Film Workflow 📦 Workflows, model information, and download links:   / smartvisionofficial   You can join my Patreon for free to access available downloads, updates, and additional ComfyUI resources. --- 🎙️ Write prompts and scripts faster with Typeless Typeless turns natural speech into clean, polished text. 👉 Try Typeless: https://www.typeless.com/?via=smartvi... Affiliate disclosure: Smart Vision may earn a commission from purchases made through this link. --- Hardware used for this video: GPU: NVIDIA RTX 4070 Ti Super 16GB System RAM: 64GB Output resolution: 1280 × 704 Short Clip generation time: Approximately 5 minutes Long Clip: Approximately 23 seconds using four Generation Windows Final film length: Approximately 38 seconds --- Models and tools: • MiniMax H3 Ref2VA • MiniMax H3 Ref2V Turbo LoRA • ComfyUI • Motion Context • Sage Attention --- Timeline: 00:00 Intro 01:43 PLAN THE FILM, Build a Scene, Not Just Random Clips 02:36 CHARACTERS + VOICES, Lock the Characters and Their Voices 04:30 INSTALLATION + WORKFLOW, How the Three Workflows Are Used 07:56 SHORT CLIP, Test the Important Things First 09:45 LONG CLIP The Real Challenge: A Two-Character Dialogue 12:17 FINISHING THE FILM, Complete the Scene 14:11 PERFORMANCE, How Much Time Did It Actually Take? This is not yet a perfect one-click AI movie generator. Character movements, speaking order, facial expressions, camera direction, and transitions still require testing and prompt refinement. However, it brings local AI filmmaking much closer to producing complete, coherent scenes instead of rebuilding everything every few seconds. If you enjoy local AI filmmaking, low-VRAM workflows, and ComfyUI tutorials, subscribe to Smart Vision. #MiniMaxH3 #ComfyUI #AIFilmmaking #imagetovideo #characterconsistency #videogeneration #freeaivideogenerator