The problem
I'd already built AI image pipelines for client work, USP extraction, generation, assembly, the whole chain, for Rubberfashion and internally at Midnight Majestic. Video was the obvious next step, and also the one where I had the least hands-on depth. I didn't want to be the operations person who greenlights a video tool based on a sizzle reel some vendor sent over. I wanted to actually know what I was asking a team to evaluate.
So I started building, not client work first, personal narrative projects where I could fail without anyone's budget on the line.
What I found
Most tools in this space compete on hype. New model drops, timeline fills up with demo clips, everyone moves on to the next one two weeks later. Underneath that noise, two things actually determine whether a tool survives contact with real production work.
First, whether it takes multiple reference images or only a start and end frame. Start/end interpolation is fine for a single transition. It falls apart the moment you need a character or product to stay consistent across a sequence. Second, native audio. A few well-hyped models generate great video and nothing else, which means every clip needs a separate audio pass bolted on after. That's not a minor inconvenience, it's a workflow that doesn't scale past one-off demos. I ruled out several tools on that basis alone, no matter how good the visuals looked.
What I did about it
I ran two personal narrative series as the testbed: Adventures of Elena, and The ParaNormals, following three characters, Mae, Reese, and Nomi, across a set of shots. That gave me a reason to actually push a pipeline past a single clip, into continuity, character consistency, and a real production sequence.
For core generation I settled on GPT, Grok Imagine 2, Seedream, Seedance, Nano Banana, and Veo, the tools that held up against the reference and audio criteria. Magnific handles upscale. Creatify and Higgsfield cover motion and avatar work. ElevenLabs for voice. I also tested Dola, Dreamina, Lumina, Modelark, Midjourney, Runway, Kling, and Krea against the same criteria, some earned a place in the stack, some didn't.
I also made a deliberate call on where to run all of it. A lot of platforms lock you into their own small set of native models, which is fine until that platform's roadmap stops matching what you actually need. I built the pipeline on top of access layers instead: ComfyUI, fal.ai, Wavespeed, OpenRouter, kie.ai, Wireflow, Figma Weave, so the flow isn't married to any one vendor's decisions about what to support. Swap the underlying model, the pipeline still runs. That's the same reasoning behind every SOP I've written: don't rebuild the process from memory every time, build it once so it survives the tool underneath it changing.
I also ran one live piece against actual client material, a product video for Rubberfashion, to check the pipeline against something with real constraints instead of only personal creative freedom.
What happened
I'm a couple months into this, and I'll say that straight. Two of the personal pieces are up on YouTube and picked up real views as experimental UGC-style tests, but I haven't pulled hard performance numbers on those yet, and I'm not going to claim a metric I haven't actually measured. What I have is a working pipeline, a clear model-evaluation framework, and one client piece that proves it holds up outside personal projects.
I've spent twenty-three years being the person who builds the process other people run through.
That only works if I actually understand what I'm asking them to do. This is me doing the hands-on work myself, first, so that when I'm leading a team through the next wave of tools, I'm evaluating them the same way I evaluated these: on whether they hold up in production, not on how good the demo looked.
