I used to think every stage needed heavy guidance.
Then I started treating the still and the video as two different animals. The still demands precision. The video often demands restraint. That split changed how I work.
This is what actually works in my current process.
The still needs craft. The video needs breathing room.
I no longer begin with vague or open prompts for the base image. Especially in Midjourney, a well-defined prompt is usually required if I want usable results without massive cherry-picking. I describe subject, lighting, composition, mood, and key details clearly. Only then do I get a strong starting still worth sending to video.
Lab note: a “clean still” simply means a strong, coherent base image with good anatomy, usable lighting, and no major artifacts. Nothing mystical. Just something solid that the video model can interpret well.
Step 1: craft the still with intention
I write a detailed but focused prompt for the still generator. Subject description, pose, clothing, lighting, camera angle, art direction. I iterate on the still until I have a few strong candidates. This is where the real prompt work happens.
Step 2: send it to image-to-video with silence
Once I have a good still, I upload it and leave the text prompt field blank or with one or two minimal words at most. No camera moves, no emotion cues, no detailed motion descriptions unless absolutely necessary.
Step 3: let the model interpret
The video tool frequently adds natural micro-movements, breathing, subtle shifts in posture, and realistic physics on its own. These zero-prompt or near-zero-prompt clips often feel more alive than the ones where I tried to direct every second.
Step 4: light iteration only when needed
If the motion goes off track, I add a short corrective prompt on the next generation. Usually something brief like “subtle natural movement” or “gentle breathing.” Then I go silent again. Heavy prompting at the video stage still tends to stiffen the result.
Step 5: comparison testing
I run the same strong still through both heavy-prompt video and minimal-prompt video. The minimal versions win more often. They respect the original image’s energy instead of fighting it.
Lab note: the still and video stages have different optimal prompting strategies. Forcing the same approach on both weakens the final piece.
Tools & creative stack
Midjourney for well-crafted stills.
Grok Imagine for quick video tests.
Image-to-video tools (Grok and others) with empty or minimal prompt fields.
Iterative prompting only when the motion needs correction.
Discovery / takeaway
The still needs a well-defined prompt or careful cherry-picking. The video interpretation usually does not. Often the AI knows what to do when you give it a solid starting image and then get out of the way.
Key lessons that now guide my workflow:
Write detailed prompts for the still stage.
Use minimal or zero prompting for image-to-video.
A strong “clean still” is the real foundation.
Over-directing the video often kills the natural life the model wants to add.
Respect the different strengths of each tool stage.
TL;DR: Craft the still with care. Then shut up and let the video model do its thing. Minimal or no prompting on image-to-video frequently delivers the best, most natural motion.
I wasted a lot of time over-prompting every single step. Now I concentrate my prompt energy where it matters most (the still) and trust the video model with the rest. The work flows better, the motion feels more organic, and I spend less time fighting the tools.
Steve Teare
video alchemist
TerminallyBored.Monster
Palouse, Washington USA
