There are songs that you listen to.
And then there are songs that seem to listen back.
I started with a cover of “Kiss from a Rose” by Seal. After splitting the instrumental stem, I slowed the backing track down, and asked myself a dangerous question: What if I stopped trying to illustrate the song and instead let it become a story?
That question became The Rose and the Grey.
And, naturally, I made it much more complicated than it needed to be.
2/9
Step 1: Start with the feeling.
The first decision was to slow the song down.
The original track sits around 132 BPM. I brought the backing track down to roughly 105 BPM. That small change made a surprisingly large difference.
Suddenly the song had room.
It became less like a pop song and more like a memory — something suspended between waking and dreaming.
Lab note:
Tempo is emotional architecture.
You don’t just make a song slower. You change the amount of psychological space available inside it.
That gave me the atmosphere I wanted: gray ocean, fog, a distant tower, snow, a rose, and a woman standing somewhere between them.
3/9
Step 2: I decided not to make a music video.
At least, not in the conventional sense.
I wanted the female voice to carry the first part almost entirely by herself.
The lyrics become a kind of confession:
There used to be a greying tower alone on the sea.
And you became the light on the dark side of me.
So I built a visual world around those ideas rather than trying to illustrate every line literally.
The tower became loneliness.
The light became the mysterious person who changes that loneliness.
The rose became desire, beauty, vulnerability — and perhaps something more dangerous.
And the grey?
Well, the grey is everything we don’t understand.
That seemed appropriate.
4/9
Step 3: Build the world before building the shots.
One thing I’ve learned from working with AI image generation is that continuity doesn’t happen because you politely ask for it.
You have to engineer it.
I established a visual vocabulary first:
A solitary weather-beaten stone lighthouse.
A cold gray ocean.
A desolate northern coastline.
A deep crimson rose.
Cold blue-gray darkness.
One warm white light.
That last one mattered.
I wanted the lighthouse beam to become almost a character in the film.
It is always the same basic idea: cold world, warm light.
The rose is the other exception. Almost everything is desaturated except that deep crimson.
Lab note:
AI loves novelty.
Storytelling needs repetition.
If every image invents something new, you don’t have a visual language. You have 25 unrelated AI pictures.
5/9
Step 4: Meet Spark.
The biggest continuity challenge was the woman.
I’ve learned not to rely entirely on character reference images in Midjourney for this kind of work. In my experience, the reference can become too dominant. Instead of creating a new shot containing the same character, Midjourney sometimes creates another variation of the reference image.
So I took the boring approach.
Which, in AI filmmaking, is often the clever approach.
I created a locked character description and embedded the same description into every prompt in which she appeared.
Same age.
Same hair.
Same eyes.
Same facial structure.
Same clothing.
Same general physical presence.
Then I changed the environment and composition around her.
That gave me a much better chance of making the audience believe that this was one woman moving through one strange world.
6/9
Step 5: And then there was Forge.
Forge had to have his own locked description.
A man in his early 70s.
White beard.
Silver-white hair.
Blue-gray eyes.
Lean build.
Weathered but intelligent face.
Dark charcoal clothing.
Again, the complete description goes into every prompt where he appears.
But I made an important structural decision.
Forge doesn’t appear until after she has finished speaking.
He doesn’t interrupt her.
He listens.
That’s the whole point.
The woman’s voice gets almost a minute and a half to create the dream. Then we leave it hanging for a moment.
And Forge enters.
Not to explain it.
To respond to it.
That distinction became the emotional center of the whole project.
7/9
Step 6: The Ken Burns experiment.
I decided to build the entire film out of still images first.
Twenty-five of them.
Each image gets roughly seven seconds of movement through a Ken Burns-style effect.
Twenty-five images times seven seconds gives me about 2 minutes and 55 seconds.
Almost exactly the length of the piece.
This turned out to be a very good constraint.
Instead of asking AI to generate motion simply because I could, I had to make every still worth looking at.
Then I could add motion only where actual motion added something.
Snow falling.
Ocean moving.
Fog drifting.
A lighthouse beam sweeping across the water.
A rose floating on black water.
Maybe the rose opening.
Maybe.
I’m learning to be suspicious of “maybe” when AI is involved.
Sometimes the beautiful still is better than the technically impressive moving image.
Lab note:
Motion is not automatically animation.
Sometimes the most powerful movement in a film is a camera slowly moving toward a completely still face.
8/9
Step 7: Let the symbols do some work.
The further I got into the project, the more I realized I didn’t need to explain everything.
The woman doesn’t need to tell us what the rose means.
The lighthouse doesn’t need a backstory.
Forge doesn’t need to say what their relationship is.
Instead, the images keep returning to the same things.
The tower.
The light.
The rose.
The grey.
Then Forge finally says:
Some people arrive like answers.
You didn’t.
You arrived like a question.
And suddenly the entire film rearranges itself.
The question isn’t “Are these two people in love?”
The question is “What happens when another person changes the way you see your own darkness?”
That’s a much more interesting question.
And I don’t answer it.
9/9
The real lesson.
I started this project thinking I was making a music video.
I ended up making something closer to a short cinematic poem.
That’s one of the things I enjoy most about working with AI.
The tools don’t simply execute an idea. They push back.
Midjourney forces me to think about continuity and composition.
ElevenLabs forces me to think about the emotional difference between reading words and inhabiting them.
Editing forces me to discover whether the beautiful image actually belongs in the story.
And sometimes the machine gives me something I didn’t know I wanted.
That’s the happy accident.
That’s the alchemy.
The final image is Forge standing alone beneath the night sky. The lighthouse beam finally reaches him. A single red rose rests nearby.
And that’s it.
No embrace.
No explanation.
No tidy little bow.
Just light arriving in the darkness.
Maybe that’s enough.
TL;DR:
I slowed a 132-BPM song, built a world of towers, roses, snow and light, created 25 AI-generated stills, gave them slow cinematic movement, and let two voices tell a story about connection without ever quite defining it.
Sometimes the best way to tell a story is to leave a little grey around the edges.
Steve Teare
video alchemist
TerminallyBored.Monster
Palouse, Washington USA
