
Midjourney to Video with GPT Image 2 and Seedance
The three-model Midjourney to video workflow: build the look in Midjourney, rebuild it as a controlled board in GPT Image 2, then animate it in Seedance.
A Midjourney still is a dead end. The image is gorgeous, the client wants motion, and the moment you feed that frame to a video model the face drifts, the palette shifts, and the thing that made the still worth keeping is gone.
The workflow that actually survives this is not one model. It is three, each doing the job it is good at. A Midjourney to video pipeline posted in August 2026 by Kōda, a creative technologist, ran exactly that stack — Midjourney v8.2, GPT Image 2, then Seedance 2.5 — and drew 57.5K views, 719 likes and 481 bookmarks[1]. The interesting part is not the output. It is why a middle step exists at all.
TL;DR
- Midjourney is an aesthetics engine. It does not follow literal instructions well and it cannot reliably render text, so it is the wrong tool for a layout-controlled reference sheet[2].
- The fix is an instruction-following image model in the middle. GPT Image 2 turns one Midjourney frame into a controlled character or environment board that a video model can read.
- Seedance 2.5 takes it to motion: text, image or reference input, up to 30 seconds, with native audio.
- Seedance 2.0 Mini covers the same three input modes at 480p and 720p and is the cheaper tier for iteration passes.
- Character consistency on Midjourney v8 comes from a style reference plus a fixed seed, not from
--cref, which Midjourney's own compatibility chart marks unsupported on v8[2]. - Every model in this Midjourney to video chain runs on this site, so the handoffs are file transfers rather than three separate subscriptions.
Why one model cannot do the whole job
Every model in this chain is missing something the next one has.
Midjourney has the strongest visual identity of any current image model, which is exactly why art directors keep using it. That identity is also why it ignores you. Ask for a four-panel sheet with labelled views at fixed positions and you get four beautiful images that are not a sheet. Midjourney cannot reliably put readable text in an image either — short words sometimes land, longer strings come out garbled.
An instruction-following model has the opposite profile. It will hold a layout, place elements where you asked, and render legible labels. Its default look is plainer.
A video model needs a clean, unambiguous starting frame. Feed it a moody Midjourney render with three characters at different depths and it has to guess what moves. Feed it a controlled board and the guesswork drops.
So the division is not arbitrary: Midjourney for the look, an instruction model for the structure, a video model for the motion.
Step one: generate the look in Midjourney
Start where the aesthetic is decided. In the workflow above the Midjourney prompt was two words plus flags[1]:
anime girl --ar 2:3 --profile geutxpd x2o8dm8Two things to copy here. The aspect ratio is set at generation time rather than cropped later, because ratio changes composition in Midjourney rather than just trimming edges. And consistency is carried by a Personalization profile, not by a character reference.
That second point trips people up. --cref and --cw are marked unsupported on v8.1 and v8.2 in Midjourney's official compatibility chart; they are V6-era features[2]. On v8 you hold a set together with a style reference, a reused seed, and a written description you refuse to reword. If you want the full version of that argument, I wrote it up separately in Midjourney V8 character consistency.
One Midjourney run returns up to four images. Pick the frame with the clearest silhouette, not the prettiest one. The video model will thank you later. You can run this step at seedance2.so/midjourney.
Step two: rebuild the frame as a controlled board
This is the step people skip, and skipping it is why their Midjourney to video attempts drift.
Take the chosen Midjourney frame into GPT Image 2 and ask for a reference board. The instruction from the workflow above was to "create a 2:1 premium cinematic CHARACTER IDENTITY BOARD using Image A as the sole character and style reference"[1], with the character named and the source image pinned as the only authority for appearance.
Three things make this work:
One image is the sole authority. Naming the reference explicitly stops the model from averaging your character with its own priors.
The layout is specified. A 2:1 board with defined regions is a structural instruction, which is what instruction-following models are for and what Midjourney is not.
Labels are readable. If the board carries view names or notes, they come out legible.
The output is a single frame that encodes the character from more than one angle, which is a much better prompt for a video model than a single moody portrait. Run it at seedance2.so/text-to-image.
Skip this step when your shot is a landscape, a product, or anything without a character who has to stay recognisable. Environment boards follow the same pattern when the location has to repeat across shots.
Step three: put it in motion with Seedance
Seedance 2.5 accepts text, image, or reference input and produces clips up to 30 seconds with native audio. For this pipeline the image path is the one you want: the board from step two becomes the visual anchor, and the prompt describes motion rather than appearance.
That split matters. Once the frame carries the appearance, your prompt should stop describing the character and start describing what happens — camera move, subject action, pacing. Prompts that re-describe a face already present in the reference tend to fight it.
Seedance 2.5 or 2.0 Mini? By capability, not by habit. Seedance 2.5 is the one to use when a shot has to carry sound or run long, since native audio and the 30-second ceiling are its differences. Seedance 2.0 Mini covers the same three input modes at 480p and 720p and is the sensible tier for iteration passes where you are testing motion ideas and will regenerate a dozen times.
A practical order: block out the motion on the cheaper tier until the movement reads correctly, then run the keeper on 2.5 for the finish. Start at seedance2.so/image-to-video.
Where a Midjourney to video pipeline usually breaks
Drift between step one and step three. Almost always caused by skipping step two and handing the video model an ambiguous frame.
A prompt that fights the reference. If the reference already establishes the character, describing the character again in the video prompt gives the model two sources of truth.
Ratio changed mid-pipeline. Decide the aspect ratio at the Midjourney stage and keep it through the board and the clip. Changing it later means cropping, and cropping a composition that was built for a different shape rarely survives.
Expecting a face lock. No current setup gives you a bit-exact character across shots. A style reference plus a fixed seed plus a controlled board gets you a coherent set, which is what production actually needs.
FAQ
Can I go straight from Midjourney to video and skip GPT Image 2?
Yes, and for simple shots you should. A landscape, an object, or a single-character close-up with an unambiguous silhouette animates fine from a raw Midjourney frame. The middle step earns its place when a character has to stay recognisable across several shots.
Why not do the whole thing in one image model?
Because the two halves want opposite behaviour. Midjourney's strong aesthetic is a feature when the look carries the work and an obstacle when you need a fixed layout with readable labels. Using both is cheaper than fighting either one.
Does the Midjourney version matter for this workflow?
Yes. v8.2 and v8.1 are the current renders and behave differently from v7 and the v5.x line. It matters most for consistency: Character Reference is unsupported on v8, so a v6-era tutorial will send you after a flag that no longer exists[2].
How long can the final clip be?
Seedance 2.5 goes up to 30 seconds with native audio. For anything longer the normal approach is several clips cut together, which is also why holding the character consistent across shots matters.
Can I use a reference video instead of a still?
Yes. Seedance 2.5 and 2.0 Mini both support a reference-to-video mode alongside text and image input, which is useful when you want to carry motion or framing from existing footage.
Which step should I iterate on first?
The Midjourney frame. It is the cheapest to regenerate and it determines everything downstream. Getting motion right on a frame you do not like is wasted work.
Does this need three separate subscriptions?
Not here. Midjourney, GPT Image 2 and both Seedance tiers run on this site, so moving between steps is a file handoff. Midjourney has no official API of its own, which is why a hosted route exists at all.
What resolution should I generate at?
Match the last step. Seedance runs at 480p and 720p, so a 4K board is wasted effort. Generate the board at a resolution close to your delivery target and spend the difference on more iterations.
Picking your battles in the chain
The instinct with a three-model pipeline is to optimise every step. In practice one step decides the outcome: the frame you hand the video model. Midjourney decides whether the look is worth animating, the board decides whether the video model understands what it is looking at, and Seedance mostly does what the frame tells it.
If you only change one thing about your current Midjourney to video process, add the board step. It is the difference between a clip that drifts and a clip that holds.
References
- Kōda [@aimikoda]. Midjourney v8.2 + GPT Image 2 + Seedance 2.5 workflow. Posted August 8, 2026. Retrieved August 2026 from x.com/aimikoda/status/2085803816553685091
- Midjourney. Version — feature compatibility and comparison chart. Retrieved August 2026 from docs.midjourney.com/hc/en-us/articles/32199405667853-Version
Author

Categories
More Posts

Seedance 2.5 vs Hailuo H3: A Creator's Test Plan
Compare Seedance 2.5 vs Hailuo H3 on 30-second scenes, 2K output, native audio, reference inputs, API availability, pricing, and production risk.


How to make AI product videos from product photos
Make AI product videos from a product photo: first-and-last-frame control, slow camera-move recipes, a detail-fidelity checklist, and UGC-style variations.


Is Seedance 2.0 Available in the USA? The Real Answer
Yes, Seedance 2.0 is available in the USA, with no waitlist and no VPN. Here is why the official API has no US region, and why that never reaches your workflow.

