Midjourney Image-to-Prompt Guide: How to Turn Any Image Into a Prompt That Actually Works
The first time I tried this, I did the dumb thing.
I found a gorgeous portrait on Pinterest, ran
it through Midjourney's `/describe` command, copied the first
suggestion word-for-word, hit enter, and waited for magic.
What I got back was… fine. Technically
competent. Completely different from the image I loved.
That's the moment most people give up on
image-to-prompt and decide it "doesn't work." It does work — but not the way you think. Midjourney
doesn't reverse-engineer an image. It guesses at it. Your
job is to take that guess and turn it into something useful.
This guide is the
Midjourney-specific version of the process. If you want the platform-agnostic version — the one that
works across Midjourney, DALL·E, Stable Diffusion, and Flux — start with the pillar guide: Image to Prompt: The Ultimate Guide + Free
Framework. Then come back here for the Midjourney-only tools, parameters, and
quirks.
Let's get into it.
What "Image to Prompt" Actually Means in Midjourney
Here's the thing nobody explains clearly: "image to prompt" in Midjourney is really two completely different jobs, and mixing them up is why people get frustrated.
| Job | What it does | Tool | ||
|---|---|---|---|---|
| Read the image → get text | Midjourney looks at your image and writes a text prompt describing it | `/describe` | ||
| Use the image → guide the output | Midjourney uses the image itself as visual input during generation | Image prompt (`--iw`), `--sref`, `--oref` |
Job one gives you vocabulary. Job two gives you visual control.
The best results come from using both together.
`/describe` teaches you how
Midjourney thinks about your image, and reference parameters make sure the model actually
looks at it while generating.
Most tutorials only teach you one half. That's why their results look
loose.
Method 1: The `/describe` Command (Midjourney Built-In Image-to-Prompt)
This is the native image-to-prompt tool, and it's genuinely useful — as long as you know what you're getting.
How to run `/describe` on Discord
- Type `/describe` in any Midjourney bot channel or your DM with the bot.
- Choose the `image` option and upload your file (or use the `link` option and paste a direct image URL).
- Hit enter.
- Midjourney returns four prompt suggestions, numbered 1–4.
Under those suggestions you'll see buttons:
- 1️⃣ 2️⃣ 3️⃣ 4️⃣ — generate a grid from that specific
prompt
- 🔄 — re-roll and get four fresh descriptions
- Imagine all — run all four at once
How to do it on the Midjourney web app
If you're working on midjourney.com instead of Discord, upload your image into the prompt bar (the image
icon on the left), then look for the Describe option in the image's menu. Midjourney
redesigns this panel fairly often, so if the button has moved, it's usually inside the same dropdown
where you assign the image as a style or character reference.
Same output, nicer interface.
What those four results actually are
They're not four attempts at the same answer. They're four different interpretations, usually with different emphasis:
- One leans photographic (camera, lens, lighting terms)
- One leans artistic (medium, artist-adjacent vibes, art movements)
- One leans descriptive (subject, colors, composition)
- One is a bit of a wildcard
Read all four. Don't pick one — harvest all four.
What `/describe` gets right
- Medium and technique. It's good at spotting oil painting vs. 3D render vs. film photography.
- Color and mood language. It'll hand you words like "muted earth tones," "chiaroscuro," "hazy golden light."
- Aspect ratio. It usually appends the correct `--ar`.
- Art vocabulary you didn't know. Honestly the biggest benefit. It'll say "tenebrism" or "isometric" or "cinemagraph" and suddenly you have a search term.
What `/describe` gets wrong (every single time)
- It doesn't know your subject. A photo of your specific product becomes "a sleek modern bottle." A photo of your friend becomes "a young woman."
- It hallucinates specifics. Camera bodies, lens focal lengths, film stocks, artist names — often invented to sound plausible.
- It flattens composition. It rarely tells you where things sit in the frame, which is often the single most important thing about the image.
- It over-indexes on style, under-indexes on structure. You'll get the mood back. You won't get the layout.
So: treat `/describe` output as a word bank, not a finished prompt. That single mindset shift fixes 80% of the frustration.
Method 2: Image Prompts (`--iw`)
Instead of describing the image, you can just… give Midjourney the image. How it works: paste a direct image URL (or upload the image) at the start of your prompt, then add your text after it.
https://example.com/reference.jpg a ceramic coffee cup on a windowsill, soft morning light --iw 1.5 --ar 3:2
The key parameter is `--iw` (image weight). It controls how much the image influences the result versus your text.
- Lower values (0.25–0.5) → your text dominates, image is a whisper
- Default (1) → balanced
- Higher values (2–3) → the image dominates, text becomes a nudge
The exact accepted range shifts between model versions, so if you get an error, drop the number and try
again.
When to use image prompts: when you want the overall composition and content echoed — a similar layout, similar
objects, similar framing.
When not to use them: when you only want the look. For that, you want style reference instead.
I go
deeper into choosing between these in How to Use
Reference Images in AI Generation — worth reading if you're regularly working from
source material.
Method 3: Style Reference (`--sref` + `--sw`)
This is the one that changed how I work.
`--sref` tells
Midjourney: ignore what's in this image, just steal how it looks. Color grading, brush texture,
grain, lighting quality, rendering style.
a fox standing in a snowy field --sref https://example.com/style.jpg --sw 300 --ar 16:9
Style weight (`--sw`) controls intensity:
- 0–50 — a faint influence
- 100 — default, noticeable but balanced
- 300–500 — strong stylization
- 700–1000 — the style takes over almost completely
You can stack multiple references:
--sref https://img1.jpg https://img2.jpgAnd you can weight them individually with `::`:
--sref https://img1.jpg::2 https://img2.jpg::1Pro move: `--sref random` gives you a random
style seed. When it produces something you love, Midjourney shows you the seed number, and you can reuse
it forever with `--sref 1234567890`. It's like building your
own private style library without hosting a single image.
When to use `--sref`: brand consistency, series work, matching
a visual identity across dozens of images. This is the closest thing Midjourney has to a "house style"
button.
Method 4: Omni Reference (`--oref` + `--ow`)
In older versions (V6), character consistency ran on `--cref` and `--cw`. In V7, Midjourney replaced that with Omni Reference — `--oref` — which is broader. It can carry over a person, a creature, an object, or a specific prop into new scenes.
a knight standing at the edge of a cliff at sunset --oref https://example.com/face.jpg --ow 200 --ar 2:3
Omni weight (`--ow`) controls how tightly Midjourney clings to the reference:
- (25–75) — loose likeness, lots of stylistic freedom. Good when you're pushing the subject into a very different art style.
- Default (100) — balanced.
- High (200–400) — strong likeness, good for photorealistic consistency across a series.
Midjourney's parameter ranges and defaults shift with every model release, so if a number doesn't behave, check the official docs before assuming you broke something.
Rule of thumb I use:
- Want the same person, new scene? → `--oref`
- Want a different subject, same aesthetic? → `--sref`
- Want both? → use them together, and expect a few rounds of tuning.
Method 5: Personalization and Moodboards (`--p`)
Once you've rated enough images, Midjourney unlocks personalization — a profile of your
taste that you apply with `--p`.
You can also build
moodboards: upload a set of images that represent a look, and Midjourney generates a
personalization code from that collection. It's essentially `--sref` for a whole aesthetic rather than a single
image.
This is the underrated one. If you're producing images for the same brand week after week,
spending 30 minutes building a moodboard saves you hundreds of prompt tweaks later.
The 8-Part Midjourney Prompt Formula (That `/describe` Won't Give You)
/describe` gives you adjectives. This gives you structure. Combine both and your hit rate goes way up.
- Subject — who or what, with specific detail
- Action or pose — what they're doing, how they're positioned
- Setting — where, and what's around them
- Composition and camera — framing, angle, distance, depth of field
- Lighting — direction, quality, color temperature
- Medium and style — photo, oil painting, 3D render, risograph
- Color palette — dominant colors, saturation, grading
- Mood — the emotional read
Then parameters at the end.
Watered-down version (what most people write):
a woman in a caféFormula version:
a woman in her thirties sitting alone at a marble café table, chin resting on her hand, looking out of frame, rain-streaked window behind her, medium close-up at eye level with shallow depth of field, soft diffused window light from the left with deep shadows on the right, 35mm film photograph, muted teal and warm amber palette, quiet and contemplative --ar 4:5 --style raw --s 200
Same idea. Wildly different output.
Notice how the second one hits all eight parts without
sounding like a robot reading a checklist. That's the goal — write it like you're describing the shot to
a photographer over the phone.
My 5-Step Midjourney Image-to-Prompt Workflow
This is the actual process, in order.
Step 1: Describe the image yourself first (60 seconds)
Before touching `/describe`, write down what *you* see. Just
bullet points. Subject, layout, lighting, mood.
Why? Because once you read Midjourney's
description, you'll anchor to it. Your own read catches things the model misses — especially composition
and intent.
Step 2: Run `/describe` and mine it for vocabulary
Now run it. Read all four results. Highlight every phrase that makes you go "oh, that's the word for
it."
Ignore everything else. Especially fake camera specs.
Step 3: Merge into one prompt using the 8-part formula
Your bullet points supply the structure. `/describe` supplies
the vocabulary. Combine them into one clean prompt.
Keep it under roughly 60 words if you can.
Midjourney weights early words more heavily, so front-load your subject and the single most important
visual quality.
Step 4: Add parameters and generate a grid
Set your `--ar` to match the reference. Add `--style raw` if
you're chasing realism. Add `--sref` or `--oref` if you need the actual image in the
loop.
Generate. Look at all four results, not just the one you like.
Step 5: Fix ONE thing per round
This is the discipline that separates people who get there in three rounds from people who flail for
thirty.
Look at your grid, identify the single biggest gap between what you got and what you
want, and change only the words responsible for that gap.
Wrong lighting? Rewrite the
lighting clause. Don't touch anything else.
Use Vary (Subtle) to nudge,
Vary (Strong) to explore, and Remix when you need to swap words
mid-flight.
Midjourney Parameter Cheat Sheet for Image Matching
| Parameter | What it does | Use it when | ||
|---|---|---|---|---|
| `--ar` | Aspect ratio | Always. Match your reference exactly. | ||
| `--style raw` | Reduces Midjourney's default "pretty" aesthetic | Matching photos, documentary looks, product shots | ||
| `--s` (stylize) | 0–1000, how much artistic license Midjourney takes | Lower (0–100) for literal, higher (500+) for stylized | ||
| `--c` (chaos) | 0–100, variety across the grid | Raise it when you're exploring, drop to 0 when refining | ||
| `--weird` | 0–3000, unconventional aesthetics | Rarely, when everything looks too generic | ||
| `--iw` | Image prompt weight | Using an image URL in the prompt | ||
| `--sref` / `--sw` | Style reference and strength | Matching a look across many images | ||
| `--oref` / `--ow` | Omni reference and strength | Keeping a character or object consistent | ||
| `--no` | Negative prompt | Removing recurring unwanted elements | ||
| `--q` | Quality / render time | Final renders, detailed textures | ||
| `--seed` | Reproducible starting noise | Locking in a result to make small variations |
The two I reach for most: `--style raw` and `--s`. If your output looks "too Midjourney" — too glossy, too dramatic, too illustrated — those two are almost always the fix.
Worked Example: Matching a Moody Café Portrait
Let's run the whole thing.
Reference image: a woman by a rain-streaked café
window, soft grey light, film grain, muted colors, 4:5 vertical.
Step 1 — My own read:
- Vertical, subject on the left third
- Face lit from window on the left, right side falls into shadow.
- Visible grain, slight halation on highlights
- Cool grey-blue palette with one warm accent (a coffee cup)
- Feels lonely, not sad
Step 2 — `/describe` output (paraphrased highlights):
- "cinematic film still, muted color grading"
- "shallow depth of field, bokeh background"
- "pensive atmosphere, natural window light"
- (Also gave me a made-up camera model and a lens spec I ignored)
Step 3 — Merged prompt:
a woman in her thirties seated beside a rain-streaked café window, chin resting on her hand, gazing out of frame, positioned on the left third, medium close-up at eye level, shallow depth of field, soft directional grey window light from the left with deep falloff on the right, 35mm film still with visible grain and gentle highlight halation, muted grey-blue palette with a single warm amber accent, quiet and pensive --ar 4:5 --style raw --s 150
Step 4 — First grid result: lighting was right, grain was right, but the framing was too
tight and the mood read as sad rather than contemplative.
Step 5 — One fix:
changed "medium close-up" to "medium shot with headroom above the subject" and swapped "pensive" for
"calm, self-contained." Left everything else alone.
Round three was the keeper.
Total
time: about eight minutes. Compare that to copy-pasting `/describe` and re-rolling forty times.
Troubleshooting: Why Your Midjourney Output Doesn't Match
| Problem | Likely cause | Fix | ||
|---|---|---|---|---|
| Too glossy, too "AI-looking" | Default Midjourney aesthetic | Add `--style raw`, lower `--s` to 0–100 | ||
| Style is right, subject is wrong | You used `--sref` when you needed `--oref` or an image prompt | Switch reference type | ||
| Composition keeps drifting | Prompt has no framing language | Add camera angle, shot distance, subject placement | ||
| Face changes every generation | No character anchor | Add `--oref` with your reference, raise `--ow` | ||
| Grid is too random | Chaos too high | Set `--c 0` | ||
| Colors are off | No palette specified | Name 2–3 specific colors, add a grading term | ||
| Prompt is long but nothing works | Contradictions inside the prompt | Cut it in half; conflicting instructions cancel out | ||
| Reference barely influencing output | `--iw` or `--sw` too low | Raise `--iw` toward 2, or `--sw` toward 400 | ||
| Reference overpowering your text | Weights too high | Drop `--iw` to 0.5, `--sw` to 100 |
Copy the Vibe, Not the Artwork
Quick real talk, because this matters.
Midjourney's reference tools are powerful enough now that
you *can* get uncomfortably close to someone else's specific piece. Just because the tool lets you
doesn't mean it's a good idea — creatively or legally.
The line I hold to: techniques and
aesthetics are fair game, specific artworks and living artists' signature output are
not.
"Chiaroscuro lighting, muted palette, film grain" is a description of craft.
Naming a working illustrator and asking for their exact style is something else.
I've written
about where that line sits, how to use references responsibly, and how to build something that's
genuinely yours in How to Recreate an Image
With AI (Honestly)**. If you're using image-to-prompt for client work, read that
one before you ship anything.
Mistakes I See Constantly
Copy-pasting `/describe` output raw. It's a first draft written by something that has
never seen your intent. Edit it.
Not matching the aspect ratio. Change the ratio
and you change the composition entirely. Set `--ar` before you tune anything
else.
Writing 200-word prompts. Midjourney doesn't reward length. It rewards
clarity and priority order. Long prompts dilute your key terms.
Changing five things at
once. Then you have no idea what worked. One variable per round.
Ignoring
the other three grid images. Sometimes image #3 has the lighting you wanted on the
composition you didn't. That's information.
Chasing pixel-perfect matches.
Midjourney is a re-interpreter, not a photocopier. Aim for "same feeling, same craft level" and you'll
be happy. Aim for "identical" and you'll burn a hundred credits.
Skipping the
seed. When something is *almost* right, grab the seed and iterate from there instead of
re-rolling from scratch.
FAQs
Does Midjourney have an image-to-prompt feature?
Yes — the `/describe` command. Upload an image (or paste a URL) and Midjourney returns four text prompt suggestions based on what it sees. It's available on Discord and in the web app.
How accurate is Midjourney's `/describe` command?
It's reliable for medium, style, mood, and color, and unreliable for specific subjects, composition, and technical camera details. Treat it as a vocabulary generator rather than a finished prompt.
Can Midjourney recreate an image exactly?
No. Even with an image prompt at maximum weight, Midjourney reinterprets rather than reproduces. The realistic goal is capturing the style, lighting, composition, and mood — not pixel-level duplication.
What's the difference between `--sref` and `--oref`?
`--sref` copies the *look* of a reference (color, texture, rendering style) while ignoring its content. `--oref` copies the *subject* — a person, creature, or object — and places it into a new scene.
What is `--iw` in Midjourney?
Image weight. It controls how strongly an image prompt influences the output relative to your text. Default is 1; lower values favor your text, higher values favor the image.
Should I use `/describe` or a third-party image-to-prompt tool?
`/describe` is tuned to Midjourney's own prompt language, so its vocabulary lands better inside Midjourney. Third-party tools often give more literal, structured descriptions. Running both and merging the results is genuinely the strongest approach.
Why do my Midjourney images look "too AI"?
Usually the default aesthetic. Add `--style raw` and lower `--s` to somewhere between 0 and 150. Then add real photographic language — imperfect lighting, grain, natural skin texture, off-center framing.
How do I keep the same character across multiple Midjourney images?
Use omni reference (`--oref`) with a clear, well-lit reference image, and tune `--ow` upward until the likeness holds. Keep the character description identical across prompts, and lock a seed if you need tight consistency.
Quick Recap
- `/describe` gives you vocabulary — not a finished prompt.
- Use the 8-part formula to give that vocabulary structure.
- Pick the right reference tool: image prompt for composition, `--sref` for style, `--oref` for subject.
- Set `--ar` first, `--style raw` and `--s` second.
- Fix one thing per round.
- Aim for the same feeling, not the same pixels.
Do those six things and Midjourney's image-to-prompt tools stop feeling like a slot machine and start feeling like a workflow.
Midjourney ships updates constantly — parameter names, ranges, and defaults change between model versions. If a flag in this guide throws an error, check the current official documentation, then come back and keep building.