How-to Guides 5 min read Updated 2026-04-11

How do I write a good prompt for an AI image generator like Midjourney?

Quick answer

A good AI image prompt names five things in order — subject, medium or style, composition, lighting, and technical parameters — and you build it one layer at a time, checking the output after each addition instead of pasting one giant sentence and hoping.

One clean red brushstroke on a faint grid beside a muddy tangle of overlapping strokes, lit from the left.
One deliberate mark beats forty vague ones — each added word should pull the image somewhere specific, not fight the rest. AI-generated illustration

The core decision rule is simple: start with subject + medium + lighting, generate, then add exactly one modifier (a style, a lens, a mood) and compare the two images side by side.

If the new image is better, keep the modifier; if it is worse or just different in a way you did not want, delete it. This slow, one-change-at-a-time loop is what separates a prompt that works from a prompt that merely sounds impressive. Most beginners write prompts that are too long and too vague at the same time — forty adjectives, no structure — and then blame the model when the result is muddy.

Here is the mechanism behind that structure. Image models do not read your prompt like a person reading instructions; they turn each phrase into a direction in a mathematical space of images, and every word pulls the result somewhere. That is why order and specificity matter.

"A woman in a red coat" gives the model a subject and a color anchor. "A woman in a red coat, cinematic still, shot on 35mm film, soft window light from the left" gives it a subject, a medium, a camera reference, and a light direction — four separate pulls that reinforce each other instead of fighting.

Vague words like "beautiful" or "high quality" pull almost nowhere, because they describe your reaction to an image, not the image itself. Concrete nouns and technical terms pull hard. The other half of the mechanism is the negative space: what you leave out, the model fills in with its own defaults, and those defaults are often generic.

If you do not specify a background, you will usually get a soft blur. If you do not specify a lens, you get a neutral mid-range look.

Here is a full copyable example, built layer by layer, using Midjourney's syntax because it is the clearest place to see parameters at work. Layer one, subject and medium: `a lighthouse on a rocky cliff, oil painting`. Layer two, add composition and lighting: `a lighthouse on a rocky cliff, oil painting, low angle looking up, storm light breaking through clouds`.

Layer three, add the technical flags: `--ar 16:9 --style raw`. The `--ar 16:9` flag sets the aspect ratio to widescreen, which changes how the model frames the scene — a wide canvas pushes it toward landscape composition, while `--ar 2:3` would push it toward a vertical, poster-like framing.

The `--style raw` flag tells Midjourney to lean less on its own aesthetic polish and follow your prompt more literally, which is useful when you want the oil-painting texture to show rather than Midjourney's default glossy look. According to our AI tool database, Midjourney's current version is V7, which added Draft Mode, Omni Reference, Personalization v2, and a Niji 7 anime mode — Draft Mode in particular is worth knowing about, because it lets you iterate on prompt wording quickly at lower fidelity before committing to a final render.

The same five-part structure transfers to other tools; the syntax changes but the logic does not. Adobe Photoshop's Firefly Generative Fill, for instance, takes short descriptive phrases rather than flag syntax, and the open-source Stable Diffusion WebUI (AUTOMATIC1111) uses separate positive and negative prompt boxes plus numeric weights, so you can write `(storm light:1.3)` to push one element harder.

Our database rates Midjourney 4.8 out of 5 and Photoshop 4.8 out of 5 on editorial quality, but that tells you about the tools, not about your prompt — a well-structured prompt in a mid-tier tool beats a lazy prompt in the best one.

Now the limits, because this method is not magic. First, it costs time: the one-modifier-at-a-time loop means many more generations than a single paste-and-pray attempt, and on paid plans each generation counts against your allowance. Second, it fails when your goal is genuinely abstract — if you want "a feeling of loneliness," no amount of structure will pin that down, because the model needs something visible to draw.

Translate the feeling into a scene first: an empty bus stop at dusk, one figure, long shadow. Third, prompt structure cannot fix a model that has not learned your subject; rare characters, specific real people, and niche objects often come out wrong no matter how you phrase them, and reference-image features like Omni Reference exist precisely because text alone hits a wall there.

Fourth, parameters are tool-specific and change between versions — a flag that worked last year may be renamed or retired, so check the vendor's own documentation rather than trusting an old prompt you found online. Finally, be honest about what you are optimizing: a prompt that produces one great image is not the same as a prompt that produces a consistent series, and consistency usually requires fixed seeds or reference images, not better adjectives.

How this page was produced: this answer was generated by an automated content pipeline from the sources listed in the text. It was not written or reviewed by a human editor, and it contains no first-hand product testing by us. Where a figure is stated, it comes from our own AI tool database and its verification date is noted. If something here looks wrong, tell us and we will correct or remove it.

People also ask

More in How-to Guides5 more

AI image promptMidjourney prompthow to write AI image promptsMidjourney parametersAI art prompt structure

Want to try this yourself? AI-Mind generates content from a plain description — no prompt engineering required.

Try AI-Mind
← Back to all questions