Skip to main content

Midjourney Image Generator

Midjourney is a Discord- and web-based ai image generator that launched in open beta on July 12, 2022. It transforms clear text prompts into detailed visuals, with a dedicated web interface added in August 2024.

  • Precise, positive prompts work best. Describe what you want to see rather than what you want to avoid, and refine through iteration—perfection comes from multiple cycles, not one perfect command.
  • Four core prompt pillars drive quality results: subject, lighting, viewpoint (camera angle), and aspect ratio. Master these before worrying about advanced features.
  • Parameters like --style raw, --s (stylize), --c (chaos), --w (weirdness), and style references (--sref, --sw) separate decent results from consistent, on-brand visuals.
  • This guide provides concrete, ready-to-copy prompt structures for product photos, marketing visuals, blog images, social posts, and long-form story illustration.

What Is Midjourney and How It Works

Midjourney is an AI system that generates images from natural language prompts. Originally accessed exclusively via Discord, users gained a dedicated web interface in August 2024 with the release of model version 6.1. The platform has evolved through multiple iterations, with v6 trained from scratch for better text interpretation, finer detail control, and improved coherence in complex scenes.

The basic workflow is straightforward: type a prompt using the /imagine command on Discord or directly in the web app, receive a grid of four image options, then upscale your favorites or generate variations. Image generation typically takes 30-60 seconds per batch.

Midjourney doesn’t search the web for existing images. It creates entirely new visuals based on patterns learned during training, which enables creative freedom but raises questions about training data and potential biases.

Example prompt: “a desk setup with clean, minimal design with one laptop and soft lighting”

First Discovery: Choose Precise, Concrete Language

Wording is your primary control mechanism. Vague adjectives like “beautiful” or “cool” give Midjourney too much freedom, while specific descriptors shape structure, texture, and mood directly.

Consider these contrasts:

  • Vague: “beautiful landscape”
  • Precise: “serene alpine lake at sunrise, soft golden hour lighting, low morning fog”
  • Vague: “fast car”
  • Precise: “sleek red sports car racing on a wet highway at night, neon reflections, motion blur”

Avoid vague plurals. “Flowers” opens interpretation to anything, while “twelve flowers arranged in a circle on a stone table, overhead view” enforces numerical precision and viewpoint. This reduces randomness in generation.

Minor wording tweaks create dramatic shifts. Swapping “elegant” for “gritty” or “minimal” for “ornate” changes composition, contrast, and texture because the model associates these words with distinct visual patterns. Test micro-variations—even single adjective changes reveal how sensitively Midjourney responds to your text prompt choices.

Second Discovery: Describe What You Want (and Use –no Sparingly)

Negative phrasing introduces unwanted concepts to the model. Writing “no clutter” or “without people” still primes Midjourney with those ideas, often producing ghostlike traces of the very things you wanted excluded.

Weak: “a desk setup with no clutter” Strong: “a clean, minimal desk setup with one laptop, empty white desk, soft window light, neutral wall”

The --no parameter offers a more reliable exclusion method. Use it for persistent annoyances:

  • --no text — removes unwanted words or labels
  • --no people — keeps scenes empty
  • --no clouds — clears skies

This parameter applies a lightweight exclusion mask during generation, though traces may still appear occasionally. Use it as a supplement to positive descriptions, not as your primary tool.

Social graphics example: “minimalist quote poster, centered typography, solid pastel background, --no photo“

Third Discovery: Add Details Wisely and Keep Prompts Coherent

Stacking too many styles or impossible conditions forces Midjourney to blend conflicting cues, producing muddy results.

Overloaded: “cinematic oil painting ultra realistic 3D render watercolor sketch of a city at night at noon”

Coherent: “cinematic oil painting of a rainy neon city street at night, reflections on wet pavement, dramatic contrast”

A coherent prompt typically focuses on one main subject, one medium, one mood, one time of day, and a consistent art style reference. Before hitting generate, verify your prompt speaks with one voice—no contradictory viewpoints, no conflicting light sources, no impossible combinations.

Iteration beats lengthy single prompts. Professionals refine through 3-5 cycles of adjective tweaks and parameter adjustments rather than cramming everything into one attempt.

Midjourney Prompt Building Cheat Sheet

Consider these seven components for almost every prompt:

  • Subject: Who or what appears (person, animal, object, location)
  • Medium: What form (digital photo, oil painting, illustration, 3D render, charcoal sketch)
  • Environment: Where (indoors, outdoors, underwater, outer space, city street)
  • Lighting: What quality (golden hour, soft candlelight, harsh midday sun, neon glow)
  • Color: What palette (muted pastels, bold primaries, warm tones, grayscale, monochrome)
  • Mood: What feeling (energetic, mysterious, calm, whimsical, eerie)
  • Composition: How framed (portrait, close-up, wide shot, bird’s-eye view)

Complete example touching all seven: “portrait photo of an elderly jazz musician, indoor bar environment, warm golden hour window light, rich amber and teal colors, nostalgic calm mood, close-up composition, shallow depth of field”

Treat this as a mental template. Specify one medium (not three), choose one dominant mood, and let your final output reflect intentional choices rather than accumulated options.

Lighting Is The Fastest Way to Change Mood

In Midjourney, lighting keywords change emotion, depth, and focus more dramatically than almost any other element.

Light Direction:

  • Front light creates even illumination with minimal shadows—clean, low-drama, ideal for products
  • Side light adds shadows, contrast, and texture—dramatic, sculptural, excellent for portraits
  • Backlight produces glow and silhouette effects—atmospheric, cinematic, separates subject from background

Shadow Contrast:

  • Soft shadows come from large, diffused sources (cloudy sky, softbox)—gentle, natural, approachable
  • Sharp shadows come from small, direct sources (direct sun, spotlight)—high drama, crisp definition

Shadow Length:

  • Short shadows from high sun create bright, minimal mood
  • Long shadows from low horizon add depth, direction, and cinematic feel

Same scene, different lighting:

  • “city street on a cloudy day with soft diffused light” — neutral, approachable
  • “city street at golden hour with long shadows and strong rim light” — dramatic, narrative pull

Aspect Ratios and Image Size in Midjourney

Image size in Midjourney primarily means aspect ratio, set via --ar. This is one of the first constraints the model uses to imagine your scene—it shapes composition before any details render.

Classic Ratios
RatioUse Case
1:1Centered products, avatars, grid posts
3:2Classic landscapes, blog headers, web galleries
4:3Traditional photos, emails, presentations
2:3 / 3:4Portraits, posters, editorial covers
Creative Ratios
RatioUse Case
16:9Cinematic banners, YouTube thumbnails, panoramic landscapes 
9:16 / 5:6Reels, TikTok, Stories, vertical ads 
1:2Tall fashion shots, scroll-stopping posters

Examples:

  • “minimalist product photo of a glass skincare bottle, studio lighting, neutral background, --ar 1:1“
  • “panoramic mountain landscape at sunrise, --ar 2:1“

Choose your final usage platform first (web hero, Instagram grid, vertical story) and set aspect ratio in the prompt from the start.

Directing the Camera: Viewpoints and Angles

Viewpoint tells Midjourney where the viewer stands in the scene. Camera angle shapes emotion before lighting or color takes effect.

Main angles:

  • Eye-level — natural, realistic
  • Low angle — powerful, heroic, dramatic
  • High angle — overview, flat-lay feel
  • Wide shot — context, environment, storytelling
  • Close-up — emotion, intimacy, detail
  • Bird’s-eye view — structure, patterns, layouts
Targeted recommendations
SubjectBest Viewpoint
ProductsEye-level or slight top-down
PortraitsEye-level for realism, low angle for impact
ArchitectureLow angle for scale, wide for context
Social contentClose-ups, aesthetic high angles

Example: “bird’s-eye view of a minimalist office desk, overhead flat-lay, soft morning light”

Treat viewpoint as one of your first decisions, not an afterthought added to fix composition problems.

Styling Midjourney Outputs: Style, References, and Parameters

Style determines how an image feels and communicates. It emerges from the interplay of color, texture, contrast, shapes, and detail level—qualities that shape viewer perception beyond subject matter.

Midjourney offers several style systems:

  • Style References via --sref (image URL) and --sw (strength 0-1000) let you mimic a visual language or moodboard across multiple images
  • Omni Reference (--oref with --ow weight) maintains character and object consistency across scenes—crucial for series and stories

Define 1-2 visual styles per project (e.g., “soft pastel editorial illustration” vs “high-contrast cinematic photo”) and stick with them. Consistency builds brand recognition across generated images.

Core Midjourney Parameters: Realism, Stylization, Chaos, Weirdness

Parameters don’t change scene content—they change how Midjourney interprets and renders your prompt.

--style raw (or --raw): Reduces Midjourney’s default stylization for more literal interpretation. Ideal for commercial photography, product shots, and clean branding where artistic drift hurts more than helps.

--s (stylize): Default is 100. Lower values (0-50) produce more literal, realistic images generated with minimal artistic interpretation. Higher values (200-1000) create artistic, expressive results—but extreme values may drift from your prompt.

--c (chaos): Low values (0-10) keep outputs predictable and similar. Higher values introduce varied compositions—useful for marketing visuals or exploring concepts when you want surprise.

--w (weirdness): Increases experimental, unexpected interpretations. Powerful for metaphorical marketing visuals, but use carefully—weirdness should support your message, not confuse it.

Designing Strong Product Images in Midjourney

Product images require instant readability, premium feel, and zero distractions—similar to real studio photography.

Key principles:

  • One clear subject, minimal background
  • Lighting that reveals form and material (glass, metal, matte plastic)
  • Centered or deliberately offset composition, never cluttered
  • Low stylization and chaos for commercial clarity

Prompt formula: “minimalist product photo of a glass skincare bottle, placed on a clean white surface, soft diffused studio lighting, neutral background, professional studio photography, --ar 1:1 --style raw --s 50 --c 5 --no text --no props“

Lighting variations:

  • “dramatic softbox lighting from the left” — adds dimension, premium feel
  • “soft side lighting” — reveals texture, reduces harshness

Imagine a real-world studio setup first (light position, camera distance, background material) and translate that mental scene into your prompt. Commands like “centered composition” and exclusions like --no props keep focus tight.

Creating Bold Marketing Visuals That Stand Out

Marketing visuals prioritize emotional impact and stopping power over strict realism. In a noisy feed, your image has seconds to communicate value.

Core traits:

  • Strong contrast and directional lighting
  • Dynamic composition (off-center subjects, diagonals, unusual angles)
  • Clear metaphor or focal idea linked to the brand
  • Moderate chaos and weirdness for originality

Prompt template: “dramatic marketing visual of a steaming coffee cup rising like a sunrise over a city skyline, strong rim light, high contrast, warm orange and deep blue colors, low angle, --ar 16:9 --s 200 --c 25 --w 15“

Increasing --c encourages fresh compositions while controlled --w produces memorable metaphors without derailing the concept. Test 3-4 variations per concept—change angle, lighting, or color palette before locking in your campaign image.

Images for Blogs and Social Posts: Visuals That Carry the Story

Editorial and social images support reading rather than overshadowing text. They work best when visualizing metaphors instead of literally depicting headlines.

Instead of prompting “work-life balance,” describe a physical scene embodying the idea: “two contrasting workspaces, one chaotic and one calm, split frame composition.”

Examples:

  • “minimalist illustration of two contrasting city halves, one in cool blue night colors, one in warm morning light, clean vector style, --ar 3:2“
  • “overhead photo of a desk divided diagonally, one side cluttered, one side clean with a single laptop and notebook, soft daylight, --style raw“

Match aspect ratios to channels: 3:2 or 4:3 for blog headers, 1:1 for feed posts, 9:16 for Stories. Show small human details—crumbs, tools, steam, reflections—to make images feel lived-in rather than generic.

Local Social Content and Stories: Showing How It Feels to Be There

For local businesses—cafés, gyms, studios, salons—the goal is communicating atmosphere and routine rather than products alone. Show what it feels like to be there.

Focus prompts on specific moments:

  • Early mornings before opening
  • Quiet breaks between customers
  • Tools and textures (flour dust, steam, stacked chairs)
  • Warm interior lights against cool exterior dawn

Examples:

  • “cozy neighborhood bakery at 6am, baker’s hands dusted with flour shaping bread, warm golden light from the oven, shallow depth of field, --ar 9:16“
  • “barista preparing espresso before opening, empty café, chairs upside down on tables, cool blue dawn light outside, warm interior lights, --style raw“

Stories and vertical content benefit from close-ups and human-scale viewpoints. Reuse successful prompts with small tweaks (time of day, season, props) to build visual narrative over weeks without starting from scratch each time.

Building Visual Worlds and Stories with Omni Reference

Long-form illustration—novellas, comics, ongoing content series—requires consistency across many images generated over time. Characters and world must feel recognizable from scene to scene.

Omni Reference workflow:

  1. Generate an anchor portrait defining your protagonist with specific style and detail level
  2. Use that image as omni reference when prompting new scenes via --ow (Omni weight, default 100, range 0-1000)
  3. As scenes become more expressive (stronger lighting, higher chaos), increase --ow to protect character identity

Build a “prompt backbone” you keep reusing: same style language, character reference, detail level. Only vary mood, viewpoint, and environment from chapter to chapter. When readers recognize a character before consciously thinking about it, your system is working.

Iteration and Refinement: Turning First Drafts into Final Images

Professional-looking outputs come from several prompt iterations, not single perfect commands.

Simple iterative loop:

  1. First run: Broad concept with basic subject, lighting, aspect ratio
  2. Second run: Refine adjectives, specify vantage point, add --no for unwanted clutter
  3. Third run: Adjust stylization, chaos, or weirdness to fine tune realism or experimentation

Save strong results as style or character references for future prompts. Even small word changes (“soft light” vs “harsh light,” “muted palette” vs “vibrant colors”) are worth testing.

Iteration isn’t failure—it’s the core creative process when working with any ai image generator.

Practical Midjourney Prompt Examples You Can Reuse

Copy these templates and modify one element at a time to build confidence:

Product photo: “minimalist product photo of a ceramic coffee mug, placed on a clean marble surface, soft diffused studio lighting, neutral background, professional studio photography, --ar 1:1 --style raw --s 50 --c 5 --no text --no props“

Marketing visual: “dramatic marketing visual of a runner breaking through morning fog at sunrise, strong rim light, high contrast, warm orange and deep purple colors, low angle, --ar 16:9 --s 200 --c 25 --w 15“

Editorial/blog metaphor: “overhead photo of a desk divided diagonally, one side cluttered with papers and coffee cups, one side clean with a single laptop, soft window daylight, --ar 3:2 --style raw“

Local business atmosphere: “cozy coffee shop at 6:30am, barista arranging pastries in display case, warm interior lights against blue dawn outside, shallow depth of field, --ar 9:16“

Story/character illustration: “fantasy portrait of a young alchemist with silver hair and amber eyes, warm candlelit laboratory, glass vials glowing green, painterly style, --ar 2:3 --s 150“ (use as anchor, then reference with --oref for subsequent scenes)

Treat these as starting points. Modify lighting, angle, or aspect ratio one element at a time.

FAQ

These questions cover practical workflow concerns not fully addressed in the main sections.

How do I choose between the Discord bot and the Midjourney web interface?

Discord remains powerful for quick experimentation, community interaction, and traditional /imagine workflows where you want to explore ideas rapidly among millions of other users’ prompts. The web interface centralizes editing tools—panning, zooming, inpainting—in a visual workspace better suited for detailed refinement.

Beginners should start on the web for clarity, then adopt Discord for fast prompting or collaboration if needed. Both platforms sync via linked accounts, so images and settings stay consistent.

How many words should a good Midjourney prompt have?

There’s no strict limit, but 15-40 focused words usually work best—enough to cover subject, medium, lighting, color, mood, and composition without introducing contradictions. If your prompt exceeds 2-3 short sentences, consider whether it contains redundant or conflicting details.

Split complex ideas into multiple images rather than overloading one prompt. Precision and coherence matter more than length.

Can I use Midjourney images commercially?

Midjourney’s terms of service govern commercial use. Paid subscriptions generally allow commercial applications, but always review the most recent license terms on the official website before using images in ads, packaging, or products.

Policies change over time, and use involving trademarks, likenesses, or sensitive themes may require additional review. This guide focuses on creation and control rather than legal advice.

How do I make Midjourney follow a brand’s visual identity?

Define a brand “style recipe” first—color palette, contrast level, texture preferences, typical lighting—and bake those consistently into every prompt. Use style references (--sref and --sw) for visual consistency.

Create an internal library of 5-10 “approved look” images and reuse them as visual anchors. For recurring mascots or characters, Omni references with adjusting --ow values help maintain identity across variations.

What should I do if Midjourney keeps adding unwanted objects or text?

First, flip your language to positive specificity (“empty white background” instead of “no background”). Then add --no parameters for persistent issues: --no text, --no logos, --no extra people.

Lower chaos and stylization slightly (--c and --s) to reduce unexpected embellishments. Some artifacts may still appear—inpainting or simple post-editing in external software like Adobe Firefly can handle final cleanup faster than endless regeneration.

Changed

Vision Newsletter

Subscribe

* indicates required
Languaje *
Choose the languaje for the newsletter.