How to Build a Custom Reverse-Prompt Gem in Google Gemini ⚡
Stop begging for prompts in comment sections. Learn how to configure a custom Google Gemini “Gem” that deconstructs any trending visual, viral ad, or cinematic photo into exact, copy-paste prompts for Midjourney v6.1, FLUX.1, and DALL-E 3.
⚡ TL;DR
The Gemini Reverse-Prompt Gem is a custom AI agent configured inside Google Gemini that functions as an automated image-to-prompt translator. Instead of manual guessing, you upload any image and the Gem uses Gemini’s native multimodal vision engine to deconstruct lighting, focal optics, camera angles, color grading, and artistic textures.
- No More Guessing: Converts raw image inputs into clean, production-ready prompts tailored specifically for Midjourney v6.1, FLUX.1 [dev], and DALL-E 3.
- Multi-Dimensional Analysis: Automatically identifies lighting setups (Kelvin temp, rim lighting, softboxes), lenses (35mm, 85mm f/1.4, macro), and rendering engines (Octane, Unreal Engine 5, 35mm film grain).
- Permanent Setup: Created once in the Gemini “Gems” menu and accessible across all your devices with zero coding.
Why Google Gemini is the Ultimate Reverse-Prompt Engine
While models like GPT-4o have vision capabilities, Google Gemini (powered by Gemini 1.5 Pro and 2.0 Flash) is uniquely suited for reverse prompt engineering due to its native multimodal foundation:
🎯 Sub-Pixel Spatial Awareness
Gemini decodes exact spatial relationships, depth of field, perspective vanishing points, and compositional grids with unmatched precision.
📸 Photographic Vocabulary
Understands camera bodies (Hasselblad H6D, Sony A7R V), sensor sizes, anamorphic lens flares, and film stocks (Kodak Portra 400, Cinestill 800T).
⚡ Sub-Second Turnaround
Processes high-resolution 4K images and outputs multi-model prompt variants in under 4 seconds without timeout errors.
Step-by-Step Setup: Building Your Gem in 60 Seconds
Creating your personal Reverse-Prompt Gem takes less than a minute inside Google Gemini. Follow these 4 exact steps:
Step 1: Open Google Gemini & Navigate to “Gems”
Go to gemini.google.com. In the left navigation sidebar, click on Gem manager (or click Explore Gems). Then click the blue “New Gem” button.
Step 2: Name Your Custom Gem
Set the name to “Reverse Prompt Architect” (or “Visual Deconstructor Pro”). Add a brief description such as: “Upload any image to receive exact Midjourney, FLUX, and DALL-E prompt formulas.”
Step 3: Paste the Master System Instruction
Copy the complete system prompt from the box below and paste it directly into the “Instructions” field. This prompt programs Gemini to behave as an elite Director of Photography, lighting supervisor, and AI prompt engineer.
Step 4: Save & Upload Any Image
Click “Save” in the bottom right corner. Your Gem will now appear in your pinned sidebar. Whenever you see a viral photo or render, simply drag-and-drop the image into this Gem to get the full prompt breakdown instantly.
The Master System Instruction (Copy & Paste)
Here is the exact prompt engineering payload to paste into your Gem’s instruction box. It forces Gemini to systematically analyze visual elements and output syntax tailored to the latest diffusion engines:
Visual Deconstructor & Reverse-Prompt Engineer
This prompt instructs Gemini to avoid superficial commentary and immediately return structured technical breakdowns and engine-specific copy-paste prompts.
You are the world's foremost Director of Photography, CGI Supervisor, and Reverse AI Prompt Engineer. YOUR PURPOSE: Whenever the user uploads an image, photo, digital artwork, or UI mockup, your sole job is to reverse-engineer its visual DNA and reconstruct the exact prompt syntax required to recreate or stylistically match that image across modern AI generators (Midjourney v6.1, FLUX.1, and DALL-E 3). EXECUTION PROTOCOL: When an image is provided, immediately generate a structured response following this exact 5-section schema: ### 1. VISUAL DNA BREAKDOWN - **Subject & Focus:** Identify the primary subject, garment/material textures, facial expressions, micro-details, and surrounding props. - **Lighting Architecture:** Identify the key light, fill light, rim light, ambient Kelvin temperature (e.g., warm 3200K vs cool 6500K), volumetric haze, and shadow hardness. - **Photographic & Camera Optics:** Estimate camera type (e.g., 35mm film, Hasselblad medium format), lens focal length (e.g., 24mm wide, 85mm f/1.4 portrait, macro), aperture, depth of field (shallow bokeh vs deep focus), and shutter speed. - **Art Medium / Engine:** Identify the medium (e.g., authentic 35mm Kodak Portra 400 film grain, Unreal Engine 5 render, Octane 3D, matte oil painting, pastel claymation, or vector flat art). - **Color Grading & Mood:** Palette colors (hex codes or tonal ranges), contrast level, saturation, and overall emotional atmosphere. --- ### 2. MIDJOURNEY v6.1 MASTER PROMPT Provide a highly optimized Midjourney v6.1 prompt formatted in standard Midjourney syntax: ```text [Subject description with rich micro-details], [camera angle and framing], [lighting conditions and Kelvin temperature], [environment and atmospheric particles], [photographic medium or render engine], [color grading style] --ar [aspect-ratio] --v 6.1 --style raw --stylize [value] ``` --- ### 3. FLUX.1 [DEV / SCHNELL] MASTER PROMPT Provide a natural language descriptive prompt optimized for the FLUX diffusion transformer architecture (dense, sentence-based narrative description without comma stuffing): ```text [A detailed, continuous 3-to-4 sentence narrative description focusing on photorealistic textures, exact physical interactions, lighting source placement, and environmental realism]. ``` --- ### 4. DALL-E 3 / CHATGPT MASTER PROMPT Provide a high-coherence, visually distinct prompt tailored for DALL-E 3 that enforces strict composition, avoids text glitches, and maintains photorealism: ```text [Photorealistic shot of ... with sharp focus on ... captured using an 85mm f/1.4 lens, natural daylight pouring in from the side, rich micro-textures, cinematic color grading]. ``` --- ### 5. PROMPT ITERATION CONTROLS (TWEAKABLE VARIABLES) List 3 to 4 modular tags or parameters the user can swap to customize the output (e.g., change lighting from Golden Hour to Cyberpunk Neon, switch lens from 85mm portrait to 14mm ultrawide, or replace subject). OPERATING RULES: - Never say "I can't generate images" or give conversational apologies. - Never output vague phrases like "beautiful lighting" or "high quality". Use specific technical terminology: "chiaroscuro lighting", "subsurface scattering", "crepuscular rays", "chromatic aberration", "anamorphic bokeh". - If the user asks for a prompt variation, adapt the existing deconstructed DNA accordingly.
How the 5-Dimension Deconstruction Framework Works
Standard AI prompts fail because they use generic adjectives like “hyperrealistic, 8K, highly detailed”—words that modern diffusion models like FLUX and Midjourney ignore. The Reverse-Prompt Gem forces the model to extract concrete photographic parameters:
| Dimension | What the Gem Analyzes | Why It Matters in Diffusion Models |
|---|---|---|
| Lighting Architecture | Direction, hardness, color temperature (3200K vs 5600K), rim lights, bounce fill | Controls contrast, drama, and material reflections across surfaces. |
| Optics & Focal Length | 24mm wide angle vs 50mm natural vs 85mm portrait vs 100mm macro | Dictates background compression, perspective distortion, and bokeh depth. |
| Material Textures | Subsurface scattering on skin, brushed metal, woven linen, wet asphalt | Eliminates the “plastic AI skin” look and produces realistic tactile quality. |
| Medium / Film Stock | Kodak Portra 400 grain, Fujifilm Velvia saturation, Octane 3D render, Claymation | Establishes the rendering engine or physical substrate of the visual. |
| Engine Parameters | Midjourney --ar, --stylize, --style raw |
Calibrates model adherence to the prompt vs default stylized aesthetics. |
Real-World Testing: Deconstructing Viral Aesthetics
Here are three common visual categories and how the Gemini Reverse-Prompt Gem accurately translates them into working prompts:
Case 1: Luxury Dark-Mode Perfume Commercial Shot
Input Image: A minimalist glass cologne bottle resting on wet slate with dynamic water splashes and moody amber backlighting.
Commercial studio product photography of a heavy luxury matte black glass cologne bottle resting on wet obsidian slate. Subtle high-speed water splash droplets frozen in mid-air around the base. Moody low-key chiaroscuro lighting, intense warm amber back-rim lighting outlining the bottle silhouette, delicate softbox fill on the embossed silver typography. Macro 90mm lens, f/2.8, crisp refraction through water droplets, Hasselblad H6D-100c capture --ar 4:5 --v 6.1 --style raw
Case 2: Cyberpunk Neon Rain Street Portrait
Input Image: A portrait of a subject holding a clear umbrella under magenta and cyan neon store signs on a rain-slicked Tokyo street.
A cinematic medium close-up portrait of a woman under a transparent vinyl umbrella on a rain-drenched Shinjuku street at night. Vivid dual-tone lighting with neon magenta rim light on the left cheek and cool cyan street signage reflection on wet shoulders. Droplets of rain dripping from the umbrella edge, shallow depth of field with creamy circular bokeh, shot on 35mm film, subtle grain, Cinestill 800T color grading --ar 16:9 --v 6.1 --stylize 250
Case 3: 3D Minimalist Tech App Mockup
Input Image: Floating isometric clay-textured mobile screen UI with soft pastel gradients and smooth ambient occlusion shadows.
Isometric 3D render of a floating glassmorphism smartphone showcasing a sleek financial analytics dashboard. Frosted glass panels, soft pastel lilac and mint green color palette, velvety matte clay material, diffused studio global illumination, clean contact shadows, Octane render, C4D aesthetic, hyper-minimalist composition --ar 1:1 --v 6.1
Frequently Asked Questions
What is the difference between a Gemini Gem and a ChatGPT Custom GPT?
Both allow persistent system instructions and specialized capabilities. However, Gemini Gems leverage Google’s native 1M+ token context window and superior multimodal vision engine, making them faster and more perceptive when analyzing fine photographic details and lighting nuances.
Can I use this Gem on mobile inside the Gemini Android / iOS app?
Yes! Once you create and save the Gem on your desktop browser, it immediately synchronizes with your Google account. You can open the official Google Gemini app on Android or iOS, tap on your Gem, take a photo with your phone camera, and receive the prompt breakdown instantly on the go.
How does this prompt work with FLUX.1 vs Midjourney?
Midjourney prefers keyword-dense syntax with specific parameter flags (like --ar 16:9 and --style raw). FLUX.1 uses a modern diffusion transformer architecture that responds best to natural, flowing descriptive prose. The system instruction specifically outputs separate, properly formatted prompts for both engines.
Can I upload multiple reference images at once?
Yes. You can upload 2 or 3 images simultaneously and ask the Gem: “Combine the lighting from Image 1 with the subject from Image 2 and the camera angle of Image 3 into a single unified prompt.” Gemini will synthesize the visual DNA from all three sources into one cohesive prompt.
Can I use these prompts in AI video generators?
Yes. To convert any of the image prompts into a video generation prompt (for Google Flow, Runway Gen-3, Luma Dream Machine, or Kling AI), simply prepend a camera movement instruction (such as “Slow cinematic camera dolly push forward toward the subject” or “Dynamic 360-degree orbital pan at 24fps”).