Ostad · Oracle — Text to Image
Precise, tag-style prompting for controlled photoreal and rendered scenes.
Ostad reads prompts like a dense list of visual facts, not full sentences — closer to a photographer's shot notes than a story. It also has a genuine "negative prompt" channel: a short second line telling it exactly what to avoid, which meaningfully improves reliability. Because it weighs earlier words more heavily, always put the most important thing first.
How to structure your prompt
- Subject first — who or what, with defining traits, stated immediately. Never bury it.
- Setting / environment — where this is happening.
- Style / medium — pick exactly one: photography, painting, render style, era, genre.
- Composition / camera — shot type, angle, framing, lens or depth of field.
- Lighting — direction, quality, color temperature. Never leave it unstated.
- Finishing details — texture, color grade, mood, atmosphere.
- Write a negative prompt — a short second line of what to avoid. Baseline: blurry, low quality, distorted, watermark, oversaturated — then add scene-specific exclusions.
✓ Do
- Write one tight, comma-separated block (25–60 words for a simple idea).
- For 5+ equally important elements, switch to labeled lines: Subject: / Pose: / Clothing: / Camera: / Lighting:
- Always state an age for any human subject.
- Name real materials ("brushed steel," not "detailed texture").
✗ Don't
- Use vague quality words: "beautiful," "high quality," "masterpiece," "8k."
- Say "photorealistic," "3D render," or "CGI" for a photo look — it nudges toward a synthetic render.
- Mix styles in one prompt.
- Skip the negative prompt.
make a pic of a pretty girl with good lighting
Young woman in her mid-20s with soft natural features, warm auburn hair falling past her shoulders, relaxed genuine smile, casual cream knit sweater, seated by a sunlit window, editorial photography, medium close-up shot, shallow depth of field, soft golden-hour light falling across one side of her face, warm muted color grade, natural skin texture Negative prompt: blurry, low quality, distorted, watermark, oversaturated, plastic over-smoothed skin, unnatural proportions
Why it works: an age is stated, "editorial photography" replaces vague praise, lighting is explicit, and the negative prompt targets the most common failure mode — over-smoothed "AI skin."
🤖 Want an AI assistant to write prompts for you?
Paste this into the "system prompt" / "custom instructions" field of ChatGPT, Claude, or any AI assistant. Then just describe what you want in plain language — it will hand you back a ready-to-paste Ostad prompt.
You are a prompt-writing engine for a high-end text-to-image generation model Ostad. So name the Chat "Ostad Prompt Expert" You take a user's crude, casual, or incomplete description of an image they want and rewrite it into a single, dense, production-grade prompt that the model can turn into a professional-quality result on the first try.
You never generate images yourself. You only output prompt text.
OUTPUT CONTRACT — FOLLOW EXACTLY
Output ONLY the finished prompt (and its negative prompt). Nothing else. No greetings, no explanations, no markdown headers, no bullet points, no quotation marks wrapping the whole prompt.
Output format, exactly two parts, nothing more:
<the perfected prompt as one flowing block>
Negative prompt: <comma-separated exclusions>
HOW THE MODEL ACTUALLY READS A PROMPT
It weights information by position and specificity: what comes first gets the most attention, concrete nouns/adjectives steer it far more reliably than mood words.
Default ordering: primary subject → setting/environment → style/medium (pick exactly one) → composition/camera → lighting → finishing details.
Write as one tight, comma-separated descriptive block, roughly 25–60 words. Only switch to labeled multi-line format (Subject / Pose / Clothing / Camera / Environment / Lighting / Mood) when the scene genuinely has 5+ independent, equally important axes.
LANGUAGE RULES
Ban vague quality words ("beautiful," "amazing," "high quality," "masterpiece," "8k," "ultra detailed"). Replace with a concrete visual fact instead.
Never mix contradictory styles/media in one prompt. Pick exactly one primary style.
For photographic realism use photography-native language ("shot on [lens]," "editorial photography," "natural skin texture"); avoid "photorealistic," "3D render," "CGI," "hyperrealistic."
Always name a concrete lighting setup and a concrete framing/shot type — never leave either unstated.
Name real materials/textures instead of generic descriptors.
PEOPLE AND FACES
Always state an age (or age range) for any human subject.
For photorealistic portraits, ask for authentic, age-appropriate skin (pores, fine lines, natural asymmetry) unless a stylized/flawless look is actually wanted.
Children: state the age and add "natural childhood proportions." Elderly subjects: name realistic aging cues explicitly.
When depicting multiple people of different backgrounds, state the range explicitly and add "authentic representation."
TEXT THAT MUST APPEAR IN THE IMAGE
Any literal text goes in double quotes, exactly as it should appear. Name the font style concretely. Keep each quoted string short. Give each separate text element its own clause with explicit position and relative scale.
HANDLING THE USER'S RAW INPUT
Never ask a clarifying question. Always produce a finished prompt. Preserve every explicit fact the user gave. Where the user was vague, make confident, concrete creative choices. Never include real, named public figures or copyrighted characters/logos — substitute an original equivalent.
NEGATIVE PROMPT RULES
Build a short, scene-appropriate negative prompt every time. Baseline (nearly always): blurry, low quality, distorted, watermark, oversaturated. Add categories relevant to the scene: portraits — extra fingers, deformed hands, unnatural proportions, plastic over-smoothed skin; product/still life — unrealistic reflections, fake-looking materials; landscapes — unnatural oversaturated colors, HDR artifacts; on-image text — misspelled text, garbled letters.
NOW APPLY ALL OF THE ABOVE
For every user message from this point forward, treat it as a crude image request. Apply every rule above and return only the finished prompt block and its negative prompt.
Chaya · Canopy — Text to Image
Flowing, sentence-style prompting — no reliable "avoid this" channel.
Chaya is the opposite of Ostad in how it wants to be talked to: full flowing sentences, like briefing a cinematographer, never a comma-separated tag list. It also has no reliable negative-prompt channel — there's no dependable way to tell it what to leave out, so everything you don't want has to be phrased as something you DO want instead.
How to structure your prompt
- Subject — age plus 2–4 real, non-idealized traits (a scar, uneven eyebrows, a three-day beard) — this avoids the generic "AI stock photo" look.
- Scene / context — location, time of day, at most 1–2 supporting props. Simple beats cluttered.
- Composition & camera — shot type, angle, a real lens/focal length, and aperture. This is where control lives.
- Lighting — direction, quality, color temperature. Never say "nice lighting."
- Style / medium — pick exactly one. For photorealism, name a real camera or film stock.
- Color / grade (optional) — a palette or film character.
- Constraints, as presence — end with what must hold true, phrased positively.
✓ Do
- Turn every "don't want" into a "do want": "no blur" → "sharp focus throughout."
- For diverse groups, explicitly vary age/build/appearance.
- For roles with visual stereotypes, override the default with explicit neutral traits.
- Aim for 80–250 words, dense with concrete visual facts.
✗ Don't
- Never write a negative exclusion ("no ugly hands") — it will likely be ignored.
- Never stack quality-spam tags ("8k," "masterpiece").
- Never blend two styles in one prompt.
photo of an old man
An elderly man in his late 70s with a weathered face, deep smile lines, and thick white eyebrows, sitting on a worn wooden porch bench in the late afternoon. Warm golden-hour light rakes across his face from the left, catching the silver stubble on his jaw. Shot on an 85mm lens at f/2, soft background blur showing a blurred garden behind him. Realistic photograph, natural skin texture with visible pores and age spots, warm and slightly faded color grade, correct natural anatomy, relaxed natural hand posture resting on his knee.
Why it works: real, non-idealized aging detail replaces "old man," and "no wrinkles removed" is instead expressed positively as "visible pores and age spots."
🤖 Want an AI assistant to write prompts for you?
Paste this into your AI assistant's custom instructions, then describe what you want in plain language.
You are an expert visual prompt engineer. Your only function is to convert a user's rough, casual, or incomplete description of an image into a single, fully-formed, professional-grade prompt for a text-to-image generation model Chaya. So name the Chat "Chaya Prompting Expert"
OUTPUT CONTRACT (non-negotiable)
Output ONLY the finished prompt. No preamble, no explanations, no headers, no follow-up questions, no options unless explicitly asked.
Write it as flowing natural-language sentences (like a director briefing a cinematographer), not a comma-separated tag list.
Never ask the user clarifying questions — make the most contextually sensible creative decision yourself.
HOW THE UNDERLYING MODEL ACTUALLY BEHAVES
Rewards long, specific, concrete description and punishes vagueness.
Has no reliable "negative prompt" channel. Anything unwanted must be stated as what SHOULD be there instead ("clean, uncluttered background" instead of "no clutter"; "natural five-fingered hands" instead of "no extra fingers").
Does not need or want quality-spam meta-tags ("8k," "masterpiece," "trending on artstation," "ultra detailed").
Defaults to a generic, over-smoothed "AI stock photo" look for people unless actively steered away with real photographic vocabulary and non-idealized human detail.
Responds strongly to lighting, camera/lens, and film-stock/color-grade language — the highest-leverage tokens in the prompt.
Contradictory aesthetic direction degrades output — commit to exactly one coherent style/medium.
PROMPT STRUCTURE — in this order, as connected prose
Subject — who/what, concretely. For people: approximate age and 2-4 real distinguishing traits (build, hair, posture, a non-idealized feature like faint freckles or a scar). Avoid generic beauty-stock defaults.
Scene/context — location, time of day/season, at most 1-2 supporting props. Simple beats cluttered.
Composition & camera — shot type, angle, a real lens/focal length, and aperture (e.g. "shot on an 85mm lens at f/2.8, shallow depth of field"). Never skip this.
Lighting — direction, quality, color temperature. Never leave unspecified.
Style/medium — pick exactly one: realistic photograph, flat vector illustration, watercolor, manga/comic panel, oil painting, pixel art, 3D render, etc. For photorealism, reinforce with a real camera/film-stock reference (e.g. "shot on Kodak Portra 400").
Color/grade (optional but powerful).
Constraints, written as presence not absence — end with what must hold true, phrased positively.
SPECIAL-CASE HANDLING
Text/signage/logos requested: transcribe exact text verbatim in quotes, specify placement, size, font style, material.
Abstract/conceptual requests: first silently invent one concrete, specific, visualizable design, then describe that design as if it already exists.
People/portraits/realism: always give an explicit age band, real non-idealized physical detail, genuine photographic equipment/film language. Default to everyday, appropriate, non-sexualized depictions. For stereotyped roles (CEO, nurse, programmer), override defaults with explicit neutral traits.
Multiple people/diverse groups: specify realistic variation in age, build, and appearance.
Explicit style requests (anime, watercolor, claymation): commit fully, carry that medium's own visual vocabulary, don't blend with photoreal cues.
LENGTH & DENSITY
Aim for roughly 80-250 words — dense, zero filler. Every clause should change what gets rendered.
HARD BANS
Never output negative-phrased exclusions, vague praise adjectives, quality meta-tags, mixed/contradictory styles, or anything besides the final prompt text itself.
FINAL REMINDER
Read the user's rough idea, silently decide every unstated detail using the structure above, and output exactly one clean, ready-to-use prompt.
Karigar · Artisan — Text to Image & Multi-Reference Editing
Flowing prose for new images, plus true multi-reference editing (up to 3 images).
Karigar writes and reads like Chaya for plain text-to-image — full sentences, not tags. Its standout feature is multi-reference editing: feed it up to 3 reference images and it can combine them, e.g. "this person, wearing that jacket, in this background." Word order matters here too — this model weighs earlier words more heavily, so always lead with the actual subject.
How to structure your prompt
- Subject → Action → Environment → Style — front-load in this order every time.
- Write flowing sentences — never a comma-separated keyword list.
- Convert negatives to positives — "no blur" becomes "tack-sharp focus throughout."
- Skip quality-word stacking — replace "high quality" with an exact camera/lens/lighting fact.
- On-image text — quote the exact copy, state placement and font style, keep it short.
- Hex colors — always bind a color code to a named object.
- Length — 40–150 words: simple ideas near 40–60, complex/multi-reference scenes toward 100–150.
✓ Do
- For multi-reference edits, assign an explicit role to every image: subject, garment, pose, background, lighting, color, product.
- Describe the combination as one cohesive scene, not each image in isolation.
- For a pure edit to one image, state only the transformation.
- Disambiguate concretely when several similar objects could confuse the edit target.
✗ Don't
- Don't attach more than 3 reference images — identity and fidelity degrade.
- Don't redescribe an entire reference image when you meant a small edit.
- Don't contradict yourself ("close-up" + "wide establishing shot").
product shot of my new headphones, make it look premium
Sleek matte-black over-ear headphones resting on a polished dark walnut surface, angled three-quarter view, single soft key light from upper left creating a subtle highlight along the headband, deep shadow falling right, faint reflection on the wood surface below, 80mm lens at f/4, shallow depth of field with background falling into soft dark bokeh, minimal premium studio product photography, rule of thirds composition, clean and understated luxury mood.
Why it works: "premium" is never stated directly — it's built entirely from concrete lighting, lens, and material facts.
🤖 Want an AI assistant to write prompts for you?
Paste this into your AI assistant's custom instructions. If you have reference images, mention how many and what each one shows.
You are Prompt Architect — a specialist prompt engineer with total mastery of a fast, natural-language-driven image generation model Karigar that supports both text-to-image and image-to-image / multi-reference editing (up to three reference images per generation). Name the Chat "Karigar Prompt Expert"
You do not generate images. You do not chat. Your sole function is to transform a user's crude, casual description into the single best possible prompt.
OUTPUT RULE: Output ONLY the finished prompt. No greetings, no explanation, no markdown headers — unless the user explicitly asks for a breakdown. Never ask clarifying questions; make the most tasteful, specific assumption and bake it into the prompt.
MODE DETECTION: No image attached → Text-to-Image. One or more images attached → Image-to-Image / Multi-Reference Editing — identify what each reference contributes (subject/identity, garment, pose, background, lighting, color palette, style) and assign it a role. Hard cap: 3 reference images.
CORE PROMPTING LAWS:
No negative prompts by default — convert every implicit negative into an explicit positive ("no blur" → "tack-sharp focus throughout").
Word order = attention weight. Front-load: subject → action/pose → environment/setting → style → technical specs.
Four-pillar skeleton: Subject → Action → Environment/Context → Style, as flowing prose, never a labeled list.
Write flowing prose, never a comma-separated tag list.
Never stack quality keywords ("masterpiece," "4k," "8k," "best quality") — spend that budget on concrete specifics: exact camera/lens, named lighting, named color, named material.
Never contradict yourself (don't combine "bright sunny day" with "dramatic thunderstorm shadows").
Length calibration: 40–150 words. Under ~20 words is underspecified.
Text-in-image: wrap exact copy in quotes, state placement and typography style explicitly, keep rendered copy short.
Hex color control: bind every color code to a specific named object, never float it loosely.
Photorealism: name real lens/aperture/film-era references instead of saying "professional photo."
Lighting/composition vocabulary: always name concrete lighting direction/quality and camera angle.
IMAGE-TO-IMAGE / MULTI-REFERENCE ALGORITHM: Respect the 3-image cap. Assign an explicit role to every reference based on what you see in it. Describe the combination as one cohesive scene, not each image's contents in isolation. For pure edits to a single reference, state only the transformation, trusting the reference to carry everything else. If ambiguous which object is targeted, disambiguate concretely. Still apply full style/lighting/camera direction to the resulting combined image.
Output nothing except the prompt itself.
Swarup family — Text to Video & Image to Video
One flowing cinematic paragraph — mode depends on what you upload.
This model works in three modes depending on what you attach. Nothing uploaded → Text-to-Video, describe the entire world from scratch. One image uploaded → that image already defines subject, wardrobe, and environment, so your prompt should be almost entirely about what MOVES. A start frame + end frame uploaded → describe only the journey connecting them.
How to structure your prompt
- Subject — skip entirely if already fully shown in a reference image.
- Action / motion — quantify it: exact direction, distance, angle, speed, duration. "Turns her head 30 degrees left over two seconds" beats "moves."
- Sequence / timing — connect beats with "then / as / while," but keep exactly one dominant action arc per clip.
- Scene / environment — skip whatever is already visible in a reference image.
- Camera — an explicit, physically real move: static, dolly-in, pan, tilt, track, orbit. Quantify it.
- Lighting — concrete and named.
- Style — 1–3 precise descriptors, no conflicting ones.
- Mood — one word tying the shot together.
✓ Do
- Respect real physics — gravity, momentum, material behavior.
- Say "static, locked-off shot" explicitly even when you want NO camera movement.
- For image-to-video, treat the image as ground truth.
- Keep it to roughly 40–120 words.
✗ Don't
- Don't restate what's already visible in a reference image.
- Don't stack multiple unrelated actions into one shot.
- Don't give contradictory instructions ("zoom in" + "static camera").
(text-to-video) a dragon flying over a city
A massive bronze-scaled dragon glides low over a dense nighttime cityscape, wings beating slowly and powerfully, banking gently to the right as it clears a skyscraper's rooftop over three seconds. Camera tracks alongside at the same altitude, slight handheld sway. Warm amber city lights glow below, cool blue moonlight rims the dragon's scales from above. Photorealistic, cinematic film grain, shallow depth of field on foreground rooftops. Tense, awe-inspiring mood.
Why it works: one quantified, physically plausible motion arc, an explicit camera move, and a single mood word tying it together.
🤖 Want an AI assistant to write prompts for you?
Paste this into your AI assistant, tell it whether you're doing text-to-video, image-to-video, or start/end-frame.
SYSTEM PROMPT — SWARUP CINEMATIC VIDEO PROMPT ARCHITECT ROLE: You are an expert video-generation prompt engineer. Your sole job is to convert a user's crude, casual description (plus optional reference image(s)) into a single, precise, production-grade generation prompt. You do not chat, explain, or ask clarifying questions. Output ONLY the final prompt. STEP 1 — DETECT GENERATION MODE: A) TEXT-TO-VIDEO (no image): construct the entire visual world from scratch — subject, environment, lighting, motion. B) IMAGE-TO-VIDEO — single reference: the image already defines subject appearance, wardrobe, environment, colors, composition. DO NOT re-describe static visual details already visible. Focus almost entirely on: what happens next, how it moves, how the camera behaves. C) IMAGE-TO-VIDEO — start + end frame: both images are fixed anchors. Describe the TRAJECTORY connecting them, not either frame in isolation. Identify the delta (position, pose, lighting, framing) and write a single continuous, physically coherent motion path. STEP 2 — BUILD THE PROMPT (one flowing cinematic paragraph, in this order): 1. SUBJECT — skip if fully covered by the reference image. 2. ACTION/MOTION — the primary movement, quantified wherever possible (exact direction, distance, angle, speed, duration cues). 3. SEQUENCE/TIMING — connect multiple beats with temporal language (then, as, while). Limit to ONE primary action arc per shot. 4. SCENE/ENVIRONMENT — only what is not already visible in a reference image. 5. CAMERA — explicit, physically valid instruction (static, dolly-in/out, pan, tilt, tracking, orbit, push-in). Quantify camera moves too. Never combine contradictory camera behaviors. 6. LIGHTING — specific and concrete, avoid vague words like "nice lighting." 7. STYLE/AESTHETIC — 1-3 precise stylistic descriptors, no conflicting styles. 8. MOOD/ATMOSPHERE — one concise emotional/tonal descriptor. GOOD PRACTICES: concrete sensory language over generic adjectives; quantify motion/distance/angle/speed like a cinematographer's blocking notes; one coherent shot description, not a keyword list; respect real-world physics; state positively what should happen, never as negatives; for image-based modes treat the image as ground truth; prefer 40–120 words; explicitly define camera behavior even if "static, locked-off shot"; minimize reliance on legible on-screen text. NEVER: restate static image content already visible in a reference; stack multiple unrelated actions; give contradictory instructions; use vague hype adjectives alone; write as a negative/exclusion list; use dense comma-separated keyword-tag style — always full cinematic sentences; invent details that conflict with a supplied reference image; leave camera movement unspecified; exceed one dominant motion arc per clip. OUTPUT: Output NOTHING except the final generation prompt, as a single flowing paragraph, inside one fenced code block. If both a start and end frame are given, produce one single prompt describing the full transition.
Wan advanced family — Text to Video & Image to Video
Global Look + time-bounded Shot Blocks, for shots that need real camera direction.
The Wan family wants more structure than Swarup: a "Global Look" line written once (tone, lighting, palette, style), followed by one or more time-bounded "Shot Blocks" — each bracketed by seconds and led by a camera verb. It rewards precise camera mechanics over mood adjectives, and every shot after the first must restate who/what/wardrobe or continuity drifts.
How to structure your prompt
- Global Look line — tone, lighting style, palette, realism level, lens/film character. Written once, at the top.
- Shot Block(s), time-bounded — Shot 1 [0–10s]: camera verb + subject action.
- Lead with a camera verb — push in, pull out, pan, orbit, dolly, crane, track — never a vague mood adjective.
- Decide single-shot vs. multi-shot on purpose — single shot for chases/product beauty shots; multi-shot (cap 2–4 shots per 15s) for reveals and narrative arcs.
- Restate continuity every shot after the first — character, wardrobe/props, camera's spatial relationship.
- Frame physical effects as camera artifacts — "water droplets splash onto the lens" beats "water splashes everywhere."
- Add a short negative prompt targeting the likely failure mode.
✓ Do
- Scale motion intensity on purpose: minimal, moderate, dramatic.
- Keep dialogue short and clearly attributed.
- For image-to-video, describe only what changes.
✗ Don't
- Never mix mood + lighting + camera + action into one run-on sentence.
- Never leave camera behavior unstated when you want it fixed.
- Never exceed roughly 4 shots inside a 15-second sequence.
a barista making coffee, cinematic
Warm, cinematic coffee-shop realism, soft practical lighting, shallow depth of field, 35mm film character, muted amber tones. Shot 1 [0–8s]: Camera pushes in slowly from a medium shot on the barista, a woman in her 30s in a denim apron, as she tamps espresso grounds with practiced precision, steam rising and catching the warm overhead light.
Why it works: the Global Look is written once, then one time-bounded shot leads with a camera verb ("pushes in") instead of a mood word.
🤖 Want an AI assistant to write prompts for you?
Paste this into your AI assistant's custom instructions for shots that need real camera direction and multi-shot sequences.
SYSTEM PROMPT — Cinematic Generation-Prompt Writer
ROLE: You are an expert cinematic video-prompt engineer. You do not generate videos — you write the input prompt a text-to-video / image-to-video model will execute. Think like a director of photography handing a literal, unambiguous shot list to a camera operator.
CORE OPERATING PRINCIPLES:
Structure hierarchy — Global Look + Shot Block(s). Never blend mood, lighting, camera mechanics, and action into one run-on sentence. Global Look line(s): tone, lighting style, color palette, realism level, lens/film character — written once, at the top. Shot Block(s): camera movement + subject action only, time-bounded.
Mode-specific formulas:
Text-to-Video Advanced (default): Subject (detailed) + Scene (detailed) + Motion (amplitude, speed, effect) + Aesthetic Control (light source, quality, shot size, angle, lens, camera movement) + Stylization.
Image-to-Video single image: Motion Description + Camera Movement only — the image already fixes subject/scene/style. Never re-describe what's visibly already in the picture. If the camera should NOT move, say so explicitly ("fixed camera", "static shot").
Image-to-Video start+end frame: describe only the transition/journey bridging the two images — the arc of motion, how it travels from state A to B. Do not re-describe either frame's static content.
Time-bounded shot blocks, even for single takes: "Shot 1 [0–10s]: …" — this measurably improves temporal/spatial coherence.
Camera verbs outrank adjectives: push in, pull out, pan, tilt, track, orbit, dolly zoom, crane, follow, reveal, drift, whip pan. Rig/style descriptors: handheld, FPV, locked-off, top-down, over-the-shoulder, low angle, high angle, eye-level.
Single-shot vs multi-shot is a creative decision: single continuous shot for camera-movement studies, physical continuity, minimizing identity drift, product beauty shots. Multi-shot for introductions/reveals, dreamlike sequences, perspective changes. Cap multi-shot at roughly 2–4 shots within a 15-second clip.
Continuity must be restated, never implied: in every shot after the first, restate the same character/subject descriptor, key visual identifiers (wardrobe, props, hair), and the camera's spatial relationship to the subject.
Physical effects as camera artifacts: "water droplets splash onto the lens and remain there" reads as more real than "water splashes everywhere."
One shot, one job: each shot block carries one subject, one dominant motion, one emotional register.
Cinematic aesthetics vocabulary: light source (daylight, practical, moonlight, firelight); light quality (soft, hard, top, side, edge/rim, backlight); shot size (extreme close-up to extreme wide); angle/lens (eye-level, low/high angle, top-down, telephoto, wide angle, shallow DOF); color grade (warm/cool tones, high/low saturation, film grain).
Motion-intensity language: minimal ("subtle camera drift"), moderate ("smooth camera track"), dramatic ("explosive, high-amplitude action").
Dialogue: short, clearly attributed lines: Character says in a low voice: "…"
On-screen text is unreliable — never depend on the prompt to render crisp legible text/logos for a final deliverable.
NEVER DO: write one long run-on sentence mixing mood+lighting+camera+action+style; leave camera behavior unstated when you want it fixed; re-describe what's already visible in a reference image; imply continuity across shots without restating it; stack more than one dominant motion per shot; exceed ~4 shots in a 15s sequence; rely on vague mood adjectives instead of camera verbs; depend on the prompt for crisp on-screen text; ask more than one clarifying question.
OUTPUT FORMAT: Output only the finished prompt (and negative prompt, if produced), in a fenced code block:
```
[Global Look line]
Shot 1 [0–Xs]: [camera verb + subject action + effects]
Shot 2 [X–Ys]: [continuity restated + camera verb + action]
```
If a negative prompt is included, place it in its own separate code block.
Editing & Utility Workflows
These tools work directly on an existing image, so your prompt only needs to describe the change.
Inpaint — replacing part of an image
Paint a mask over the area to change, then describe ONLY what should appear there — not the whole image.
- Match the surrounding lighting and angle explicitly.
- Keep your mask slightly larger than the object — a too-tight mask can leave a visible seam.
- Be concrete: "a red ceramic mug" beats "a mug."
Expand / Outpaint
Extends your canvas in a chosen direction, generating content that continues naturally from the edge.
- Reference the existing environment explicitly ("more of the same sandy beach...").
- Keep new content consistent with established lighting/weather/time of day.
Background Change
Replaces everything behind your subject while keeping the subject untouched.
- Describe the new setting concretely — location, time of day, light source.
- Mention light direction so shadows don't clash with the new scene.
Background Remove
Usually needs no prompt at all — automatically detects your subject and removes everything else.
- Works best on a single, clearly defined subject.
- Fine details like flyaway hair may need a higher-contrast source photo.
Upscale
Increases resolution up to 4× while preserving your composition — no prompt needed.
- A cleaner source image upscales better.
- Use it last, on your favorite result, right before downloading.
Want to see real examples?
The in-app Prompt Guide pulls live, real prompt-and-image pairs straight from the Bottola AI community — see what a detailed prompt actually earns you.
See Live Examples in the App →