SDXL Prompt Structure for Realistic Photography
Structure SDXL prompts with dual-encoder conditioning, templates, artist tags, and ComfyUI presets for photorealistic results.
On this page
- TL;DR
- Why Prompt Structure Matters in SDXL
- The Three-Conditioning-Channel Model
- CR SDXL Prompt Mix Presets — What Each Does
- Structured Prompt Templates That Work
- Artist and Style Tokens: Using the SDXL Art Style Explorer
- Camera and Lighting Language: What the Community Uses
- Common Errors and Fixes
- Where Automated Tag Generators Help (and Where They Don't)
- Minimal ComfyUI Workflow for Dual-Channel Prompting
- FAQ
- Sources
- Internal Links
TL;DR
- Split prompts across SDXL's two text encoders (CLIP-G and CLIP-L) using ComfyUI's
pos_gandpos_lchannels — routing style tokens to one channel and subject tokens to the other produces stronger separation than concatenating everything. - Use the CR SDXL Prompt Mix Presets node to map prompts to the three conditioning slots (G, L, refiner) via five presets; "style boost 1" and "style text to refiner" are the most practically useful.
- Leverage the SDXL Art Style Explorer (200+ curated tags, 500–700 artists with 6+ tags each) to discover which artist tokens the model actually recognizes — flagged artists won't produce their signature look.
- Follow structured templates (Persona → Clothing → Scene for characters; Main Description → Style → Environment/Mood/Color → Tags for PixelWave) instead of freeform prompts.
Why Prompt Structure Matters in SDXL
SDXL uses two text encoders. The base model's cross-attention receives conditioning from both, producing separate pooled outputs. A refiner model adds a third conditioning path. How you distribute prompt text across these channels changes the image. (https://arxiv.org/abs/2307.01952)
The Comfyroll Custom Nodes expose this routing through the CR SDXL Prompt Mix Presets node, which maps your positive/negative prompts and optional style text into pos_g, pos_l, pos_r (and negative counterparts). Five presets handle common routing patterns. (https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py) (https://civitai.com/articles/1835)
The Three-Conditioning-Channel Model
| Channel | Typical Use | Source |
|---|---|---|
pos_g | Global style, aesthetic direction | https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py |
pos_l | Subject, composition, spatial detail | https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py |
pos_r | Refiner conditioning (if used) | https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py |
Each channel is tokenized and encoded independently. If token lengths differ, ComfyUI pads with empty tokens to match. (https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py)
CR SDXL Prompt Mix Presets — What Each Does
| Preset | pos_g | pos_l | pos_r | Best For | |
|---|---|---|---|---|---|
| Default (no style) | prompt | prompt | prompt | Baseline, single-prompt workflows | |
| Default (with style) | prompt + style | prompt + style | prompt + style | Simple style blending | |
| Style boost 1 | style | prompt | style | Strong style transfer; subject isolated in L | |
| Style boost 2 | prompt + style | prompt + style | prompt + style | Maximum style saturation | |
| Style text to refiner | prompt | prompt | style | Refiner-only style injection | https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py |
Practical tip: Start with Style boost 1. Put your photographic/technical language (camera, lens, lighting) in the prompt field; put aesthetic style tokens (artist names, medium, era) in the style field. The presets route style text to G and R while the subject stays in L — routing, not blending. (https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py)
Structured Prompt Templates That Work
Character Portraits: Persona → Clothing → Scene
This three-block template (community-reported from Civitai 11432) structures character prompts into clear categories: (https://civitai.com/articles/11432)
[Persona]: [Age], [Gender], [Facial features], [Hair & eyes], [Personality], [Pose]
[Clothing]: [Type], [Material], [Fit], [Color & pattern], [Accessories], [Footwear]
[Scene & Mood]: [Environment], [Lighting], [Photography style]
Example:
[Persona]: 28-year-old woman, wavy brown hair, hazel eyes, confident smile, standing with hand on hip
[Clothing]: Beige trench coat, cotton gabardine, tailored fit, black turtleneck underneath, skinny jeans, ankle boots, leather handbag, sunglasses
[Scene & Mood]: Cobblestone Paris street, golden hour, shallow depth of field f/1.8, 85mm lens
General Purpose: PixelWave 4-Part Template
The PixelWave SDXL fine-tune was captioned with ChatGPT using this structure — it's the model author's design for how prompts should be written. (https://civitai.com/articles/5005)(https://civitai.com/models/141592/pixelwave)
- Main Description — Primary subject, action, key features
- Artistic Style — Medium, genre, or named style (e.g., "watercolor", "oil on canvas", "digital art")
- Additional Details — Environment → Mood/Atmosphere → Color Palette
- Tags — Comma-separated keywords for residual concepts
Example:
A lighthouse on a rocky cliff at dawn, watercolor, misty ocean with gentle waves, peaceful morning atmosphere, soft pastels pink orange light blue, sunrise ocean peaceful rocks light
Artist and Style Tokens: Using the SDXL Art Style Explorer
The SDXL Art Style Explorer (Hugging Face Space by terrariyum) provides: (https://huggingface.co/spaces/terrariyum/SDXL-artists-browser)(https://civitai.com/articles/1730)
- 200+ curated style tags organized by medium, style, theme, period, subject
- 500–700 artists (the Civitai article claims 700+; the HF Space notes a temporary beta limit of 500), each with ≥6 descriptive tags
- Flagging for artists SDXL doesn't recognize — these tokens produce generic output
- AND/OR filter logic, favorites export, offline-capable
Workflow: Search a style (e.g., "cyberpunk") → filter artists → pick 2–3 with strong tag overlap → insert their names into your style field (for Style boost 1, that's the style_positive input). Avoid flagged artists unless you're testing.
Camera and Lighting Language: What the Community Uses
The following terms appear in community guides (Civitai 11432) and produce visible shifts in SDXL output. They are reported as community practice — no public ablation study confirms a causal link to photorealism. Treat them as style markers, not optical simulations. (https://civitai.com/articles/11432)
| Category | Tokens from source 11432 |
|---|---|
| Camera body | Canon EOS 5D Mark IV, Sony A7 III |
| Lens | 85mm lens, 50mm f/1.8 lens |
| Depth of field | shallow depth of field f/1.8, deep focus f/11, bokeh background |
| Lighting | golden hour, studio lighting softbox, overcast diffused light |
| Composition | close-up portrait, over-the-shoulder shot, bird's-eye view |
Negative prompt baseline (community folklore, no controlled test): cartoon, illustration, anime, painting, CGI, 3D render, unrealistic proportions, extra fingers, low quality (https://civitai.com/articles/11432)
Common Errors and Fixes
| Symptom | Likely Cause | Fix |
|---|---|---|
| Style tokens ignored | All text in one channel (pos_g or pos_l only) | Use CR Prompt Mix Presets → Style boost 1 |
| Refiner adds unwanted style | Style text not routed to pos_r | Use Style text to refiner preset |
| Artist name gives generic result | Artist not in SDXL vocabulary | Check Art Style Explorer — flagged artists are marked |
| PixelWave output drifts from intent | Freeform prompt instead of 4-part template | Apply Main Description → Style → Environment/Mood/Color → Tags |
Where Automated Tag Generators Help (and Where They Don't)
The SDXL Art Style Explorer is itself a tag discovery tool — it maps 200+ curated style tags to artists the model recognizes, and flags those it doesn't. (https://huggingface.co/spaces/terrariyum/SDXL-artists-browser) For tag generation from images or reference styles, community tools exist, but they share a limitation: they output flat strings with no awareness of SDXL's dual-channel structure.
Best practice: Generate tags from any source → split into subject (L channel) vs. style/aesthetic (G channel) → route via Prompt Mix Presets.
Minimal ComfyUI Workflow for Dual-Channel Prompting
- Load Checkpoint → SDXL base (and refiner if used)
- CR SDXL Prompt Mix Presets → connect
prompt_positive,prompt_negative,style_positive,style_negative; select Style boost 1 - CR SDXL Base Prompt Encoder → feed
pos_g,pos_l,neg_g,neg_lfrom step 2; set resolution/crop params - KSampler → use encoded conditioning from step 3
- (Optional) Refiner → feed
pos_r/neg_rfrom step 2 to refiner sampler
FAQ
How do I split prompts across SDXL's two text encoders? Use the CR SDXL Prompt Mix Presets node in ComfyUI. Set preset to Style boost 1. Put subject/technical language in the main prompt field; put aesthetic/style tokens in the style field. The node routes them to pos_l and pos_g respectively. (https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py)
What's the difference between pos_g, pos_l, and pos_r in ComfyUI? pos_g routes to the G encoder conditioning (global style). pos_l routes to the L encoder conditioning (spatial/subject detail). pos_r routes to the refiner model. Each is independently tokenized and encoded. (https://arxiv.org/abs/2307.01952)(https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py)
Which artist names does SDXL actually recognize? Open the SDXL Art Style Explorer (Hugging Face Space). Search or browse — artists SDXL doesn't know are explicitly flagged. Only use unflagged names in your style field. (https://huggingface.co/spaces/terrariyum/SDXL-artists-browser)
Does PixelWave need a special prompt format? Yes. The model was fine-tuned on ChatGPT-generated captions following a 4-part structure: Main Description, Artistic Style, Additional Details (Environment → Mood → Color), Tags. (https://civitai.com/articles/5005)
Why do camera model tokens improve realism in SDXL? Community observation: they act as style markers associated with photographic training data. No public research confirms a causal optical mechanism. (https://civitai.com/articles/11432)
Sources
- Podell et al., "SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis," arXiv:2307.01952, 2023. (https://arxiv.org/abs/2307.01952)
- Comfyroll Custom Nodes source code,
nodes_sdxl.py. (https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py) - Civitai article 11432, "Ultimate Guide to Creating Realistic SDXL Prompts," xtremedamage, Feb 2025. (https://civitai.com/articles/11432)
- Civitai article 5005, "The PixelWave Prompt Template," humblemikey, Apr 2024. (https://civitai.com/articles/5005)
- Hugging Face Space: terrariyum/SDXL-artists-browser. (https://huggingface.co/spaces/terrariyum/SDXL-artists-browser)
- Civitai article 1835, "SDXL Prompt Mixer Presets," Akatsuzi, Aug 2023. (https://github.com/Suzie1/ComfyUI_Comfyroll_CustomNodes/blob/main/nodes/nodes_sdxl.py)
- Civitai article 1730, "Discover styles and artists for your prompts! SDXL Artist style explorer," terrariyum, May 2024. (https://huggingface.co/spaces/terrariyum/SDXL-artists-browser)
- PixelWave model page, Civitai model 141592. (https://civitai.com/articles/5005)
Internal Links
- SDXL vs. Flux: Which Model for Photorealism?
- ComfyUI Advanced Conditioning: Time-Stepped Prompts
- Training Your Own SDXL LoRA for Style Consistency
- Negative Prompting Patterns That Actually Work
- Upscaling SDXL Outputs Without Artifacts