How to Use ComfyUI: A Practical, Step‑by‑Step Guide for Beginners
If you’ve heard that ComfyUI is “node-based and super powerful” but felt intimidated by all the boxes and wires, you’re not alone. The good news: once you learn a few core concepts—checkpoints, encoders, samplers, and decoders—you’ll be building image workflows like a pro. This practical guide walks you through how to use ComfyUI from installation to your first SDXL images, plus workflows for ControlNet, LoRAs, and quality/performance tuning.
By the end, you’ll know exactly how to use ComfyUI to make consistent, repeatable, and flexible image generations without guesswork.
What Is ComfyUI and Why Use It?
ComfyUI is a visual, node-based interface for Stable Diffusion that lets you design your image pipeline step by step. Instead of a single “Generate” button, you connect nodes—each handling a distinct task such as loading a model, encoding text, sampling latents, or decoding the final image. It’s fast, modular, and transparent—perfect for learning, experimentation, and production workflows.,.
Quick Start: Install and Launch ComfyUI
- Windows/macOS/Linux: Follow the official repo and community installation guides. You can use manual installation (Python + dependencies) or packaged methods depending on your platform and GPU. The ComfyUI wiki provides step-by-step setup for Windows, macOS (including Apple Silicon), and Linux.,,.
- Models: Place your Stable Diffusion checkpoints (e.g., SDXL base/refiner or SD 1.5) in the
models/checkpoints folder. Put VAE files in models/vae, LoRAs in models/loras, ControlNet models in models/controlnet.
- Launch: Run the start script for your OS; ComfyUI opens in your browser. The canvas is where you’ll wire nodes together.
Tip: Keep your GPU drivers and CUDA toolkit up to date for best performance.
Core Concept: The Minimal Text‑to‑Image Workflow
ComfyUI’s basic text-to-image flow (SD 1.5 style) looks like this:
- Output: UNet, CLIP, and VAE components
- Node: CLIP Text Encode (Positive)
- Node: CLIP Text Encode (Negative)
- Output: Conditioning embeddings for guidance
- Inputs: UNet, positive/negative conditioning, seed, steps, sampler (e.g., DPM++ 2M Karras), and CFG scale
This basic graph—Checkpoint → CLIP (pos/neg) → KSampler → VAE Decode → Save—is the foundation for almost everything you’ll do in ComfyUI.,.
SDXL Workflow: Base + (Optional) Refiner
SDXL uses dual text encoders and often benefits from a refiner pass.
- Load SDXL Base: Use an SDXL‑compatible checkpoint. Many SDXL templates include two CLIP encoders (for large/small context). Feed both positive and negative prompts.
- KSampler (Base): Generate latents at 1024×1024 (or your target). Save latents or decoded images.
- Optional Refiner: Load the SDXL Refiner checkpoint and run an additional KSampler pass conditioned on the base output, then decode with VAE.
This two-stage process can significantly improve detail and coherence at higher resolutions.
Hands‑On: Build Your First ComfyUI Graph
- Start from a template: In the sidebar, load a built-in text-to-image example.
- Replace the checkpoint: Select your SDXL or SD 1.5 model.
- Write your prompt: Use the Positive and Negative CLIP nodes. Example:
- Positive: “cinematic portrait, soft studio lighting, 85mm lens, highly detailed, film grain”
- Negative: “blurry, low-res, deformed, extra fingers, watermark”
- Steps: 20–35 for speed/quality balance
- Sampler: DPM++ 2M Karras (reliable) or Euler a (fast)
- CFG: 4.5–7.5 (higher pushes prompt harder, but can oversaturate)
- Seed: Fix it for reproducibility; vary for exploration
- Resolution: For SD 1.5, start at 512×512 or 768×768. For SDXL, 1024×1024 works well.
- Decode and Save: Add VAE Decode → Save Image. Click Queue Prompt to generate.
Understanding the Key Nodes (In Plain English)
- Checkpoint Loader: Loads your diffusion model (UNet), text encoder(s) (CLIP), and VAE. Think of it as your “engine + language brain + image translator.”
- CLIP Text Encode: Converts your prompt into numerical embeddings the model understands. Use both positive and negative text encoders.
- KSampler: The heart of image synthesis. It denoises latent noise guided by your prompt and sampler method across a number of steps.
- VAE Decode: Translates final latents into a viewable image. Swapping VAEs changes color/contrast fidelity.
- Save Image: Writes output to disk with metadata so you can recreate results later.
For a deeper dive on these building blocks, see beginner-friendly breakdowns and node explainers.,.
Power‑Ups: LoRA, ControlNet, and Image‑to‑Image
Use LoRA for Style or Subject Control
- Add a LoRA Loader node and connect it to your model branch.
- Strength: Start around 0.6–0.8; adjust based on style intensity or overfitting.
- Multiple LoRAs: Chain or merge, but watch for conflicts; lower strengths when stacking.
Add ControlNet for Precise Composition
- ControlNet nodes let you steer composition using an input map (Canny, Depth, OpenPose, etc.).
- Typical flow: Load ControlNet model → Preprocess your guide image (e.g., Canny edge) → Feed ControlNet conditioning into KSampler along with your text conditioning.
- Weight: 0.5–1.2 is a good start. Too high can overpower your prompt.
Image‑to‑Image or Inpainting
- Replace the initial noise with an image latent via VAE Encode.
- Adjust denoise strength in KSampler to control how much of the original image remains.
- For inpainting, use a mask input and an inpaint‑aware sampler pipeline.
Quality Tuning: Prompts, CFG, Samplers, and Seeds
- Prompt engineering: Use concise descriptors, not paragraphs. Order matters less than clarity, but keep critical attributes up front.
- Low (3–5): More creative, less prompt adherence
- High (9–12): Strong adherence, can create artifacts
- DPM++ 2M Karras: Clean, reliable
- Euler a: Fast and expressive, great for previews
- UniPC / Heun / DDIM: Worth testing; results vary by model
- Fixed seed = reproducible results
- Vary seed = explore diversity
Performance Tips for Smooth Renders
- VRAM budgeting: Lower resolution, steps, or batch size if you hit OOM. SDXL at 1024×1024 can require 8–12 GB VRAM depending on nodes.
- Half precision: Enable fp16 where supported for big memory savings with negligible quality loss.
- Tiling and latent upscalers: Generate smaller, then upscale via a latent upscaler node or image upscaler model to save VRAM.
- Caching: Re‑use CLIP encodings and decoded VAEs across runs when prompts don’t change.
- Avoid unnecessary branches: Extra disconnected nodes still consume memory when executed in the same queue.
Organizing Workflows Like a Pro
- Group nodes: Use frames/labels to organize sections (Prompt, Model, Sampler, Output, etc.).
- Parameter panels: Create “control” nodes (e.g., empty prompt boxes, sliders) at the top for easy tuning.
- Save/share: Export your workflow JSON and keep a
models used note for reproducibility.
- Versioning: Keep separate graphs for SD 1.5, SDXL, and specialty pipelines (anime, photoreal, depth-to-image, etc.).
Troubleshooting Common Issues
- Wrong VAE or missing VAE Decode
- Denoise too low (e.g., <0.2 in img2img)
- Try another VAE; some VAEs improve contrast noticeably
- Lower CFG or change sampler
- Nothing changes across runs:
- Seed is fixed; enable randomize or set a new seed
- Reduce resolution, steps, or batch size; switch to fp16
- Close other GPU apps; simplify ControlNet/LoRA stacks
- Model not found / red node:
- Verify file paths and model folders; confirm file extensions
Learn Faster With Pre‑Built Workflows
Video walkthroughs and beginner series can accelerate your learning curve with ready-to-run graphs you can pause and dissect.,,. Written tutorials and wikis provide node explanations and updated installation steps to keep you current.,,.
Advanced: Modularizing and Extending Your Graphs
- API/External nodes: Some tutorials cover connecting ComfyUI to external AI services via special nodes, enabling hybrid pipelines and offloading heavy tasks.
- Node libraries and extensions: Explore community nodes for schedulers, upscalers, and preprocessing (pose, depth, segmentation). Always check compatibility with your ComfyUI version.
- SDXL refiners and chained samplers: Run staged denoising (base → refiner) or even multiple samplers for stylistic blending.
Worth Noting: Speeding Up Prompting With Sider.AI
If you frequently iterate on prompts, references, or descriptions, you may want a sidekick to brainstorm and refine variations. By the way, Sider.AI can help you quickly draft structured prompts, generate negative prompt lists, and summarize your workflow experiments so you don’t lose track between runs. You can try it here: A Simple SDXL Starter Workflow (Copy This Pattern)
- Checkpoint Loader (SDXL Base)
- CLIP Text Encode (Positive) — “ultra-detailed product photo, softbox lighting, 50mm lens, reflective surface”
- CLIP Text Encode (Negative) — “low-res, motion blur, watermark, background clutter”
- KSampler: 1024×1024, 28 steps, DPM++ 2M Karras, CFG 5.5, fixed seed
Optional add-ons:
- Refiner pass with SDXL Refiner checkpoint at 10–15 steps
- ControlNet (Depth) with a simple object silhouette for layout
- LoRA at 0.6 for a specific brand or art style
Key Takeaways
- ComfyUI’s power comes from its transparency—build your pipeline node by node.
- The core text‑to‑image chain is simple: Checkpoint → CLIP (pos/neg) → KSampler → VAE Decode → Save.
- SDXL benefits from dual encoders and an optional refiner pass for detail.
- LoRAs and ControlNet give you style control and composition precision.
- Tune CFG, sampler, and seed for quality and consistency; manage VRAM with fp16 and sensible resolutions.
- Organize workflows and version them for painless iteration.
Next Steps
- Install ComfyUI following the repo/wiki instructions and launch a sample workflow.,
- Rebuild the minimal chain from scratch to cement the basics.
- Add ControlNet and a LoRA, then A/B test sampler and CFG settings.
- Save and share your workflow JSON with notes on models, seeds, and parameters.
Happy generating—and welcome to the calm, controllable world of ComfyUI.
FAQ
Q1:How do I install and run ComfyUI on Windows, macOS, or Linux?
Follow the official repo and the community wiki for platform-specific steps, model folder locations, and dependencies. After installation, launch the local server and open ComfyUI in your browser to start wiring nodes.,.
Q2:What’s the simplest ComfyUI workflow for text-to-image?
Load a checkpoint, encode positive and negative prompts with CLIP, run a KSampler, decode with VAE, then save the image. This chain is the foundation for how to use ComfyUI effectively for most generations.,.
Q3:How do I use SDXL in ComfyUI?
Use an SDXL checkpoint with dual text encoders, then optionally add a refiner pass for better detail. Run at 1024×1024 with balanced CFG (around 5–7) and an efficient sampler like DPM++ 2M Karras..
Q4:Can I add ControlNet and LoRA in the same ComfyUI workflow?
Yes. Load your LoRA and ControlNet nodes, connect them to the model and KSampler conditionings, and tune weights (e.g., 0.6–0.8 for LoRA, ~0.5–1.2 for ControlNet). Watch VRAM usage and reduce resolution or steps if you hit OOM.
Q5:Why are my ComfyUI images low‑contrast or washed out?
Try a different VAE, lower CFG, or switch samplers. Some VAEs produce more faithful color and contrast; small adjustments can fix washed-out results quickly.