ComfyUI Tutorial: Node-Based Workflows for Stable Diffusion

# ComfyUI: Workflow Berbasis Node untuk Stable Diffusion ComfyUI adalah lingkungan grafis berbasis node untuk menjalankan Stable Diffusion dan model difusi terkait. Alih-alih menyembunyikan pipeline...

By Ruby Abdullah · · tutorial
ComfyUIStable DiffusionImage GenerationGenerative AIWorkflowPython

ComfyUI: Node-Based Workflows for Stable Diffusion

ComfyUI is a graphical, node-based environment for running Stable Diffusion and related diffusion models. Instead of hiding the generation pipeline behind a single "Generate" button, it exposes every step as a node you can wire together, inspect, and reuse. This tutorial explains how the graph works, how to build the common workflows, and how to drive ComfyUI programmatically through its HTTP API.

Table of Contents

  • What ComfyUI Is and Why a Graph Helps
  • Installation
  • The Default Text-to-Image Workflow
  • How Latents Flow Through the Graph
  • Image-to-Image and Inpainting
  • LoRA and ControlNet Nodes
  • Upscaling Workflows
  • ComfyUI Manager and Custom Nodes
  • Saving and Loading Workflows
  • Driving ComfyUI from the HTTP API
  • SDXL vs SD1.5 Notes
  • Best Practices and Tips
  • Conclusion and Key Takeaways
  • What ComfyUI Is and Why a Graph Helps

    A diffusion image generation pipeline is a sequence of distinct operations: load a model, encode text prompts, prepare an empty latent, run the sampling loop, decode the latent back to pixels, and save the result. Most tools wrap all of this into a form with sliders. ComfyUI instead represents each operation as a node, and you connect the outputs of one node to the inputs of the next.

    This design has practical advantages over a form-driven UI like Automatic1111:

    • Transparency. You can see exactly which model, which conditioning, and which sampler produced an image. There are no hidden defaults applied behind the scenes.
    • Reproducibility. The graph itself is the recipe. Sharing the workflow JSON shares the complete, runnable pipeline, not a screenshot of settings someone has to re-enter.
    • Complex pipelines. Multi-stage flows, such as a base pass followed by a refiner and then an upscale, are awkward to express with a single form. In a graph they are just more nodes wired in sequence.
    • Caching. ComfyUI caches the output of every node. If you change only the seed, it re-runs the sampler but reuses the already-loaded model and encoded prompts, so iteration is fast.

    Automatic1111 is friendly for quick single-image generation and has a large extension ecosystem. ComfyUI trades some of that immediacy for control and is the better fit when the pipeline matters as much as the output, or when you intend to automate it.

    Installation

    ComfyUI runs on Windows, Linux, and macOS. You need Python 3.10 or newer and, for reasonable speed, an NVIDIA GPU with at least 6 GB of VRAM for SD1.5 or 10 GB for SDXL. CPU-only operation works but is slow.

    Manual installation with git

    git clone https://github.com/comfyanonymous/ComfyUI.git
    

    cd ComfyUI

    Create an isolated environment

    python -m venv venv

    source venv/bin/activate # On Windows: venv\Scripts\activate

    Install PyTorch matching your CUDA version first

    pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu124

    Then the remaining dependencies

    pip install -r requirements.txt

    Start the server:

    python main.py
    

    By default it serves the web interface at http://127.0.0.1:8188. Useful flags include --listen 0.0.0.0 to expose it on the network, --port 8000 to change the port, and --lowvram or --cpu for constrained hardware.

    Portable build (Windows)

    For Windows users who do not want to manage Python, the project ships a portable archive containing an embedded Python and all dependencies. Download it from the releases page, extract it, and launch with runnvidiagpu.bat or runcpu.bat. This is the lowest-friction way to get started.

    Placing models

    ComfyUI does not download checkpoints for you. Place your model files in the appropriate subdirectory under models/:

    ComfyUI/
    

    models/

    checkpoints/ # Full SD1.5 / SDXL checkpoints (.safetensors, .ckpt)

    loras/ # LoRA files

    vae/ # Standalone VAE weights

    controlnet/ # ControlNet models

    upscalemodels/ # ESRGAN and similar upscalers

    clipvision/ # CLIP vision encoders

    A checkpoint dropped into models/checkpoints becomes selectable in the Load Checkpoint node after a refresh. Prefer the .safetensors format over .ckpt, since it cannot execute arbitrary code on load.

    The Default Text-to-Image Workflow

    When you open ComfyUI for the first time it loads a default text-to-image graph. Walking through it node by node is the fastest way to understand the system. The nodes, in execution order, are:

    Load Checkpoint

    Loads a checkpoint and exposes three outputs:

    • MODEL — the UNet, used by the sampler.
    • CLIP — the text encoder, used to turn prompts into conditioning.
    • VAE — the autoencoder, used to convert between pixel and latent space.

    Splitting one file into three typed outputs is what lets you mix components later, for example swapping in a standalone VAE.

    CLIP Text Encode (Prompt)

    There are two of these nodes, one for the positive prompt and one for the negative prompt. Each takes the CLIP output and produces a CONDITIONING value. The positive conditioning describes what you want; the negative describes what to steer away from.

    Positive: a photograph of a red fox in a snowy forest, soft morning light, sharp focus
    

    Negative: blurry, low quality, watermark, text, deformed

    Empty Latent Image

    Creates an empty latent tensor of a given width, height, and batch size. This defines the output resolution. For SD1.5 use a base of 512x512; for SDXL use 1024x1024. The batch size controls how many images are generated in one run.

    KSampler

    The core of the pipeline. It takes the MODEL, the positive and negative CONDITIONING, and the LATENT, then runs the denoising loop. Its parameters are the ones you will tune most:

    • seed — the random starting noise. The same seed with identical inputs reproduces the same image.
    • steps — number of denoising iterations. 20 to 30 is typical; more steps cost time with diminishing returns.
    • cfg — classifier-free guidance scale. Higher values follow the prompt more strictly. 6 to 8 is a sensible range for SD1.5.
    • samplername — the sampling algorithm, for example euler, dpmpp2m, or dpmpp2msde.
    • scheduler — the noise schedule, for example normal, karras, or exponential.
    • denoise — how much of the input latent to overwrite with new content. For text-to-image this stays at 1.0.

    VAE Decode

    Takes the denoised LATENT from the sampler plus the VAE, and converts the latent back into a viewable RGB image.

    Save Image

    Writes the decoded image to output/ and embeds the full workflow in the PNG metadata. The preview also appears in the node.

    How Latents Flow Through the Graph

    Stable Diffusion does not denoise pixels directly. It works in a compressed latent space roughly eight times smaller per side, which is why a 512x512 image corresponds to a 64x64x4 latent tensor. Understanding this flow clarifies why the nodes are arranged as they are.

  • Empty Latent Image (or VAE Encode for img2img) produces a LATENT.
  • KSampler iteratively denoises that LATENT, guided by the conditioning.
  • VAE Decode expands the final LATENT back into pixels.
  • Because the sampler operates entirely in latent space, any node that consumes or produces a LATENT can be chained, which is exactly what makes latent upscaling and multi-pass refinement possible. The VAE is only touched at the boundaries, when you enter or leave latent space.

    Image-to-Image and Inpainting

    Image-to-image

    To transform an existing image rather than start from noise, replace Empty Latent Image with two nodes:

    • Load Image — reads a file from input/ and outputs an IMAGE (and a MASK).
    • VAE Encode — converts that IMAGE into a LATENT using the VAE.

    Feed the encoded LATENT into the KSampler and lower the denoise value. This is the key parameter for img2img: it controls how far the result drifts from the input.

    denoise = 0.3   -> subtle changes, keeps composition and most detail
    

    denoise = 0.6 -> noticeable reinterpretation, structure preserved

    denoise = 0.9 -> only loose guidance from the original

    Inpainting

    Inpainting regenerates only a masked region. The Load Image node provides a MASK output, or you can paint a mask in the built-in mask editor. Use VAE Encode (for Inpainting), which accepts both the image and the mask, then send its LATENT to the sampler. Only the masked area is resampled; the rest is preserved. Inpainting-specific checkpoints generally produce cleaner seams than base checkpoints, especially at the mask boundary.

    LoRA and ControlNet Nodes

    LoRA

    A LoRA is a small set of weights that adjusts a base model toward a style, character, or concept. The Load LoRA node sits between Load Checkpoint and the rest of the graph. It takes MODEL and CLIP as inputs and returns modified MODEL and CLIP outputs, which you then route into the sampler and the text encoders.

    Load Checkpoint -> Load LoRA -> KSampler (MODEL)
    

    Load Checkpoint -> Load LoRA -> CLIP Text Encode (CLIP)

    Each Load LoRA node has independent strengthmodel and strengthclip values, typically between 0.5 and 1.0. To stack multiple LoRAs, chain several Load LoRA nodes so the modified outputs of one feed the next.

    ControlNet

    ControlNet conditions generation on a structural input such as a depth map, pose skeleton, or edge map. The relevant nodes are:

    • Load ControlNet Model — loads the ControlNet weights.
    • A preprocessor node (Canny, OpenPose, Depth, and similar) — converts a source image into the control map.
    • Apply ControlNet — takes the positive conditioning, the ControlNet model, and the control image, and returns conditioning augmented with the structural constraint.

    The augmented conditioning then goes into the KSampler in place of the plain positive conditioning, so the generated image follows both the prompt and the supplied structure.

    Upscaling Workflows

    ComfyUI supports two distinct upscaling strategies, and they are often combined.

    Latent upscale

    The Upscale Latent node enlarges the latent tensor before a second sampling pass. Because it happens in latent space, a follow-up KSampler with a moderate denoise (around 0.5) regenerates detail at the higher resolution rather than just stretching pixels. This is the classic "hires fix" pattern.

    KSampler (pass 1) -> Upscale Latent -> KSampler (pass 2, denoise 0.5) -> VAE Decode
    

    Model upscale

    The Upscale Image (using Model) node applies a dedicated upscaling network such as a 4x ESRGAN model from models/upscalemodels. This operates on the decoded RGB image and is purely a super-resolution step, with no diffusion involved. It is fast and sharpens fine detail but invents less new content than a latent pass.

    A common high-quality pattern uses both: a latent upscale with a light second sampling pass to add coherent detail, followed by a model upscale to reach the final resolution.

    ComfyUI Manager and Custom Nodes

    Much of ComfyUI's capability comes from community custom nodes. The ComfyUI Manager is the standard tool for discovering and installing them. Install it by cloning into the customnodes directory:

    cd ComfyUI/customnodes
    

    git clone https://github.com/ltdrdata/ComfyUI-Manager.git

    Restart ComfyUI and a "Manager" button appears in the menu. From there you can install custom node packs, update them, and install missing nodes referenced by a workflow you loaded. The "Install Missing Custom Nodes" feature is especially useful when you open someone else's workflow and several nodes show up red because the packs are not present.

    Be selective with custom nodes. Each pack is third-party code that runs in your environment, and a poorly maintained pack can break on a ComfyUI update. Install what a workflow actually needs and keep the set small.

    Saving and Loading Workflows

    A workflow can be exported to a JSON file from the menu. There are two formats to be aware of:

    • The editor format (saved via Save) preserves node positions and the full visual graph for re-opening in the UI.
    • The API format (Save (API Format), which requires enabling developer options in settings) is a flattened representation keyed by node ID, intended for the HTTP API.

    ComfyUI also embeds the editor workflow directly in the metadata of every PNG it saves. This means you can drag a generated PNG back onto the canvas and the exact graph that produced it is restored, which makes shared images self-documenting. Be aware that this metadata is stripped if the image is re-encoded or run through a tool that discards PNG text chunks.

    Driving ComfyUI from the HTTP API

    ComfyUI exposes an HTTP and WebSocket API on the same port as the UI, which makes it straightforward to integrate into a backend service. The flow is: load a workflow in API format, edit the fields you care about, POST it to /prompt, then wait for completion via the WebSocket and fetch the result.

    First, export your graph using Save (API Format). The result is a dictionary keyed by node ID. The example below assumes node 4 is Load Checkpoint, node 6 is the positive CLIP Text Encode, and node 3 is the KSampler. Inspect your own exported JSON to confirm the IDs.

    import json
    

    import uuid

    import urllib.request

    import urllib.parse

    import websocket # pip install websocket-client

    SERVER = "127.0.0.1:8188"

    CLIENTID = str(uuid.uuid4())

    def loadworkflow(path):

    with open(path, "r", encoding="utf-8") as f:

    return json.load(f)

    def queueprompt(workflow):

    """Submit a workflow and return the promptid."""

    payload = {"prompt": workflow, "clientid": CLIENTID}

    data = json.dumps(payload).encode("utf-8")

    req = urllib.request.Request(f"http://{SERVER}/prompt", data=data)

    req.addheader("Content-Type", "application/json")

    with urllib.request.urlopen(req) as resp:

    return json.load(resp)["promptid"]

    def gethistory(promptid):

    url = f"http://{SERVER}/history/{promptid}"

    with urllib.request.urlopen(url) as resp:

    return json.load(resp)

    def getimage(filename, subfolder, foldertype):

    params = urllib.parse.urlencode(

    {"filename": filename, "subfolder": subfolder, "type": foldertype}

    )

    with urllib.request.urlopen(f"http://{SERVER}/view?{params}") as resp:

    return resp.read()

    Now edit the prompt and seed, submit, and wait on the WebSocket for the executing signal that marks completion:

    def waitforcompletion(promptid):
    

    """Block until the given prompt finishes, using the WebSocket feed."""

    ws = websocket.WebSocket()

    ws.connect(f"ws://{SERVER}/ws?clientId={CLIENTID}")

    try:

    while True:

    message = ws.recv()

    if isinstance(message, str):

    event = json.loads(message)

    if event["type"] == "executing":

    data = event["data"]

    # node is None and matching promptid => this prompt is done

    if data["node"] is None and data["promptid"] == promptid:

    break

    finally:

    ws.close()

    def run(workflowpath, positiveprompt, seed):

    workflow = loadworkflow(workflowpath)

    # Edit the fields. Node IDs come from your exported API JSON.

    workflow["6"]["inputs"]["text"] = positiveprompt

    workflow["3"]["inputs"]["seed"] = seed

    promptid = queueprompt(workflow)

    waitforcompletion(promptid)

    # Pull the output filenames from history and download the images.

    history = gethistory(promptid)[promptid]

    images = []

    for nodeoutput in history["outputs"].values():

    for image in nodeoutput.get("images", []):

    data = getimage(image["filename"], image["subfolder"], image["type"])

    images.append((image["filename"], data))

    return images

    if name == "main":

    results = run(

    "workflowapi.json",

    positiveprompt="a lighthouse on a cliff at dusk, dramatic clouds",

    seed=12345,

    )

    for filename, data in results:

    with open(filename, "wb") as f:

    f.write(data)

    print("Saved", filename)

    If you prefer polling over the WebSocket, you can repeatedly call /history/{promptid} and treat the prompt as finished once that ID appears in the response with populated outputs. The WebSocket is preferable in production because it also streams progress and avoids unnecessary requests. The API format that /prompt expects looks like this for a single node:

    {
    

    "3": {

    "classtype": "KSampler",

    "inputs": {

    "seed": 12345,

    "steps": 25,

    "cfg": 7.0,

    "samplername": "dpmpp2m",

    "scheduler": "karras",

    "denoise": 1.0,

    "model": ["4", 0],

    "positive": ["6", 0],

    "negative": ["7", 0],

    "latentimage": ["5", 0]

    }

    }

    }

    Each input is either a literal value or a two-element array [nodeid, outputindex] representing a wire from another node's output. This is the same connection model you see visually in the editor.

    SDXL vs SD1.5 Notes

    The two model families differ in ways that affect your graphs:

    • Native resolution. SD1.5 is trained at 512x512; SDXL at 1024x1024. Set the Empty Latent Image accordingly. Asking SD1.5 for 1024x1024 directly tends to produce duplicated subjects.
    • Text encoders. SDXL uses two CLIP encoders, so its conditioning nodes (CLIP Text Encode SDXL) expose extra fields, and a full SDXL pipeline often pairs a base model with a refiner model in a two-stage sampler setup.
    • VRAM. SDXL needs noticeably more memory. On 8 GB cards use --lowvram or rely on tiled VAE decode nodes.
    • Samplers and CFG. SDXL generally works well at slightly lower CFG (around 5 to 7) than SD1.5.

    For learning the node model, SD1.5 is lighter and faster to iterate on. Move to SDXL when output quality is the priority and your hardware allows.

    Best Practices and Tips

    • Name and group nodes. Right-click to title nodes and use group boxes to label sections such as "base pass" or "upscale." A labelled graph is far easier to revisit weeks later.
    • Use primitive and reroute nodes. Expose shared values like seed or resolution through a Primitive node so you change them in one place. Reroute nodes keep long wires tidy.
    • Version your workflow JSON. Keep API-format and editor-format exports in a git repository alongside notes on which checkpoints they expect.
    • Pin the seed while iterating. Fix the seed when tuning prompts or CFG so you compare like with like, then switch back to random when exploring.
    • Keep the model directory organized. Use clear filenames and subfolders; ComfyUI lists whatever is present, and a cluttered checkpoints folder slows you down.
    • Separate the server from your client code. When using the API, run ComfyUI as a long-lived service and let your application submit jobs. This reuses the loaded model across requests and avoids repeated startup cost.

    Conclusion and Key Takeaways

    ComfyUI treats image generation as an explicit, inspectable graph rather than a single opaque action. That structure is what gives it its main strengths: you can see and reproduce exactly how an image was made, build multi-stage pipelines that would be cumbersome in a form-based UI, and automate the whole thing through a clean HTTP API.

    Key points to carry forward:

    • The default text-to-image graph is the foundation: Load Checkpoint, CLIP Text Encode, Empty Latent Image, KSampler, VAE Decode, Save Image.
    • Everything happens in latent space until VAE Decode, which is why latent upscaling and multi-pass refinement compose so naturally.
    • LoRA and ControlNet are added by inserting nodes into the existing flow, not by reconfiguring it.
    • Workflows are portable as JSON, and ComfyUI embeds them in saved PNGs for self-documenting outputs.
    • The /prompt, /history, and WebSocket endpoints let you load an API-format workflow, edit prompt and seed, submit it, and retrieve images from your own code.

    Start with the SD1.5 default graph, learn how the wires carry MODEL, CONDITIONING, and LATENT between nodes, and grow from there into LoRA, ControlNet, upscaling, and finally API-driven automation.

    Related Articles

    Stable Diffusion Tutorial: Generative AI for Image Generation

    Stable Diffusion: Tutorial Komprehensif Daftar Isi Pendahuluan Prasyarat Memahami Arsitektur Stable Diffusion 4....

    Temporal Tutorial: Durable Execution for Reliable Workflows

    Temporal dengan Python: Durable Execution untuk Workflow yang Andal Temporal adalah platform untuk durable execution: ia...

    PaddleOCR: High Accuracy Text Extraction from Images and Documents

    PaddleOCR: Ekstraksi Teks dari Gambar dan Dokumen dengan Akurasi Tinggi Halo temen-temen, kali ini kita bahas salah satu...

    DSPy: Stop Hand-Tuning Prompts, Let the Compiler Optimize Them

    DSPy: Berhenti Ngoprek Prompt Manual, Biarkan Compiler yang Optimasi Halo temen-temen, kali ini aku mau ngenalin satu li...