Ideogram 4.0: ComfyUI Workflow, Prompt Tips and vs Krea 2
A hands-on look at Ideogram 4.0 in ComfyUI: the five model files and where they go, how structured JSON prompts work, prompt habits that hold up, and a side-by-side with Krea 2 on control, text and price, plus settings for running both on PicassoIA.
Most people meet Ideogram 4.0 through a hosted app and never see the part that makes it unusual: the weights are open, and ComfyUI ships a ready-made workflow that runs them on your own machine. That changes what you download, how you write prompts, and how the model stacks up against Krea 2, which goes the other way with a single creativity dial and a pile of style references. Below you get the ComfyUI setup file by file, the JSON prompt format, the prompt habits that pay off, a plain comparison with Krea 2, and a shortcut for running both in the browser on PicassoIA.
What Ideogram 4.0 Actually Is
Ideogram built its name on one thing: text that reads correctly inside an image. Version 4.0 keeps that strength and adds a structured way to describe a scene, so a poster, a product shot or a portrait can be laid out like a page instead of left to chance.
Open Weights, Not Just an API
The notable shift is distribution. The model was released with open weights, and ComfyUI added support on day one, which means you can run it locally instead of paying per image through someone else's server. ComfyUI's documentation also lists a separate partner node route that calls the hosted service, so you choose between local files and a remote call depending on your hardware.
Two Prompt Formats
You can write plain natural language, which suits quick ideas, or a structured JSON caption for precise control over layout, colors and style. The JSON route is what the model was built around, and most community talk centers on it. On Ideogram v4 Quality the same split shows up as two separate inputs: prompt for natural language and json_prompt for the structured version. Use one or the other, never both.
đź’ˇ Quick rule: if the image has a layout that matters (a poster, a label, a product on a defined background), reach for JSON. If it is a mood or a portrait, plain text is faster.
ComfyUI Setup, File by File
Local setup is mostly a download job. The official workflow expects five files in specific folders, and putting them in the wrong place is the most common reason the workflow refuses to load.
The Five Model Files
File
Folder
Size
ideogram4_fp8_scaled.safetensors
models/diffusion_models/
about 13.8 GB
ideogram4_unconditional_fp8_scaled.safetensors
models/diffusion_models/
about 13.8 GB
qwen3vl_8b_fp8_scaled.safetensors
models/text_encoders/
about 8 GB
gemma4_e4b_it_fp8_scaled.safetensors
models/text_encoders/
about 2 GB
flux2-vae.safetensors
models/vae/
about 335 MB
That adds up to roughly 38 GB on disk. The files come from the Comfy-Org/Ideogram-4 repository on Hugging Face. The second diffusion file handles the unconditional pass, which is the reason the model does not need a negative prompt: the unconditional side simply drops the text tokens, so there is no separate "avoid this" string to fill in.
Loading the Workflow
Download the workflow file from the ComfyUI documentation page and drag it onto the canvas. The graph is short: an Ideogram4 subgraph node that does the heavy lifting, a ResolutionSelector for output size, and a Save Image node. If ComfyUI shows red missing nodes, update ComfyUI first, since the workflow relies on recent additions.
The default workflow ships with a structured JSON prompt already filled in. Edit that example instead of starting from a blank box, because it shows exactly which fields the node expects. If you prefer building prompts visually, the Ideogram 4 Prompt Builder in KJNodes lets you place elements on a canvas and writes the JSON for you.
Hardware Reality Check
Here is the honest gap: the documentation does not state a VRAM figure or a minimum system spec. The file sizes tell you more than the docs do. Two 13.8 GB diffusion files plus an 8 GB text encoder mean a card with a lot of memory, fast storage, and patience on the first load. Run the default workflow once, watch your memory use, and only then build a larger pipeline around it. If your GPU is tight, the hosted route is the sensible fallback.
Writing JSON Prompts That Work
The structured caption is built from a scene summary, a style block, a background, and optional per-object descriptions with bounding boxes and hex color palettes. The default ComfyUI workflow organizes this under three top-level names.
The Three Top Blocks
high_level_description: one or two sentences that name the whole scene.
style_description: lens, lighting, film stock, mood. This is where "documentary photograph, soft window light, fine grain" lives.
compositional_deconstruction: the list of elements, each with a box, a description and a palette.
Boxes and Colors
Here is the shape of a prompt, trimmed down to one element:
{
"high_level_description": "A white ceramic mug on a walnut desk beside an open notebook, morning window light",
"style_description": "Documentary photograph, 50mm lens, soft shadows, fine film grain, natural colors",
"compositional_deconstruction": [
{
"description": "White ceramic mug with a thin blue rim, steam rising",
"bounding_box": [0.55, 0.40, 0.80, 0.75],
"colors": ["#F4F1EA", "#3C6E9E"]
}
]
}
Treat this as a shape, not a schema. The exact field names inside each element and the way box coordinates are written come from the default workflow, so copy them from there rather than from this sketch.
Colors are worth the extra typing. A hex value pins the mug to a specific blue, where "blue" alone drifts from run to run. Boxes do the same job for position: instead of hoping "mug on the right" lands where you want, you hand the model a rectangle.
There is also a practical reason to prefer JSON even for simple scenes. ComfyUI's notes say plain-text prompts see a higher false-positive rate from the safety filter, and structured prompts reduce that blocking.
Prompt Tips That Hold Up
JSON gets the attention, but most failed images trace back to vague wording, not missing structure. Three habits fix the majority of them.
Say What the Camera Sees
Describe the photograph, not the idea. PicassoIA's own example prompts for the model read like a camera note: a color film portrait, shallow depth of field, sharp focus on the eye, blurred background bokeh, high ISO film grain, candid documentary style. Lens, light direction and film stock give the model something concrete to render, and they keep results out of the glossy, over-polished look.
Keep Text Short
If the image needs words, put them in quotes and say where they sit. The example prompts on the model pages do exactly that, with a short quoted label and a position such as "centered". One to three words per text element is the sweet spot. Long sentences inside an image are where every model, this one included, starts to slip.
Skip the Negative Prompt
The architecture already handles the "avoid this" side, so write what you want and stop there. On the hosted version, passing a natural-language prompt switches on Magic Prompt automatically, which expands a rough description before generating. That helps beginners and gets in the way if you wrote a precise prompt, so check the output against your intent. If Magic Prompt rewrites too much, move the same idea into json_prompt.
Ideogram 4.0 vs Krea 2
Krea 2 comes in two sizes on PicassoIA. Krea 2 Large is listed for photorealism and artistic range in one model, and Krea 2 Medium is positioned for anime and painterly work.
One caveat before the table: this comparison rests on model specs, ComfyUI documentation and public listings. It is not a pixel-level shootout, so run the same prompt on both before you commit a project to either.
Control Versus Taste
The two models ask you to steer in different ways.
Ideogram 4.0 gives you structure. Boxes place elements, hex values pin colors, and quoted text lands where you said it should. You decide, the model executes.
Krea 2 gives you a dial. The creativity setting runs from raw, which renders only what you describe, through low and medium, up to high, where the model takes real liberty with light, atmosphere and composition. You can also attach up to 10 style reference images and set a strength from 0 to 1 (the default is 0.5) to decide how strongly their look shapes the result.
Cost and Speed
A public comparison site lists Krea 2 Medium Turbo at about $15 per 1,000 images and Ideogram 4.0 at about $60 per 1,000. In blind preference voting the two land almost level, at 1,223 and 1,217 Elo, with no clear leader. Note that the cheap figure belongs to the Medium Turbo variant, not Large.
On speed, the example runs listed on PicassoIA's model pages give a rough feel: one Ideogram v4 Quality sample took about 67 seconds, an Ideogram v4 Balanced sample about 41 seconds, and two Krea 2 Large samples about 103 and 143 seconds. Those are single runs, not a benchmark, but they hint that Balanced is the iteration model and Quality is the finishing pass.
Ideogram 4.0
Krea 2
Prompt input
Natural language or structured JSON
Natural language plus a creativity setting
Layout control
Bounding boxes and color palettes
Aspect ratio presets and style references
Text inside images
Headline strength, legible labels
Possible, but not the main selling point
Style transfer
Written into the style block
Up to 10 reference images
Negative prompt
Not needed
Not part of the inputs
Public price figure
About $60 per 1,000 images
About $15 per 1,000 (Medium Turbo)
The short verdict: pick Ideogram when the image is a layout (posters, packaging, thumbnails with words). Pick Krea 2 when the image is a mood (portraits, editorial scenes, anything where you want the model to bring taste).
Run Both on PicassoIA
If 38 GB of downloads and an unknown VRAM requirement sound like a bad Tuesday, both families run in the browser on PicassoIA with no setup, no watermarks and clean downloads.
Paste a natural-language description into prompt, or a JSON object into json_prompt. Fill only one.
Choose a resolution. There are over 20 sizes, from 2048x2048 up to 3328x1248. Leave it on None and the model picks the aspect ratio itself. For a blog header, 2560x1440 is a clean 16:9.
Turn on copyright detection if you plan to publish commercially. It is opt-in and off by default.
Generate, then repeat the same prompt on Ideogram v4 Balanced when you want faster drafts before the final render.
creativity: use raw for product shots that must match your words, medium as the everyday default, high for concept art.
aspect_ratio: 16:9 for banners, 2.35:1 for a cinematic strip, 9:16 for vertical posts. Eight ratios are available.
seed: lock a composition you like, then change one phrase at a time and compare.
Style references: start at strength 0.5 and move up only if the reference barely shows.
Picking the Right Model
A quick way to decide:
Text and layout come first: Ideogram v4 Quality.
Fast drafts of the same idea: Ideogram v4 Balanced.
Photographic mood or a specific look from references: Krea 2 Large.
Painterly or anime direction: Krea 2 Medium.
If neither fits, FLUX 2 Pro, Seedream 5 Pro and GPT Image 2 are worth a test with the same prompt. And when a finished graphic needs its text pulled out as editable pieces, Layerize separates the text layers.
Make Your Own Images Today
The fastest way to settle the Ideogram versus Krea question is to run your own prompt through both. Open Picasso IA, load Ideogram v4 Quality and Krea 2 Large in two tabs, paste the same one-sentence idea into each, and compare the results side by side. Then rewrite the prompt as JSON for the first and nudge the creativity setting on the second.
Ten minutes of that tells you more than any spec sheet. Pick a scene from your own work, run it twice, and keep whichever model gets closer on the first try.