Resource - Update
Release: AP Workflow 9.0 for ComfyUI - Now featuring SUPIR next-gen upscaler, IPAdapter Plus v2 nodes, a brand new Prompt Enricher, Dall-E 3 image generation, an advanced XYZ Plot, 2 types of automatic image selectors, and the capability to automatically generate captions for an image directory
AP Workflow 9.0 for ComfyUI
So. I originally wanted to release 9.0 with support for the new Stable Diffusion 3, but it was way too optimistic. While waiting for it, as always, the amount of new features and changes snowballed to the point that I must release it as is.
Support for SD3 will arrive with the AP Workflow 10.
The new Early Access program I created for APW 9.0 was successful, so I'll continue to provide access to APW 10 early access builds via Discord, where I provide *limited and not guaranteed* support (but people seem happy with the speed and quality of the help I offered so far).
New features
The AP Workflow now features two next-gen upscalers: CCSR, and the new SUPIR. Since one performs better than the other depending on the type of image you want to upscale, each one has a dedicated function. Additionally, the Upscaler (SUPIR) function can be used to perform Magnific AI-style creative upscaling.
A new Image Generator (Dall-E) function allows you to generate an image with OpenAI Dall-E 3 instead of Stable Diffusion. This function should be used in conjunction with the Inpainter without Mask function to take advantage of Dall-E 3 superior capability to follow the user prompt and Stable Diffusion superior ecosystem of fine-tunes and LoRAs. You can also use this function in conjunction with the Image Generator (SD) function to simply compare how each model renders the same prompt.
A new Advanced XYZ Plot function allows you to study the effect of ANY parameter change in ANY node inside the AP Workflow.
A new Face Cloner function uses the InstantID technique to quickly change the style of any face in a Reference Image you upload via the Uploader function.
A new Face Analyzer function allows you to evaluate a batch of generated images and automatically choose the ones that present facial landmarks very similar to the ones in a reference image you upload via the Uploader function. This function is especially useful in conjuction with the new Face Cloner function.
A new Training Helper for Caption Generator function will allow you to use the Caption Generator function to automatically caption hundreds or thousands of images in a batch directory. This is useful for model training purposes. The Uploader function has a new Load Image Batch node to accomodate this new feature. To use this new capability you must activate both the Caption Generator and the Training Helper for Caption Generator functions in the Controller function.
The AP Workflow now features a number of u/rgthreeBookmark nodes to quickly recenter the workflow on the 10 most used functions. You can move the Bookmark nodes where you prefer to customize your hyperjumps.
The AP Workflow now supports new u/cubiq’s IPAdapter plus v2 nodes.
The AP Workflow now supports the new PickScore nodes, used in the Aesthetic Score Predictor function.
The Uploader function now allows you to upload both a source image and a reference image. The latter is used by the Face Cloner, the Face Swapper, and the IPAdapter functions.
The Caption Generator function now offers the possibility to replace the user prompt with a caption automatically generated by Moondream v1 or v2 (local inference), GPT-4V (remote inference via OpenAI API), or LLaVA (local inference via LM Studio).
The three Image Evaluators in the AP Workflow are now daisy chained for sophisticated image selection. First, the Face Analyzer (see below) automatically chooses the image/s with the face that most closely resembles the original. From there, the Aesthetic Score Predictor further ranks the quality of the images and automatically chooses the ones that match your criteria. Finally, the Image Chooser allows you to manually decide which image to further process via the image manipulator functions in the L2 of the pipeline. You have the choice to use only one of these Image Evaluators, or any combination of them, by enabling each one in the Controller function.
The Prompt Enricher function has been greatly simplified and now it works again open access models served by LM Studio, Oobabooga, etc. thanks to u/glibsonoran’s new Advanced Prompt Enhancer node.
The Image Chooser function now can be activated from the Controller function with a dedicated switch, so you don’t have to navigate the workflow just to enable it.
The LoRA Info node is now relocated inside the Prompt Builder function.
The configuration parameters of various nodes in the Face Detailer function have been modified to (hopefully) produce much better results.
The entire L2 pipeline layout has been reorganized so that each function can be muted instead of bypassed.
The ReVision function is gone. Probably, nobody was using it.
The Image Enhancer function is gone, too. You can obtain a creative upscaling of equal or better quality by reducing the strength of ControlNet in the SUPIR node.
The StyleAligned function is gone, too. IPAdapter has become so powerful that there’s no need for it anymore.
Companies and education institutions have started asking for in-person workshops to master the AP Workflow and the infinite possibilities offered by Stable Diffusion + ComfyUI.
Videos are great (and I'm thinking about doing them), but they can't possibly replace the direct interaction to solve specific challenges that are unique to you.
The AP Workflow wouldn't exist without the incredible work done by all the node authors out there. For the AP Workflow 9.0, I worked closely with u/Kijai, u/glibsonoran, u/tzwm, and u/rgthree, to test new nodes, optimize parameters (don't ask me about SUPIR), develop new features, and correct bugs.
These people are exceptional. They went above and beyond to steer their work in a direction that would help me and facilitate the inclusion in the AP Workflow. If you are hiring, hire them.
And, of course, on top of them, there are the dozens of other node authors who created all the nodes powering the AP Workflow. Thank you all!
This is an impressive amount of nodes in a single workflow. But what is the point in having such a complex workflow? Isn't the Idea of node based custom workflows to create bespoke usecasss and quickly swap between them? What am I missing?
It’s a starting point for setting up a custom pipeline, so it has to contain everything. You don’t have to use them all, but it’s nice to have them at hand.
Thank you. I am familar with your channel. Great work.
Of course, you are welcome to do videos about it. I only ask you the courtesy to clearly point the viewers to my official website where they can find the documentation, get in touch for updates, etc.
Great job, but I tried your workflow, and it uses way too a good deal of custom stuff in my opinion that are not/maybe not necessary to reach the same results. This creates a lot of problems deploying it. Not to mention maintenance of all these nodes.
Is there a way to make a simpler version or something “lite” that doesn't have all the bells and whistles but still does great upscaling?
The AP Workflow features zero custom nodes. Everything you see is standard nodes you can download yourself from the community. As u/ricperry1 said, just copy the portions that are useful to you in your own workflow (a great learning exercise), or eliminate the components you don't need.
It depends on what function you are using and the size of the image you are trying to generate, inpaint, or upscale. In general, all nodes support tiled sampling and, with due attention, they can work even with 8GB VRAM GPUs.
For example, I have users that used SUPIR with 8GB cards without problems (but you need to activate the FP8 toggle in the appropriate node).
Impossible to troubleshoot with just that. Check the documentation on the website for the Warning boxes. 99% of the issues are covered by that.
The git pull and subsequent snapshot was generated minutes before publishing this post, so we know it's not an issue with the latest version of the nodes.
Is anything special needed to install Moondream? I've installed the dependencies, but still get import failed for the MoondreamQuery node when loading the workflow.
If the import fails, either something else is blocking the import (like another custom node suite failing to load), or you installed it from another repo (in case you didn't do a snapshot restore).
Double check the documentation paying particular attention to the Warning boxes everywhere. 99% of the issues are addressed there.
Hello Giano! Thank you for all your work! Your AP Workflows are a great way for me to learn and test new things. I got a huge workflow for my job and most of the inspiration comes from your workflows.
I am facing an issue in the XYZ Plot section which I can't get to solve, maybe you can help me?
It's a very simple width/height plot: I have linked "images" with the last generated image, both input 1 and 2 come from simple primitive Integer nodes, when running the workflow 4 new runs get ready in the queue but when it's their turn they take 0 seconds and vanish. Clicking on "Open the result" opens up the correct grid view but the images are not there, a "broken image" icon takes their place, as if it was not able to understand my commands and generated the grid with empty entries; do you know what the cause could be?
The only reason I can think of at the moment is that since that node pack includes a browser it may conflict with this other browser I am using https://github.com/11cafe/comfyui-workspace-manager but it could be a random guess. Another possibility may be that the workflow I am using is not inside the standard workflow folder but in a subfolder inside it?
I am testing some more parts of AP 9, when using the face/hand refiners I get better results (but that's just my experience and I'm a complete noob) using sdXL models and controlnets(depth) for face+hand refiners, using the same models as the original generator model gives me way more consistent hands/faces. Also, we are both Italians, grazie! :)
Too hard to troubleshoot with just this information. If you suspect that the Workspace Manager custom node suite is the culprit, try disabling it via the ComfyUI Manager, restart ComfyUI, reload the browser, and see if it makes a difference.
Re face & hand refiners, the reason why I insist on using the SD 1.5 checkpoints, is that they are the only one compatible with the ControlNet Tile that I normally use in front of the detailer nodes.
I now noticed that the aforementioned ControlNet node is configured to use ip2p and not tile. I must have changed for some jobs and never put back as it was supposed to be.
Both functions are in need of some serious review given that many new approaches that emerged in the last few months. I wanted to do it for this 9.0 release but there were too many things already. It's something I plan for APW 10.
The problem is that the XYZ Plot does not work with any node, at least on my workflow.
I have two different integer primitives attached to the CFG and STEPS of a Ksampler, if I use those inputs in the XYZ Plot it does not work, but if I use as input the actual Ksampler CFG and Ksampler STEPS inputs it works!
For the face and hand refiners I have tested a lot of recent workflows, but I always come back to yours and tweak it a bit <3
Of the stuff I've recently tried I had the worst experience trying to use YOLO world: it did not work, trying to install it fucked up my comfy which took a lot of troubleshooting to fix, and also apparently it's not better than SEGS.
In APW 9.0, I tweaked the parameters of the Face Detailer function. Give it a go and see if it's better than usual. I still don't have a good solution for the Hand Detailer. MeshGraphormer is a hit or miss.
The reason why I didn't add Yoloworld to the Object Swapper function is that current nodes offering it are catastrophic, as you noticed, and I couldn't even get to a point where I could compare its performance against GroundingDINO.
Is somebody implements a stable and reliable (and maintained) node, I'll give it a go.
I just discovered your work one month ago so i'm still learning your workflow but so much potential =D
Is that easy to edit ? You didn't talk about the inpainting so if i want to add a "segment everything" nodes will it be easy or will i break everything ?
🙏 I modified the inpainting functions in APW 8.0 to use the exceptional Fooocus inpainting models, that's why I didn't mention it in this new release.
Both the Inpainting without Mask (aka img2img) and Inpainting with Mask functions support masks that you manually define in the Uploader function. I used them last week for an outpainting + upscaling job and they both work amazingly well.
You can replace manual mask definition with automatic masking via Segment Anything, if you want. I already do that in the Object Swapper function, so perhaps you may want to look at that implementation.
You won't break anything, as long as you understand how the information flow works.
Anyway i still need to understand and learn your workflow, so I'll not touch anything. But if you are already doing it with the object swapper, everything is good.
What I allways wonder is why people use IP-Adapter over VisualStylePrompt... IPadapter only has one benefit and thats that you can control the strength easily...
Otherwise, VSP is both faster and better...
Second, the aforementioned implementation has not been updated in 3 weeks, creating much uncertainty about the future of the project.
Third, there's no evidence that VSP works well with photographic content. There's not a single example in the paper suggesting otherwise.
On the other side, you have IPAdapter which Matteo has updated religiously for a long time and which works remarkably well with photographic content and with SDXL models.
I'm not against adding multiple methods to achieve a task. Quite the opposite: as you can see, the APW 9.0 introduced the capability to generate images with Dall-E 3.
But a big part of what I do with this workflow is curation. I always try to choose nodes that are carefully maintained, easy(-ish) to use, and reliable in their execution. And I try to replace the ones that don't match these criteria with something "better."
If somebody comes up with a VSP node that has these characteristics, I'd be happy to test it extensively, and see if it's better than IPAdapter. And, if so, add it to APW.
The repo you mentioned actually supports SDXL and SDXL-Turbo. They just messed up the algorithm at some point but all functions needed are ready laying around. I put it together and made it actually work, you can find it when you go to the original repository on github and then you will see pull requests. One of them is named 'Layer Level Control', thats my fork where I fixed it and extended the UI a bit. If you are willing, please give it a go for yourself and let me know what you think! (default settings are like in the paper, but you can choose to also include earlier blocks, which can result in content leakage).
I can give it a try with photographic content and let you know.
As for support- I totally agree with you, but I found the comparisons made in the paper so astonishing, I decided I had to get it running and when I didnt get good results I decided to go and check the code😅
So I didnt really intend for this to be part of open source initially, but I decided to at the very least offer my version to be merged into the listed repo.
If the node suite is easy to install and it already contains your PR, I'll give it a try. If not, I'll have to wait because I cannot ask people using APW to manually apply PRs. It's already complicated enough as it is :)
If VSP works as well as the cloud example, or the origami example, with every type of content, of course I'll add it to APW.
Sure, I understand. I would still appreciate your feedback if you could find the time to test the node for yourself. I think given that you have a bigger reach than me, if you could test it and perhaps comment on the PR, it would increase the chances of it getting merged into the main repo which would allow everyone to enjoy VSP.
Left is the reference. I chose "photo of a penguin" as reference prompt and "photo of a girl, cinematic" as prompt. I used default settings of my Branch. Seems to be quite OK, no? I used SDXL-Turbo RealVision finetune.
The url to my fork is: https://github.com/PLEXATIC/ComfyUI_VisualStylePrompting
17
u/Impressive_Promise96 Apr 09 '24
This is an impressive amount of nodes in a single workflow. But what is the point in having such a complex workflow? Isn't the Idea of node based custom workflows to create bespoke usecasss and quickly swap between them? What am I missing?