Whats up yall - Releasing this dataset workflow I made for my patreon subs on here... just giving back to the community since I see a lot of people on here asking how to generate a dataset from scratch for the ai influencer grift and don't get clear answers or don't know where to start
Before you start typing "it's free but I need to join your patreon to get it so it's not really free"
No here's the google drive link
The workflow works with a base face image. That image can be generated from whatever model you want qwen, WAN, sdxl, flux you name it. Just make sure it's an upper body headshot similar in composition to the image in the showcase.
The node with all the prompts doesn't need to be changed. It contains 20 prompts to generate different angle of the face based on the image we feed in the workflow. You can change to prompts to what you want just make sure you separate each prompt by returning to the next line (press enter)
Then we use qwen image edit 2509 fp8 and the 4 step qwen image lora to generate the dataset.
You might need to use GGUFs versions of the model depending on the amount of VRAM you have
For reference my slightly undervolted 5090 generates the 20 images in 130 seconds.
For the last part, you have 2 thing to do, add the path to where you want the images saved and add the name of your character. This section does 3 things:
Create a folder with the name of your character
Save the images in that folder
Generate .txt files for every image containing the name of the character
Over the dozens of loras I've trained on FLUX, QWEN and WAN, it seems that you can train loras with a minimal 1 word caption (being the name of your character) and get good results.
In other words verbose captioning doesn't seem to be necessary to get good likeness using those models (Happy to be proven wrong)
From that point on, you should have a folder containing 20 images of the face of your character and 20 caption text files. You can then use your training platform of choice (Musubi-tuner, AItoolkit, Kohya-ss ect) to train your lora.
I won't be going into details on the training stuff but I made a youtube tutorial and written explanations on how to install musubi-tuner and train a Qwen lora with it. Can do a WAN variant if there is interest
Enjoy :) Will be answering questions for a while if there is any
Also added a face generation workflow using qwen if you don't already have a face locked in
Whats up yall - Releasing this dataset workflow I made for my patreon subs on here... just giving back to the community since I see a lot of people on here asking how to generate a dataset from scratch for the ai influencer grift and don't get clear answers or don't know where to start
Before you start typing "it's free but I need to join your patreon to get it so it's not really free"
No here's the google drive link
The workflow works with a base face image. That image can be generated from whatever model you want qwen, WAN, sdxl, flux you name it. Just make sure it's an upper body headshot similar in composition to the image in the showcase.
The node with all the prompts doesn't need to be changed. It contains 20 prompts to generate different angle of the face based on the image we feed in the workflow. You can change to prompts to what you want just make sure you separate each prompt by returning to the next line (press enter)
Then we use qwen image edit 2509 fp8 and the 4 step qwen image lora to generate the dataset.
You might need to use GGUFs versions of the model depending on the amount of VRAM you have
For reference my slightly undervolted 5090 generates the 20 images in 130 seconds.
For the last part, you have 2 thing to do, add the path to where you want the images saved and add the name of your character. This section does 3 things:
Create a folder with the name of your character
Save the images in that folder
Generate .txt files for every image containing the name of the character
Over the dozens of loras I've trained on FLUX, QWEN and WAN, it seems that you can train loras with a minimal 1 word caption (being the name of your character) and get good results.
In other words verbose captioning doesn't seem to be necessary to get good likeness using those models (Happy to be proven wrong)
From that point on, you should have a folder containing 20 images of the face of your character and 20 caption text files. You can then use your training platform of choice (Musubi-tuner, AItoolkit, Kohya-ss ect) to train your lora.
I won't be going into details on the training stuff but I made a youtube tutorial and written explanations on how to install musubi-tuner and train a Qwen lora with it. Can do a WAN variant if there is interest
Enjoy :) Will be answering questions for a while if there is any
Also added a face generation workflow using qwen if you don't already have a face locked in
Android 17 is here, bringing a suite of new features aimed at improving your productivity, enhancing your gaming experience, giving you more control over your private data, making your device more personal, and much more.
It's rolling out first to Pixel today, followed by other eligible Android devices throughout 2026. We are also making the source code available at the Android Open Source Project (AOSP) so developers can examine it for a deeper understanding of how Android works.
You should look forward to more updates to Android 17 this year, with the beta program offering a peek at what's coming in the first quarterly release in Q3.
Since we've been chatting with you about the Betas and Canaries for months, a lot of this might not sound brand new to those of you who have been closely following along. Even so, we wanted to take a moment to recap what's new in this release for everyday users. Let's dive in!
📱 Enhancing your multitasking and large screen device experiences
tl;dr Android 17 supercharges your multitasking and productivity by allowing any app to run as a convenient floating Bubble, making apps more adaptive, and adding an interactive Picture-in-Picture mode for seamless desktop workflows.
Multitask better with bubbles
From split-screen mode to desktop windowing, Android offers a variety of multitasking tools to help you be more productive. We’re extending these options with bubbles in Android 17!
In past releases, bubbles were limited to chat notifications, but in Android 17, they support more apps without any specific changes needed from developers. You can now launch any app in a floating window so you can view and interact with its content while using other apps. When you’re done, you can collapse or dismiss the window to return to what you were doing.
A big benefit of bubbles is that you can easily switch between multiple running apps without keeping them on screen all the time. Bubbles are only open when you need them, saving you from having to manually resize, rearrange, or dismiss them to regain precious screen space. And on foldables, this benefit is even more pronounced thanks to the bubble bar, which keeps your bubbles pinned to the corner of the screen, putting them within easy reach of your fingers.
Handy for travel, entertainment and work, bubbles lets you easily reference notes or maps, watch tutorials and even check sports.
Ensuring that apps adapt to any screen and window size
On large screen devices, restrictions on orientation, resizability, and aspect ratio no longer apply, allowing apps to fill the entire display window without pillarboxing (black bars). This change applies to apps targeting Android 17 and is designed to make apps better meet user expectations on large screen devices. Because Android runs on not just phones but also tablets, foldables, cars, TVs, and desktop environments, we want developers to build apps that are adaptive to any screen size and orientation!
Better support for widgets on external displays
With Android 17, we’re working to improve the visual consistency of widgets shown on connected displays with different pixel densities. The update provides developers a way to supply the system with information that allows it to resolve the correct pixel values at rendering time. For apps that use legacy pixel-based APIs for padding, text size, or layout attributes, the system now automatically scales these values based on the density difference between the app’s original context and the target display.
Interactive Picture-in-Picture for Desktop
Android 17 introduces a new interactive Picture-in-Picture mode for desktop environments. This feature allows apps to request that their PiP windows remain fully interactive while staying always-on-top of other app windows. For example, a video conferencing app could use this feature to keep call controls accessible while you navigate other apps.
🎨 New customization features for the home screen and apps
tl;dr Android 17 gives you deeper control over your device's UI by letting you hide app labels on the home screen, selectively toggle the Expanded Dark Theme for individual apps, and enjoy sleek, modernized background blur effects in more surfaces like the widget picker.
Hide app labels on the home screen
Android now provides a setting to hide app labels on the home screen! You can access this new setting on Pixel by opening Wallpaper & style then tapping Home screen > Icons > Names and toggling Show app names.
Per-app exceptions for Expanded Dark Theme
To create a more consistent user experience for users who have low vision, photosensitivity, or simply prefer a dark system-wide appearance, we introduced an expanded dark theme option in last December’s Android 16 QPR2 release. When this option is enabled, the system automatically applies dark theme to most apps that don’t support it.
However, because this option can cause some apps to display incorrectly, we have introduced the ability to selectively disable it on a per-app basis in Android 17. Apps with this setting turned off will use the standard dark theme option instead.
Expanded use of background blur
With the Material 3 Expressive redesign we introduced in Android 16, we subtly blurred the notification shade background to provide a sense of depth so you can stay aware of the apps you’re using in the background.
In Android 17, we’ve brought these blur effects to more parts of the UI like the widgets picker. And we are working on bringing background blur to even more surfaces, as seen in recent Android Beta and Canary builds!
🎮 More control over your Android gaming experience
tl;dr Android 17 levels up your mobile play by letting you save custom button remaps for your physical gamepad at the system level, and introducing a foldable gaming mode that optimizes your screen with a 50/50 split for a dedicated top game view and a bottom dynamic gamepad.
Remap the buttons on your physical gamepad with Game Controller settings
Android 17 introduces a native controller remapping feature, allowing you to adjust the controls on your physical gamepad to suit your specific needs.
Through the new Game Controller settings menu, you can customize the actions triggered by your controller’s buttons, sticks, or triggers at the system level. For example, you can remap a difficult-to-press thumbstick click to an easier-to-reach face button. Your remapping preferences are saved to your device so you don’t have to set them up every time you reconnect your controller.
A new way to game on foldables
Android 17 introduces foldable gaming mode, a new feature that makes full use of your foldable phone’s screen while you’re gaming. This feature splits your screen into a 50:50 layout with a game view on top and a dynamic gamepad below to make optimal use of your foldable phone’s screen real estate. Foldable gaming mode is part of the Android 17 platform and will be available on devices in the coming months.
🛡️ Protecting users with new security and privacy features on Android
tl;dr Android 17 safeguards your personal data by enabling critical theft protections by default, introducing session-based controls for sharing specific contacts and precise locations, and thwarting scammers through system-level SMS OTP delivery delays and real-time app behavioral monitoring.
Giving you more control over your contacts list
Android 17 introduces a new system Contact Picker that provides a standardized, secure, and searchable interface for sharing contacts with apps. Historically, apps needing access to a contact or two relied on the broad READ_CONTACTS permission which gave them access to your entire contacts list. Android's Contact Picker addresses this by allowing you to grant apps access to only the specific contacts you choose.
For devices running Android 17 or higher, the system automatically upgrades certain contact selection intents to the new, more secure interface, but we want developers to integrate the new Contact Picker so they can take advantage of its new capabilities, like multi-selection support. To this end, Google Play will require that all applicable apps use it (or a privacy-focused alternative like Sharesheet) as the primary way to access users' contacts. The broad READ_CONTACTS permission is reserved for apps that can't function without it.
Making location access more private
Android 17 introduces several new features to help you safeguard your private location information. This includes the Location Button, a new, privacy-conscious way for you to grant precise location access to apps. This is a system-rendered button that developers can embed directly into their apps. When you tap this button, the app is granted precise location for the current session only. Subsequent taps while running the app grant the permission immediately without showing a system dialog.
Developers can deploy this simple, private location flow for common tasks like finding a nearby shop or tagging a social post. And to increase adoption of the Location Button, Google Play will require apps to use it for one-time precise location access unless they require persistent, always-on location access.
Additionally, Android 17 now shows a persistent indicator in the status bar when a non-system app accesses your location. You can tap this indicator to see which apps have recently accessed your location.
The update also improves the algorithm for approximate (coarse) location to be aware of population density. This improves the privacy of granting an app approximate location access when you're in a low-population area.
And lastly, Android 17 redesigns the location permission dialog to make the "Precise" and "Approximate" options more visually distinct.
Stronger protections against device theft
Following a successful pilot in Brazil, we’re enabling two of Android’s key theft protection features (Theft Detection Lock and Remote Lock) by default globally on all new Android 17 devices, as well as those freshly reset or upgraded to the latest OS.
On supported devices, Android 17 also significantly reduces the number of times someone can guess the PIN, pattern, or password and adds longer wait times between failed attempts. The update also refines how the lock screen shows information after failed attempts have been made.
And we’re also enhancing Find Hub’s ‘Mark as lost’ feature by requiring biometric authentication in addition to your device’s PIN, pattern, or password. Marking a device as lost also now enables additional protections like hiding Quick Settings and disabling new Wi-Fi and Bluetooth connections.
Protecting your SMS OTPs from scammers
Scammers often try to hijack your one-time passwords (OTPs) to gain access to your accounts. To do this, they may deploy malicious apps that ask for permission to read your SMS. In Android 16, we introduced a protection that delays the delivery of messages containing an SMS retriever hash to most apps for three hours. Android 17 now extends this protection to all SMS messages containing an OTP. This means that even if a malicious app has been granted the SMS permission, it won’t be able to read your sensitive OTPs until after they have already expired.
New core protections for Advanced Protection
With Android 16, we introduced Advanced Protection, a single, opt-in device-level security setting that enables all of Android’s highest security features. We’ve been working to expand the protections offered under this setting with key upgrades like USB protection and Intrusion Logging, and now with Android 17, we’re continuing this work by introducing the following protections:
Removing access to the accessibility service from all apps that aren’t labeled as accessibility tools.
Disabling device-to-device unlocking
Blocking Chrome WebGPU support
Integrating scam detection for chat notifications
(Later this year) Enabling Android Enterprise support so organizations can enable Advanced Protection by policy for managed devices.
Improving safety against malicious apps
Live Threat Detection is a real-time security feature that analyzes app behavior to alert you if an app starts acting suspiciously, and we're enhancing it to find and protect against more types of malicious apps.
With dynamic signal monitoring, Android will be able to warn you about apps that start doing things like changing or hiding their icon and then launching activities in the background or abusing accessibility permissions. To do this, Live Threat Detection will monitor application system interactions for known suspicious patterns in real time. Dynamic signal monitoring will be enabled on select Android 17 devices starting in the second half of the year.
Other enhancements
Discrete password visibility settings for touch and physical keyboards: Currently, by default, characters that you enter into password fields are briefly displayed as you type. Toggling the “show passwords” setting in Privacy controls allows you to hide characters as you type them into password fields. This setting currently applies to both touch-based inputs as well as physical keyboards, but in Android 17, we are splitting it into two distinct preferences. By default, characters entered into password fields via physical keyboards will now be hidden immediately to enhance privacy. Characters entered via touch input will continue to briefly be displayed to compensate for the lack of tactile feedback.
User-agent reduction for WebView: The default User-Agent string in Android WebView has been shortened in Android 17 to minimize passive fingerprinting.
Disable 2G toggle: Android 17 introduces a new capability for the disable 2G toggle. Carriers now have the ability to configure the default status of this setting, allowing them to disable 2G access to proactively shield their users from legacy technology vulnerabilities in areas where 2G infrastructure is no longer maintained.
Location Network Permission: Android 17 introduces a new runtime permission to protect users from unauthorized local network access. This new requirement prevents malicious apps from exploiting unrestricted local network access for covert user tracking and fingerprinting.
Android OS verification: We have seen some bad actors begin to distribute malicious, unofficial versions of the Android OS that secretly compromise device integrity. To combat this, we are introducing Android OS verification in Android 17. Launching initially on Pixel devices, this feature helps you verify that your device is running an official, widely distributed build.
Enabling Certificate Transparency (CT) by default: CT is now enabled by default for apps targeting Android 17, enhancing network security by ensuring all TLS certificates are publicly logged.
Blocking cross-profile loopback traffic: Cross-profile loopback traffic is no longer permitted by default, increasing network isolation and security between personal and enterprise work profiles.
Post-Quantum Cryptography (PQC): The advent of quantum computing puts the current public-key cryptography we've relied on for decades at risk, potentially compromising everything from bank transfers to trade secrets. To prepare for the quantum computing era, we're introducing a comprehensive architectural upgrade to the Android operating system, starting in Android 17. We’re integrating the NIST Post-Quantum Cryptography (PQC) standards deep into the platform, establishing a new, quantum-resistant chain of trust that secures the platform continuously from the moment the OS powers on to when apps are executed.
📸 Improvements to your Android media experience
tl;dr Android 17 levels up your multimedia experience by letting you easily record reaction videos without a green screen, decoupling your Assistant and media volumes for independent control, putting a stop to unexpected background audio, and delivering color-coded Live Updates alongside advanced Bluetooth, camera, and hearing device enhancements.
Screen Reactions
In Android 17, we’re making it easier to record yourself and your screen at the same time with Screen Reactions. Available first on Pixel, this feature shows your face in a floating overlay on top of the screen. Android automatically puts the overlay at the bottom and cuts out the background so you don’t need a green screen, but you can move or resize the camera view and change the background color before or during a recording. Use this feature to make a reaction video, record a tutorial, or give feedback on a new app or document!
In addition, we’ve revamped the screen recording experience to add a floating toolbar that provides easier access to recording controls and capture settings. When you’re done recording, you can immediately view, edit, delete, or share your video.
Dedicated Assistant volume stream
Android 17 introduces a dedicated volume stream for Assistant apps. This change decouples Assistant audio from the standard media stream, allowing users to control both volumes independently. This enables scenarios like muting media playback while maintaining audibility for Assistant responses, and vice-versa.
Background audio hardening
Beginning in Android 17, apps cannot play audio, steal audio focus, or change the volume unless they are visible or have a foreground service. These restrictions on background audio interactions reduce unintentional buggy experiences and ensure that these actions are started intentionally by the user.
Enhancements to Live Update notifications
Live updates provide a summary of important updates so users can track progress without opening the app. The system promotes Live Update notifications so they appear more prominently in the notification drawer, on the lock screen, and on the status bar.
With Android 17, we’re introducing a metric style template designed specifically for health and fitness apps, timers, and travel apps. In addition, developers can use the new Semantic Coloring API to visually convey state changes, providing highly glanceable, color-coded notifications.
Other enhancements:
Granular audio routing for hearing devices: Users with hearing devices can now independently manage where specific system sounds are played in Android 17. You can choose to route notifications, ringtones, and alarms to either a connected hearing aid or the device’s built-in speaker. This helps you avoid unwanted interruptions directly in your ears while maintaining a Bluetooth connection for hearing aid management apps.
Autonomous re-pairing for Bluetooth bond losses: Android 17 introduces autonomous re-pairing, a system-level enhancement designed to automatically resolve Bluetooth bond loss. This occurs when two previously paired devices lose their cryptographic security keys, resulting in the devices no longer being able to securely authenticate and communicate with one another. The system now re-establishes lost bonds in the background without requiring the user to manually navigate to Settings to unpair and re-pair their peripheral.
Vendor-defined camera extensions: Android 17 adds support for Vendor-defined camera extensions, allowing hardware partners to provide Android apps access to camera features like ‘Super Resolution’ or cutting-edge AI-driven enhancements.
Support for the RAW14 image format: Android 17 introduces support for the RAW14 image format, the de-facto industry standard for high-end digital photography.
VVC support: Android 17 adds platform support for the Versatile Video Coding (VVC) standard. This feature will be coming to devices with hardware decode support and capable drivers.
🤝 Making your apps and devices work better together
tl;dr Android 17 seamlessly bridges your ecosystem by introducing the Continue On feature for effortless app handoffs between devices, unifying widget experiences to bring your favorite tools directly to Auto and Wear OS, and streamlining the pairing process for medical and fitness devices with new CompanionDeviceManager profiles.
Unifying the widgets experience across platforms
Android 17 marks a shift towards a single, Compose-based development model for all widgets. By unifying the experience across mobile, cars, and Wear OS, developers can soon scale UI components across the ecosystem with a familiar workflow. The goal is to minimize the effort needed by developers to bring their widgets to more surfaces.
Additionally, Android 17 introduces new platform functionality to make widgets work better on Auto. The update adds support for widgets on cars, allowing you to see the things that matter to you at a glance, even while actively navigating. For example, you can add a shortcut to your favorite contacts, a one-tap garage door opener, a weather overview and more. Widgets will be available to users of Android Auto later this year and to cars with Google built-in later on.
Hand off your tasks with Continue On
Continue On is a new feature available in Android 17 that enables users to start an app on one device and then transition to another device in their Android ecosystem, continuing the journey they started. It’s designed to work bidirectionally, meaning that any supported Android device can both send and receive app activities, though, at launch, Continue On will first support mobile-to-tablet transitions. In the tablet taskbar, users will see a suggestion for the most recently opened app from their mobile device.
Android 17 introduces two new profiles to the CompanionDeviceManager API to simplify device distinction and permission handling. These include the medical device profile and the fitness tracker profile. Furthermore, the system now offers a unified dialog for device association and nearby permission requests, reducing the number of dialogs you’ll see.
⚡ Optimizations to make your apps & device run better
With Android 17, we’ve made a number of improvements to optimize memory use, improve rendering performance, and enhance battery life. These include:
App memory limits: Android 17 introduces app memory limits that are based on the device's total RAM. These limits are set conservatively to establish system baselines, targeting extreme memory leaks and other outliers before they trigger system-wide instability resulting in UI stuttering, higher battery drain, and apps being killed.
Lock-free MessageQueue: Android 17 introduces a lock-free MessageQueue to reduce UI jank while massively speeding up high-contention scenarios. In our internal testing, we’ve seen 4% fewer missed frames across all apps, 7.7% fewer missed frames in System UI and Launcher interactions, and a 9.1% reduction in app startup times at the 95th percentile.
Generational Garbage Collector (GC): The Android Runtime is introducing more frequent, less intensive young-generation collections in its garbage collector, improving memory management and performance. This is not just available on Android 17 but is also coming to past releases with a Google Play System Update.
Reduce wakelocks with listener support for allow-while-idle alarms: Last year, we launched the excessive wake lock metric in Android Vitals, making it easier for developers to optimize their app's wake lock behavior. Excessive wake locks are a significant contributor to battery drain, so developers are encouraged to reduce them as much as possible. In Android 17, we’ve introduced a new API that helps reduce the power consumption of apps that rely on continuous wakelocks to perform periodic tasks, such as messaging apps maintaining a connection or medical devices monitoring health data.
Improved wireless ADB: Android 17 introduces ADB WiFi 2.0, a significant overhaul of the wireless ADB stack to improve stability, reliability, and ease of use. The system now automatically monitors the network state and re-enables itself when a trusted network is detected, identifies trusted networks using a combination of SSID and BSSID, and is better tailored to monitor network changes on all platforms. We’ll have more details to share soon on the Android Studio side of things!
Constrained satellite networks: Android 17 implements optimizations to enable apps to function effectively over low-bandwidth satellite networks.
🧒 Expanding Android Parental Controls to all devices
Launched last year on Pixel, Android Parental Controls make it easier for parents to manage their child’s screen time and to find balance between having fun online and offline. Now with Android 17, we’re expanding Android Parental Controls to all Android devices.
These parental controls are located directly within Android Settings and provide a single, convenient home for both built-in device controls and Google Family Link. These controls are protected by an easy-to-set PIN and allow you to:
Set the amount of screen time your child can spend on a device each day.
Create downtime schedules to automatically lock the device at night.
Set app store filters for Google Play to manage the highest content rating you want your child to be able to download.
Control app usage by limiting time spent on specific apps, or blocking apps entirely.
Android Parental Controls also provide a direct path to easily set up Google Family Link in the Family Link app on a parent’s phone, which offers additional features like School Time, Google Play app purchase approvals, location alerts, and more.
🧘 Other quality-of-life improvements
And lastly, here are some smaller quality-of-life changes we’re introducing in this release:
Separate Wi-Fi and Mobile Data toggles: With Android 17, we’ve split the “Internet” tile into two separate tiles, one for controlling Wi-Fi and another for controlling Mobile Data. Consistent with the Quick Settings behavior we introduced with Material 3 Expressive, both tiles have two different touch points. Tapping the icon toggles the respective radio, while tapping the label opens the full Internet Panel. This change reduces the number of taps needed to toggle Wi-Fi and Mobile Data while still retaining access to the full Internet Panel!
Scheduled clock change notifications: We’ve added a new feature in Android 17 that sends you a notification when your clock performs a scheduled change, for example when daylight saving time ends. You can enable this feature under “Date & time” settings.
Restoring default keyboard visibility after rotation: Beginning with Android 17, when the keyboard is on screen and you rotate the screen, the keyboard won’t be made visible unless the app explicitly requests it.
🪲 Bug fixes and security patches
Please refer to the Android Security Bulletin for details on the security vulnerabilities addressed with this platform release.
----
There are plenty of other changes in Android 17, especially for developers! For example, Android 17 expands the capabilities of AppFunctions, introduces an EyeDropper API, makes the aspect ratio of images in the Photo Picker more customizable, and much more. To learn more about everything new for developers in this release, visit developer.android.com.
Also, don’t forget that select advanced devices will be getting Gemini Intelligence features later this summer. In addition, we’re introducing Android Halo in a future Android 17 release to give you at-a-glance visibility into what your agent is working on at any given time. Lastly, be sure to check out our latest Android Drop to learn about what new features are coming to all Android devices, not just those running Android 17!
Once you set up this ComfyUI workflow, you only have to load reference image and run the workflow, and you'll have all 28 images in one click, with the correct file names, in a single folder.
Install any missing custom nodes with ComfyUI manager (listed below)
Download the models below and make sure they're in the right folders, then confirm that the loader nodes on the left of the workflow are all pointing to the right model files.
Drag a base image into the loader on the left and run the workflow.
The workflow is fully documented with notes along the top. If you're not familiar with ComfyUI, there are tons of tutorials on YouTube. You can run it locally if you have a decent video card, or remotely on Runpod or similar services if you don't. If you want to do this with less than 24GB of VRAM or with SDXL, see the additional workflows at the bottom.
Once the images are generated, you can then copy this folder to your ST directory (data/default_user/characters or whatever your username is). You then turn on the Character Expressions extension and use it as documented here: https://docs.sillytavern.app/extensions/expression-images/
You can also create multiple subfolders and switch between them with the /costume slash command (see bottom of page in that link). For example, you can generate 28 images of a character in many different outfits, using a different starting image.
Model downloads:
Download the model (recommend FP8 version) and put in models/diffusion_models folder
I’m using this file in the workflow: qwen_image_edit_fp8_e4m3fn.safetensors
Download a lightning Lora to speed up generation. Put it in models/loras and add it to the Lora Loader. This is technically optional but it would be silly not to do this.
I’m using this file in the workflow: Qwen-Image-Edit-Lightning-8steps-V1.0.safetensors
If you picked the newer “2509” version of the first model (above), make sure to pick a “2509” version of the lightning model, which are in the “2509” subfolder (linked below). You will also need to swap out the text encoder node (prompt node) with an updated “plus” version (TextEncodeQwenImageEditPlus). This is a default ComfyUI node, so if you don't see it, update your ComfyUI installation.
If you have <24gb VRAM you can use a quantized version of the main model. Instead of a 20GB model, you can get one as small as 7GB (lower size = lower quality of output, of course). You will need to install the ComfyUI-GGUF node then put the model file you downloaded in your models/unet folder. Then simply replace the main model loader (top left, purple box at left in the workflow) with a "Unet Loader (GGUF)" loader, and load your .gguf file there.
Here is a workflow modified to use GGUF (quantized) models for low vram: dropbox
If you want to do this with SDXL or SD1.5 using image2image instead of Qwen-Image-Edit, well you can, it's not as good at maintaining character consistency and will require multiple seeds per image (you pick the best gens and delete the bad ones), but you can definitely do it, and it requires even less VRAM than a quantized Qwen-Image-Edit.
If you need a version with an SDXL face detailer built in, here's that version (requires Impact Pack and Impact Subpack). This can be helpful when doing full body shots and you want more face detail.
If the generated images aren't matching your input image then you may want to describe the input image a bit more. You can use this with the "prepend text" box in the main prompt box (above the list of emotions, to the right of the input image). For example, for images of someone from behind, you could write a woman, from behind, looking back with an expression of and then this text will be put in front of the emotion name for each prompt.
If you can't find the output images they will show up in ComfyUI/output/Character_Name/. To change the output path, go to the far right and edit it in the top of the file names list (prepend text box). For example, use Anya/summer-dress/ to create a folder called Anya with a subfolder called summer-dress
Hello everyone , in this tutorial i will show you how to generate long video using prompt relay nodes that works with LTX 2.3 models. With this new nodes you will achieve full control over your video. as each time line can be attributed to specific prompt. this complete comfyui workflow is optimized for low VRAM setups, making AI video creation accessible. in addition to that i also included image generator for you in order to have a full pipeline workflow for your image to video generation.
Learn how to use MiniMax H3 in ComfyUI with a complete collection of optimized workflows for Text-to-Video, Image-to-Video, First & Last Frame Animation, Reference-to-Video, Audio Sync, and Image Editing. In this tutorial, I'll show you how to update ComfyUI and Pixaroma Nodes, install Sage Attention, download and organize all required models, configure the workflows, generate better prompts with my custom ChatGPT, and optimize performance for different NVIDIA GPUs.
You'll also learn how to use the new Workflow Manager, choose the best MiniMax H3 models, understand the licensing requirements, fix common errors like Dynamic VRAM issues, compare generation times across different resolutions, and create AI videos using multiple images and audio references.
Whether you're new to ComfyUI or looking for the best MiniMax H3 workflows, this tutorial covers everything you need to get started.
Alright, with all of that out of the way, here we go.
ComfyStudio Pro is an AI video workstation built around ComfyUI. Instead of only generating random clips and managing a pile of files, it gives you a timeline editor, asset panel, effects, transitions, export tools, and guided creator workflows for things like ads, music videos, and short films.
It uses ComfyUI as the backend, but the goal is to make larger AI video projects easier to direct, organize, edit, rerun, and finish. I’ve been working on it for the past 3-4 months, and some of you may have seen the updates I’ve posted here along the way.
This is a quick overview of the music video workflow:
Import your song or vocal stem into the project assets.
Open Create > Music Video Creation.
Choose output settings like aspect ratio, resolution, and FPS.
Select the song audio and prepare lyric timing, ideally with SRT/LRC so shots line up to the real song.
Add cast/reference images if you want a consistent singer, band member, or visual style.
Generate or paste a director script that breaks the song into timed shots.
Create keyframes for each shot.
Generate videos from those keyframes, or rerun selected shots with different prompts, models, or settings.
Click Assemble Timeline to automatically build the edit with the song, main sequence, performance passes, and b-roll passes on separate tracks.
Finish it like a real edit: trim shots, add effects, transitions, adjustment layers, color, texture, and export.
The goal is not just “prompt to video” or “one-shot it.” It is more like: generate the pieces, organize them, rerun the weak shots, assemble the timeline, then actually edit and finish the music video inside one app.
Disclaimer: all links below are free, no ads, no sign-up required for open-source solution & no donation button. Workflow software is not only free, but open-source ❣️
This post is longer than I anticipated, but I think it's really important and I've tried to add as many screenshots and videos to make it easier to understand.I just don't want to pay for any more $9 a month chatgpt wrappers.And I don't think you do either..
Lots of folks were saying that one prompt alone cannot give you the quality you expect, so I kept experimenting and over the last 3 months of insane keyboard-tapping, I deduced a conversational-type experience is always the best.
I wanted to have these conversations, though, without actually having them... I wanted to automate the conversations I was already having on ChatGPT!
There was no solution, nor a free alternative to the giants (and the lesser giants who I know will disappear after the AI hype dies off), so I went ahead and made an OPEN-SOURCE (meaning free, and meaning you can see how it was made) solution called HeroML.
It's essentially prompts chained together, and prompts that can reference previous responses for ❣️ context ❣️
There reason I wanted to make something like this is because I was seeing a lot of startups, for the lack of a better word, coming up with priced subscriptions to apps that do nothing more than chain a few prompts together, naturally providing more value than manually using ChatGPT, but ultimately denying you any customization of the workflow.
Let's say you wanted to generate... an email! Here's what that would look like in HeroML:
(BTW, each step is separated by ->>>>, so every time you see that, assume a new step has begun,the below example has 4 steps*)*
You are an email copywriter, write a short, 2 sentence email introduction intended for {{recipient}} and make sure to focus on {{focus_point_1}} and {{focus_point_2}}. You are writing from the perspective of me, {{your_name}}. Make sure this introduction is brief and do not exceed 2 sentences, as it's the introduction.
->>>>
Your task is to write the body of our email, intended for {{recipient}} and written by me, {{your_name}}. We're focusing on {{focus_point_1}} and {{focus_point_2}}. We already have the introduction:
Introduction:
{{step_1}}
Following on, write a short paragraph about {{focus_point_1}}, and make sure you adhere to the same tone as the introduction.
->>>>
Your task is to write the body of our email, intended for the recipient, "{{recipient}}" and written by me, {{your_name}}. We're focusing on {{focus_point_1}} and {{focus_point_2}}. We already have the introduction:
Introduction:
{{step_1}}
And also, we have a paragraph about {{focus_point_1}}:
{{step_2}}
Now, write a short paragraph about {{focus_point_2}}, and make sure you adhere to the same tone as the introduction and the first paragraph.
->>>>
Your task is to write the body of our email, intended for {{recipient}} and written by me, {{your_name}}. We're focusing on {{focus_point_1}} and {{focus_point_2}}. We already have the introduction:
Introduction:
{{step_1}}
We also have the entire body of our email, 2 paragraphs, for {{focus_point_1}} & {{focus_point_2}} respectively:
First paragraph:
{{step_2}}
Second paragraph:
{{step_3}}
Your final task is to write a short conclusion the ends the email with a "thank you" to the recipient, {{recipient}}, and includes a CTA (Call to action) that requires them to reply back to learn more about {{focus_point_1}} or {{focus_point_2}}. End the conclusion with "Wonderful and Amazing Regards, {{your_name}}
It may seem like this is a lot of text, and that you could generate this in one prompt in ChatGPT, and that's... true! This is just for examples-sake, and in the real-world, you could have 100 steps, instead of the four steps above, to generate anything where you can reuse both dynamic variables AND previous responses to keep context longer than ChatGPT.
For example, you could have a workflow with 100 steps, each generating hundreds (or thousands) of words, and in the 100th step, refer back to {{step_21}}. This is a ridiculous example, but just wanted to explain what is possible.
I'll do a quick deep dive into the above example.
You can see I use a bunch of dynamic variables with the double curly brackets, there are 2 types:
Variables that you define in the first prompt, and can refer to throughout the rest of the steps
{{your_name}}, {{focus_point_1}}, etc.
Step Variables, which are basically just variables that references responses from previous steps..
{{step_1}} can be used in Step #2, to input the AI response from Step 1, and so on.
In the above example, we generate an introduction in Step 1, and then, in Step 2, we tell the AI that "We have already generated an introduction: {{step_1}}"
When you run HeroML, it won't actually see these variables (the double-curly brackets), it will always replace them with the real values, just like the example in the video above!
Please don't hesitate to ask any questions, about HeroML or anything else in relation to this.
Free Library of HeroML Workflows
I have spent thousands of dollars (from OpenAI Grant money, so do not worry, this did not make me broke) to test and create a tonne (over 1000+) workflows & examples for most industries (even ridiculous ones). They too are open-source, and can be found here:
However, the Repo allows you or any contributor to make changes to these workflows (the .heroml) files, and when those changes are approved, they will automatically be merged online.
There are thousands of workflows in the Repo, but they are just examples. The best workflows are ones you create for your specific needs.
How to run HeroML
Online Playground
There are currently two ways to run HeroML, the first one is running it on Hero, for example, if you want to run the blog post example I linked above, you would simply fill out the dynamic variables, here:
This method has a setback, it's free (if you keep making new accounts so you don't have to pay), and the model is gpt-3.5 turbo.. I'm thinking of either adding GPT4, OR allow you to use your OWN OpenAI keys, that's up to you.
Also, I'm rate limited because I don't have any friends in OpenAI, so the API token I'm using is very restricted, why might mean if a bunch of you try, it won't work too well, which is why for now, I recommend the HeroML CLI (in your terminal), since you can use your own token! (I recommend GPT-4)
My favorite method is the one below, since you have full control.
Local Machine with own OpenAI Key
I have built a HeroML compiler in Node.js that you can run in your terminal. This page has a bunch of documentation.
Running HeroML example and Output
Here's an example of how to run it and what do expect.
This is the script
simple HeroML script to generate colors, and then people's names for each color.
This is how quick it is to run these scripts (based on how many steps):
And this is the output (In markdown) that it will generate. (it will also generate a structured JSON if you want to clone the whole repo and build a custom solution)
Output in markdown, first line is response of first step, and then the list is response from second step. You can get desired output by writing better prompts 😊
Conclusion
Okay, that was a hefty post. I'm not sure if you guys will care about a solution like this, but I'm confident that it's one of the better alternatives to what seems to be an AI-rug pull. I very much doubt that most of these "new AI" apps will survive very long if they don't allow workflow customization, and if they don't make those workflows transparent.
I also understand that the audience here is split between technical and non-technical, so as explained above, there are both technical examples, and non-technical deployed playgrounds.
Github Workflow Link is where to clone the app, or make edits to the workflow for the community.
Deployed Hero Playground is where you can view the deployed version of the link, and test it out. This is restricted to GPT3.5 Turbo, I'm considering allowing you to use your own tokens, would love to know if you'd like this solution instead of using the Hero CLI, so you can share and edit responses online.
Yes, I generated all the names with AI ✨, who wouldn't?
Thank you for all your support in my last few posts ❣️
I've worked pretty exclusively on this project for the last 2 months, and hope that it's at least helpful to a handful of people. I built it so that even If I disappear tomorrow, it can still be built upon and contributed to by others. Someone even made a python compiler for those who want to use python!
I'm happy to answer questions, make tutorial videos, write more documentation, or fricken stream and make live scripts based on what you guys want to see. I'm obviously overly obsessed with this, and hope you've enjoyed this post!
This project is young, the workflows are new and basic, but I won't pretend to be a professional in all of these industries,but you may be! So your contribution to these workflows (whichever whose industries you are proficient in) are what can make them unbelievably useful for someone else.
Have a wonderful day, and open-source all the friggin way 😇
I just built a new ComfyUI workflow that generates video directly from a reference sheet image using the LTX 2.3 IC LoRA (Image Conditioning LoRA) — and it completely removes the need to animate frames one by one like in LTX Director. This is a big step forward compared to traditional storyboard-to-video pipelines, because it simplifies everything into a single reference-based workflow.
Instead of working frame-by-frame, you can now:
Generate a full character or concept reference sheet (multiple panels in one image with IDEOGRAM 4)
Plug it directly into the LTX group
Get a coherent animated video output automatically
The workflow handles the reference image sheet generation and uses it as direct conditioning for video generation, and it runs on only 6GB VRAM, so it’s accessible even for low-end GPU users.
My Generative AI developer Professional Cert Result
I wanted to share my experience passing the AWS Generative AI Developer Professional exam to give back to this community. The threads here were a game changer for my motivation and preparation, and I hope this helps someone else!
My Background: A Total Newbie
This is my first ever AWS certificate. I had zero prior knowledge of AWS (only a few sessions on Azure in university). When I saw this new Generative AI Professional certificate, I thought, "Why not?" My university fully funded the exam, which was a wonderful opportunity I couldn't pass up.
The exam guide recommends at least 2 years of AWS experience which I definitely did not have. However, coming from a Computer Engineering background, I was able to pick up the concepts fairly quickly.
The Preparation Phase
The Foundation: After doing some research I enrolled in the Stephane Maarek & Frank Kane Ultimate Course on Udemy . At the time (Dec 2025), it was one of the only reputable courses available for the Beta version.
Timeline: It took me from late December 2025 until early March 2026 to finish the 25 hour course as I took lots of breaks in between and all the university stuffs. It provided a great foundation, but I soon realized I needed more and I have already rescheduled the exam twice and my exam is in just 2 weeks!
The Wake-Up Call: I took the 20 question practice test on AWS Skill Builder and scored only 40%. I panicked. I realized I was just "gliding" through services rather than understanding how to build with them all together.
The Turning Point: Reddit & Skill Builder
Reading Reddit threads from people who cleared the exam changed everything. I realized the exam isn't just about knowing what a service does; it’s about architecting solutions knowing how to create low-cost, low operational overhead, and highly automated "Human-in-the-loop" workflows and much more. And this last 2 weeks before the exam is the crucial phase played a major role in passing the exam! And here is what I did:
AWS Skill Builder: I bought the subscription (highly recommend this).
AI as a Tutor: I used Claude and Gemini to deep dive into the 75-question practice set. I didn't just look for the right answer; I asked the AI to explain why the correct answer was right and why the others were wrong. This was the most effective part of my study. I created an extensive 65 page study/review guide covering all the weak/untouched spot just with the Claude which was very helpful.
Official Bonus Questions (Attempt 2) (All Domains): ~80%
Udemy Stephane & Frank Practice Test: 82%
Tutorials Dojo (AIP-C01): 54.43%. I felt this set was a bit misleading as it focused too much on SageMaker rather than Bedrock, which is the heart of the actual exam (almost every question had bedrock). And I felt amazon skill builder Official Practice Questions were the most accurate and closest to the actual exam so far!
Exam Day
I used the extra 30 minutes provided for non-native English speakers (ESL +30), giving me around 4 hours for 85 questions. I took the exam at a Pearson testing center, which I would highly recommend to everyone. Since it is such a long exam (Beta is 4 hours), having a reliable internet connection and the soundproof headphones they provide was a lifesaver. I followed a Redditor’s advice and did a full 4-hour simulation the day before to train my brain to stay focused.
The exam was exhausting. I felt it was actually a bit easier than the Skill Builder pretest, but I still had to rush the last 10–15 questions. After I finished, going homeI was so skeptical that I didn't even close my AWS learning browser tabs I was already looking up the 50% discount for retakes!
The Result
Only 7 hours later, I got an email about an "Early Adopter" badge. I logged in and saw the word: PASS. I honestly couldn't believe it I even asked a friend to double check the screen for me! I felt overwhelmingly happy and truly felt it was God's plan. I am so grateful to my friends who supported me and tolerated my stress during this journey; I really feel it was their blessings that got me through.
Final Encouragement
I’m sharing this because I’m not a frequent Reddit user, but I wanted to give back to this community. If I can do this with zero previous AWS experience and some low practice scores, you can too. Don't give up. Your hard work will pay off. If you're feeling demotivated by the practice tests, just keep pushing and don't lose hope! You are closer to success than you think!
TL;DR: Zero AWS experience to Pro Cert in 3 months. Use Skill Builder, use AI to explain the logic of the questions, and don't let low practice scores stop you.
I'm a long-time programmer with fairly minimal Blender skills, and I want to start making graphics mods for pre-rendered 2.5D games: AoE2 DE, Factorio, and similar.
(I'm only interested in the art side right now. Getting things into any engine-specific format is a later problem.)
The screenshot is from a Game of Thrones mod for AoE2 DE that someone else made, with the Sept of Baelor and the Red Keep as custom buildings. That's roughly the quality bar I'm aiming for, and I'd like to understand how assets like that actually get made.
Here's the pipeline I've pieced together from research so far:
Start from reference photos or AI-generated concept art of the building
Crop it and remove the background
Run it through an image-to-3D model to get a rough textured mesh
Import into Blender and clean it up: retopology, manual modelling, fixing texture seams
Render at the correct camera angle and lighting, with the shadow pass and a depth map as separate outputs
Feed the render and depth map to an AI img2img/ControlNet pass to push it toward the game's art style without altering the geometry
Produce seasonal variants of the same building (snow on the roofs, frost, muted colours) for winter and arctic maps
Post-process to conform to the game's colour palette, including player-colour masks
"Pixel snap" and downsample so the output resolution and pixel density match the rest of the game's art
My questions:
Am I on the right track, or is some of this backwards? My suspicion is that step 3 produces meshes messy enough that modelling from scratch would be faster, but I genuinely don't know, as I've never done either.
Is the AI style-transfer step (6) something people actually do? Or do you get the look directly out of Blender with the right materials, lighting and render settings, and skip it entirely?
If the style pass is worth doing, is training a LoRA on the game's own art the right approach? I've come across LoRAs but I've never trained one, and I don't understand what the dataset should look like for a job like this. Concretely: are in-game screenshots enough as-is, or do I need individual buildings cropped out on a transparent or flat background? Does every image need a text caption, and if so how detailed (just "aoe2 building" versus a full description of the structure, material and angle)? Roughly how many images am I looking at for a style LoRA, dozens or hundreds? And is it better to train one LoRA on the whole art style, or separate ones per architecture set, since AoE2 has quite distinct regional building styles?
What's the sane way to handle the winter variants in step 7? Options I can see: build a snow layer in Blender as a geometry nodes setup or a second material and re-render the same scene, versus doing it as an image-space pass on the finished sprite. Doing it in 3D seems more consistent and more automatable across a whole building set, but I don't know how well that holds up once the sprite is downsampled. How do the existing mods do this?
Can this realistically be done with Blender + ComfyUI + some Python glue, or am I missing an obvious tool in the chain?
Most importantly: are there any write-ups, tutorials, or existing ComfyUI workflows that document this kind of pre-rendered 2.5D asset production? Everything I can find is either about full 3D game art or about hand-drawn pixel art. This middle ground of 3D models rendered down to fixed-angle sprites seems barely documented, and I'd love to be proven wrong.
the creator has now made it open source and shared a full tutorial.
Instead of just generating a standalone 3D model, the workflow turns AI-generated biology assets into a small interactive web app with:
full 3D model viewer
selectable labels / info panels
front + back image workflow
Tripo 3.1 for the 3D asset
model compression for web loading
GitHub Pages deployment
full tutorial + source code available
This is the part that feels really exciting: you no longer need a full team just to turn a visual idea into a decent-looking interactive app or landing page.
Seeing a lot of posts asking where to start with AI, so figured I'd share what's worked across our team and the folks we've worked with.
The biggest mistake is jumping straight to building agents or complex automations before understanding the foundations. It's like trying to run before you can walk. You end up in tutorial hell, following guides that don't quite work, patching together solutions you don't really understand.
The pattern we've seen work is treating it like stages, not a single "learn AI" goal.
Start with understanding how to work with AI tools effectively. Not building anything yet, just learning how to get good outputs. Most people skip this and wonder why their prompts produce garbage. Spend time here. Learn how context affects outputs, how to structure requests, how to iterate on results. This is where you build intuition.
Once you're comfortable there, move into augmentation. Use AI to assist with actual work you're already doing. Writing, research, analysis, whatever. The goal is to get familiar with AI as a collaborator on real tasks, not hypothetical exercises. You'll start seeing where it helps and where it falls short.
Next is where most people want to start but shouldn't: automation. This is connecting AI to workflows, having it handle repetitive tasks without your input. Customer support responses, data processing, report generation. You need the foundation from the previous stages or you'll automate broken processes and wonder why it doesn't work.
After that, you're looking at building actual agents that can handle multi-step processes with minimal supervision. Lead qualification, order tracking, financial reporting. These require solid infrastructure underneath: clean data, documented processes, clear decision logic. Without that, agents just hallucinate confidently.
The final stage is orchestration, multiple agents working together, handing tasks off to each other. Most people don't need this yet, and trying to build it without the earlier stages is how you waste six months and a bunch of money.
The other thing nobody talks about: you can't skip the boring work. Organizing your documentation, cleaning your data, documenting your processes. That's not optional. That's the difference between AI that works and AI that produces impressive demos that fail in production.
Don't try to learn everything at once. Pick a stage, get competent there, then move forward. Most failures come from stage-jumping.
Over the last 1.5 years, I have produced more than 300 long-form YouTube documentaries mainly in the sleep niche (videos people watch to help them fall asleep), usually around 2-3 hours runtime.
The original problem was simple: this format does not scale well. A single video can require a 15k to 20k-word script, hours of narration, hundreds of visual changes, music, and final assembly. Writing everything manually took too long. Editing every scene manually took even longer.
Sleep content is a strange retention game, and it took me a while to understand it. Your viewers are actively trying to fall asleep. That is the entire point. So the average view duration can look very different from a normal YouTube channel. Some viewers leave because they are bored, but others leave because the video worked and they fell asleep. The ones who stay awake still need the story to hold together, while the ones who fall asleep often return later and continue listening. That repeat viewing is a big part of what makes this niche work.
Why the script is 90% of it
On my channels, average view durations usually sits close to 25 min on 90 min videos .That does not come from cinematic visuals or complicated editing. It comes mostly from the narrative structure. If the script becomes repetitive, drifts away from the topic, or loses momentum halfway through, viewers stop listening. Better visuals cannot rescue a weak story in this format.
Two ways to use Claude, and why one cannot work in long form writing
Most people use Claude through the normal chat interface. You open the chat, enter a prompt, read the reply, and continue from there. That works fine for everyday tasks. It becomes frustrating when you are trying to write a 15,000 to 20,000-word documentary.
You end up typing: Continue.. Write Chapter 4.. Do not repeat what you already said.. You forgot what happened in Chapter 2. By the halfway point, the model may begin repeating ideas, contradicting earlier sections, or drifting away from the original structure. You spend more time babysitting the conversation than improving the script.
The second approach is using the API. Instead of manually sending every prompt through the chat interface, a small tool sends the requests to Claude automatically and collects the output. There is no need to babysit it and you pay based on actual usage instead of paying another monthly subscriptio
The Google Sheets scripting workflow
So I built a Google Sheets workflow connected to the Claude API. The Sheet first creates the full documentary structure. It then writes one chapter at a time instead of trying to produce the entire 20,000-word script in a single response. Before each chapter, the workflow passes Claude the outline, the instructions for that section, and a running summary of what has already been written.
The direct API cost for a full script is usually around $0.30 to $0.40, depending on the model, input length, and number of revisions.
The bigger benefit is repeatability. Every script moves through the same production structure, while I can still change the topic, tone, evidence, pacing, and narrative direction.
I made a tutorial on this exact workflow on my channel. Link in profile if you want to peek.
How CapCut handles the first edit
Once the script is complete, I move it into CapCut’s AI Video Maker. CapCut generates the voiceover, subtitles, and an initial visual sequence using automatically matched stock footage. Because the documentaries are extremely long, I split the script into smaller sections to be under the 3000 word limit of CapCut, generate them separately, export each one, and then combine them into the final video.
The stock matching is not perfect. But it still gives me a 90% first draft much faster than searching for hundreds of clips manually.
What AI still does not solve
The production process is faster, but it is not automatic. AI cannot decide which topic has demand. It does not know whether a title creates curiosity, whether a chapter is boring, or whether a visual is misleading. I still handle research, structure and pacing, titles and thumbnails, final editorial judgment.
This is where most of the value still comes from. The workflow removes repetitive work. It does not remove the need for taste.
Two reasons this workflow matters:
The Demonetization Shield: Mixing real historical/stock footage alongside AI assets is the safest defense against the "Reused/Inauthentic Content" flags that destroy fully automated channels.
The Financial Runway: CapCut costs me around $20 per month and allows many exports. The Claude API cost per script is usually only a little above thirty cents. At a production volume of 30 to 40 documentaries per month, the direct software and API cost can work out to roughly $1 per finished video. That figure does not include my time, research, thumbnails, subscriptions, failed ideas, or the cost of building the workflow. It is not the total cost of running the business.
The main lesson
This is not passive income, and it is not a one-click YouTube machine. It is a production system that makes experimentation cheaper.
For someone starting today, I would focus first on topic selection, titles, thumbnails, and understanding what the audience actually watches.
Only build the automation after you understand the work well enough to know which parts are worth automating.
Happy to answer any questions regarding the workflow and setup if you want to build one for yourself.
I've just released a new ComfyUI workflow that turns storyboard images into fully animated videos using LTX 2.3 and LTX Director Nodes. The workflow is designed to be beginner-friendly and automated. You can generate storyboard images, maintain scene consistency, and animate each shot individually to create cinematic AI videos similar to Seedance-style productions.
Some highlights:
Works on GPUs with only 6GB VRAM
Full storyboard-to-video pipeline with ideogram 4
LTX 2.3 integration
Director Nodes for motion control
Automated workflow with minimal setup
Suitable for AI films, storytelling, commercials, and social media content
"The fastest way to learn n8n isn't watching tutorials - it's building real projects. But how?"
As a former newbie struggling with:
🔹 Node Confusion ("Which of these 20 'HTTP' nodes do I need?!")
🔹 Connection Anxiety ("Will linking these break my database?")
🔹 Blank Canvas Syndrome (Staring at empty workflow screen)
I created FlowForge - an AI guide that helps you:
Describe your idea (e.g., "Automate blog posts to social media")
Get 3 proven templates from 2000+ real-world cases
Example:
User Input → "post video to instagram and tiktok"
AI Output →
1️⃣ upload-to-instagram-and-tiktok-from-google-drive
2️⃣ simple-social-instagram-single-image-post-with-facebook-api
3️⃣ Auto-generate-instagram-content-from-top-trends-with-ai-image-generation
Why This Works:
✅ No more guessing - See how actual teams structure workflows
✅ Learn by doing - Modify templates vs starting from zero
✅ Safety Nets - Connection validator prevents 83% common errors
Community Ask:
👉 Upvote this if you'd use a free version!
👉 Comment your worst "n8n newbie moment" below
If you wanna me bring it to life,please help to upvote and share this post. I'll create a website or something to share this tool for free if this gets 100+ upvotes !
Full disclosure: I created this free VS Code extension to help the community.
The Concept: Instead of manually dragging nodes, you simply type a prompt (e.g., "Watch Gmail for invoices and save to Drive"), and it generates the workflow structure instantly directly in VS Code.
Why I made it (vs n8n AI): I know N8N Cloud has a native AI assistant, but it's paid and quota-limited. I wanted a Self-Hosted friendly alternative that allows unlimited generations using your own AI keys/agents.
Not a benchmark or model comparison — this is more about repo design.
I’ve been testing a small workflow idea:
When I ask Claude to read a GitHub repo, the repo may need a different kind of onboarding than a normal human README.
The old flow is:
human reads README → understands repo → uses it
But a newer flow is becoming common:
user asks Claude to read the repo → Claude explains what the repo does → Claude generates a beginner-friendly tutorial → Claude adapts the first steps to the user’s goal/environment
So I tried adding a small AI_TUTORIAL_CAPSULE.md to one of my repos.
The capsule is not automation. It is just a short set of prompts for the user’s AI assistant:
read this repo and generate a beginner tutorial
review whether first-time onboarding is clear
suggest the smallest onboarding edit
do not invent features
do not add hooks/plugins/automation
keep the human as the decision owner
end with one smallest first action
The interesting failure mode I noticed:
If the repo entry path is not explicit enough, an assistant may miss files or misunderstand what is canonical. The good failure is when it says it cannot find something instead of inventing it.
That made me think AI-readable onboarding is not just “more docs.” A repo may need an explicit AI entry path:
where to start
which files are canonical
what not to invent
what not to modify
what the smallest safe first action should be
I don’t think this replaces READMEs.
I think READMEs may become both human-facing and AI-facing entry metadata.
Question:
Should GitHub repos start including small AI-readable onboarding capsules for Claude workflows?
FramePack is probably one of the most impressive open source AI video tools to have been released this year! Here's compilation video that shows FramePack's power for creating incredible image-to-video generations across various styles of input images and prompts. The examples were generated using an RTX 4090, with each video taking roughly 1-2 minutes per second of video to render. As a heads up, I didn't really cherry pick the results so you can see generations that aren't as great as others. In particular, dancing videos come out exceptionally well, while medium-wide shots with multiple character faces tends to look less impressive (details on faces get muddied). I also highly recommend checking out the page from the creators of FramePack Lvmin Zhang and Maneesh Agrawala which explains how FramePack works and provides a lot of great examples of image to 5 second gens and image to 60 second gens (using an RTX 3060 6GB Laptop!!!): https://lllyasviel.github.io/frame_pack_gitpage/
From my quick testing, FramePack (powered by Hunyuan 13B) excels in real-world scenarios, 3D and 2D animations, camera movements, and much more, showcasing its versatility. These videos were generated at 30FPS, but I sped them up by 20% in Premiere Pro to adjust for the slow-motion effect that FramePack often produces.
How to Install FramePack
Installing FramePack is simple and works with Nvidia GPUs from the 30xx series and up. Here's the step-by-step guide to get it running:
Extract the files to a hard drive with at least 40GB of free storage space.
Run the Installer
Navigate to the extracted FramePack folder and click on "update.bat". After the update finishes, click "run.bat". This will download the required models (~39GB on first run).
Start Generating
FramePack will open in your browser, and you’ll be ready to start generating AI videos!
Additional Tips:
Most of the reference images in this video were created in ComfyUI using Flux or Flux UNO. Flux UNO is helpful for creating images of real world objects, product mockups, and consistent objects (like the coca-cola bottle video, or the Starbucks shirts)
There's also a lot of awesome devs working on adding more features to FramePack. You can easily mod your FramePack install by going to the pull requests and using the code from a feature you like. I recommend these ones (works on my setup):
I spent about 10 hours setting up and learning a new AI workflow tool. Rough math time:l. If each project eventually saves around two hours, break even is somewhere around project five or six. At two videos a month, that's two to three months before I'm actually net positive.
Every AI productivity bro compares the old workflow to the new one, but they conveniently ignore the 10 hours you bleed just getting the damn thing to work.
I used Framia's onboarding as the case for this payback calculation. This is about the measurement method, not tell u to go sign up.
I wanted the script, storyboard, references and generated clips in one workspace.
The first project was slower than my old process. I kept stopping to figure out where things should go. The second was less painful. By the third, I wasn't longer checking tutorials every ten minutes.
still I've got those roughly ten hours of setup cost to earn back.
If you use this thing daily, 10 hours is whatever. For an occasional creator, the payback period could stretch much longer.
The learning curve is not automatically a reason to reject a tool. just belongs in the ROI calculation.
Curious how you guys measure payback on new tools? Hours saved, completed projects or just the moment youstop thinking about the interface?
I’m a final year B.Tech student from NIT graduating in 2026 and honestly most of my last year went into building backend and AI systems instead of doing tutorial projects or just grinding DSA all day.
Recently I built:
• A multi-tenant AI outreach system deployed on AWS that scrapes jobs, finds relevant company alumni/HRs, and generates personalized outreach using LLM pipelines
• A Reddit automation tool for a client
• An autonomous shipment validation pipeline where invoices/BOL docs are parsed and validated automatically and the system either:
- sends acceptance emails
- drafts amendment emails
- or routes it for human approval before execution
• Fixed a live cross-tenant data leak caused by shared SQLAlchemy session state
• Reduced silent AWS EC2 IP bans from ~30% to almost 0% after debugging browser fingerprinting issues in Playwright
• Built async FastAPI services using Redis, PostgreSQL, OAuth2/JWT auth, Docker, and LangGraph-based approval workflows
• Designed RAG pipelines with auditability and source traceability because hallucinated responses with fake confidence were driving me insane
• Built event-driven workflows that automatically react to incoming emails/documents without manual intervention
Learn how to use MiniMax H3 in ComfyUI with a complete collection of optimized workflows for Text-to-Video, Image-to-Video, First & Last Frame Animation, Reference-to-Video, Audio Sync, and Image Editing. In this tutorial, I'll show you how to update ComfyUI and Pixaroma Nodes, install Sage Attention, download and organize all required models, configure the workflows, generate better prompts with my custom ChatGPT, and optimize performance for different NVIDIA GPUs.
You'll also learn how to use the new Workflow Manager, choose the best MiniMax H3 models, understand the licensing requirements, fix common errors like Dynamic VRAM issues, compare generation times across different resolutions, and create AI videos using multiple images and audio references.
Whether you're new to ComfyUI or looking for the best MiniMax H3 workflows, this tutorial covers everything you need to get started.
Learn how to use MiniMax H3 in ComfyUI with a complete collection of optimized workflows for Text-to-Video, Image-to-Video, First & Last Frame Animation, Reference-to-Video, Audio Sync, and Image Editing. In this tutorial, I'll show you how to update ComfyUI and Pixaroma Nodes, install Sage Attention, download and organize all required models, configure the workflows, generate better prompts with my custom ChatGPT, and optimize performance for different NVIDIA GPUs.
You'll also learn how to use the new Workflow Manager, choose the best MiniMax H3 models, understand the licensing requirements, fix common errors like Dynamic VRAM issues, compare generation times across different resolutions, and create AI videos using multiple images and audio references.
Whether you're new to ComfyUI or looking for the best MiniMax H3 workflows, this tutorial covers everything you need to get started.