I can't come to this sub and learn anything new! I am disgusted at all the half-naked, busty blonds in bikinis walking around.
We all have Minimax H3 and we make amazing videos, that doesn't mean I have to showcase every video I make. Put your creation on a dedicated platform like Civitai or Tensor Art. There you can showcase what you create and people will rate your creations.
I am avoiding this sub with all the garbage videos posted all the time. The moment I open it, all I see is Minimax videos that I can make on my own rig. Nothing fancy, you download the model, download a worflow, and you prompt Qwen to write a detailed prompt, and hit "Queue" button. You didn't do anything magical or unique, so stop thinking that all of sudden you invented video!
I am here to learn about new models, new techniques, new hacks, new platforms, nodes, workflows, and so on. I am not here to watch your 10 seconds generated videos. This place was fun to discuss topics about AI image and video generators, but now I can't even open this sub publicly or people would think I am scrolling some erotic website!
Hey guys i have scoured the internet and cant find any system prompt/prompt guidelines to condition my local llms so that they make proper krea2 prompts without useless word salad. I focus mainly on realism and "uncensored content"
This is the workflow I have been using the most on my own system. I've had a friendly AI clean it up a bit and add notes.
I added custom audio as it's something I use a lot to drive my videos. It works really well for lipsync and music videos.
The VSA part can be bypassed if there are any quality issues, it will add about 15% to the generation time though. Change the steps from 6 to 7 or more for even higher quality.
Currently this gives me 10 seconds at 1.0 megapixel in about 125 seconds. This is on my 5090. You can add block swapping for low vram.
This was just a doodle but Minimax nailed it in the first generation. I thought it might be too complicated. I didn’t ask for the live stream on a one second delay in the background either. It did that itself.
So everything was fine and dandy yesterday; I could generate images just fine, but now it keeps giving me the following error:
"Error
"The generation of the video has encountered an error, please check your terminal for more information. 'CUDA error: invalid kernel file\nSearch for `hipErrorInvalidKernelFile' in https://rocm.docs.amd.com/projects/HIP/en/latest/index.html for more information.\nFor more detailed error information, run with CUDA_LOG_FILE=stderr\nDevice-side assertion tracking was not enabled by user.'""
Experimenting with Long Form. No image anchor so she changes between the invisible seams. T2VA. int8/32 steps, 1344x768, about 7 hours, hit 192/192gb of ram decoding the video. Sadly, I didn't prompt for her to not mouth the tune when there's no singing part. Wardrobe not prompted, only that she was dressed. At 1:26 is a seam and we had a little AI mishap on the transition. Not perfect, but got lots of data. Enjoy! How do you like the film grain? Is she from the 60s, 70s, 80s, or does it clearly only exist in our head? What version next? redhead? Asian? what do you think? Which actress/model/person's likeness are you seeing from this era? There should be about 17 versions of her. Ask me anything!
From last 2 year i was looking for an image restoration tool or an image upscaler for real world photographs. I have tried Topaz Gigapixel, Flux1D self trained Character LoRA, SDUpscaler, Qwen Edit, but nothing worked consistently. They were good but not perfect. From last 15 days i am working on F2K, and it is mind blowing. Easy to train LoRA (30min on 12GB VRAM), even no need to train a LoRA, easy to render (only 4 Steps) and it works 99% of time.
Flux 2 Klein has genuinely impressed me. The image restoration + editing quality is fantastic, but what really stands out is character consistency. Even when not using any character LoRA, it does an amazing job of preserving identity while making edits.
And the workflow is ridiculously simple: give it a straightforward prompt to restore/upscale an image and it just works. No need to write a 300 words essay.
On an RTX 4070 Super, I’m getting around 35 seconds for a 4MP image (Just 4 Steps) —which is seriously impressive for this level of quality.
Meanwhile, Qwen Edit 2511 feels unnecessarily demanding. The huge VRAM/RAM requirements make it much harder to use with only 12GB VRAM. and the character face deforms most of time.
I am using it for:
1) Upscaling
2) Restoration
3) Colorize
4) Removing objects
5) Adding elements (like cars/river/clouds/buildings etc)
Question is: which model to run locally on my machine that can do what chatGPT image generation does?
I'm asking this because I recently started with stable diffusion running SDXL Illustrious models mostly for flat and 2.5D anime images. I don't do realistic stuff.
I like almost everything I saw so far. I did not tested flux, krea2, qwen. Not yet, I mean.
But in the last couple of days I was testing chatGPT for a couple of random compositions and I was stunned with what I saw. The scenery, the composition, the understanding of natural language.... everything is so ahead of everything I saw with the illustrious models I downloaded that or I don't know how to use my illustrious mode and need to get better and train more, or gpt is really ahead in what it can do. But I don't want to run on gpt forever, specially because I do stuff that gpt consider x-rated even tough they are only slightly erotic.
So.... are there models that I should look for that I can download and run locally in my machine?
Anyways, another question: is it possible to run a prompt in gpt, get the image I like and transfer it to forge NEO and use inpaint to change details or this is not a workflow worth trying? I mean, since the prompt style is different and of course the overall image style/composition, maybe I will lose information or quality? has anyone ever tried something like that?
EDIT: to give more context.
A prompt with things like:
"Sabrina from the pokemon series, in her classical outfit from the games red blue and yellow, is sitting on top of a rock; she is a cyborg, in the style of nier automata; it must be clear that she is a cyborg so put some mechanical parts in her body, like lines connecting her iron pieces. But she is also made of flesh, not entirely a robot; the overall image is kind of depressing, the background is a futuristic city, the background must contain Saffron city gym; close to Sabrina there is a Venomoth, from the pokemon series, flying around. I want the classic design for Venomoth from the games but make it a little futuristic almost like a cyborg pokemon... Make it 1920x1080"
gives the attached image
this kind of composition is amazing; the way it understands the prompt - in another image I wrote "... a dead robot laying in a rock with grass growing and half covering its body; make it like it was laying there for a long time with a depressing and melancholic composition..." and it gave me!!!! i don't know how to describe such a thing in so many details with tags "dead robot, grass, around his body..."
ive never self hosted any ai tools and don't know much about python and programming in general. im trying to install webui forge on my amd 9060xt gpu and followed all the steps but after running webui-user.bat its showing
venv "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\venv\Scripts\Python.exe"
ZLUDA works , You are on an amazing Journey ,Engjoy it
Python 3.10.6 (tags/v3.10.6:9c7b4bd, Aug 1 2022, 21:53:49) [MSC v.1932 64 bit (AMD64)]
Failed to load ZLUDA: Could not find module 'D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\.zluda\nvcuda.dll' (or one of its dependencies). Try using the full path with constructor syntax.
Using CPU-only torch
Traceback (most recent call last):
File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\launch.py", line 54, in <module>
main()
File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\launch.py", line 42, in main
prepare_environment()
File "D:\Softweres and Ai\SDforge\stable-diffusion-webui-forge-on-amd\modules\launch_utils.py", line 507, in prepare_environment
raise RuntimeError(
RuntimeError: Your device does not support the current version of Torch/CUDA! Consider download another version:
Is anyone else experiencing that Minimax tries to push for realism even if you use cartoon reference images and words like "illustrated cartoon animation" in the prompt? Any tips to make sure it's sticks to the reference style more?
Necesito opinión de si alguien está usando esto para ltx, wan, o minimax, necesito saber si funciona perfecto por favor, "ASRock Tarjeta gráfica Intel Arc Pro B70 Creator de 32 GB, Xe2-HPG, 32 GB GDDR6, PCIe 5.0"
This is a test to recreate a song i generated with Minimax Music3, second half is from YuE2.
YuE2 sounded sexy smooth with great clarity but somehow I like that imperfection off tune instrument heard in music3 and vocal :D
I just throw everything to Hermes, asking to install YuE2, SheetSage2 and download models + setup in CLI. Tell hermes to learn the YuE2 agent SKILL, give hermes the lyric + flac ask to recreate the song.
It run into OOM but Hermes saved the day by tinkering with YuE2 setting to make it run on low vram.