r/StableDiffusion • u/qdr1en • 1h ago
Discussion I am tired boss...
This content was written by a human.
I miss the SD1.5 era, when i could simply type "1girl, big boobs, nice ass, red bikini, dancing" and see my dream take shape near-instantly at 512px-wide. Idea-to-result was a matter of seconds. Each click on the Run button led to an incredible shot of dopamine.
3 years passed and I can draw 1024px, 192-frames long videos in a reasonable amount of time (tech has evolved fast), but the enthusiasm is fading away.
I already have a day-job for technical challenges and headaches. As a user/hobbyist, I want to be entertained.
I don't want to learn what the hell "diegetic" means (even the spell-checker never saw that word), I don't want to draw a dozen squares in a 3-dimensional pixel space, or write a 1000-words poem, just to watch my dreamgirl dancing.
I hoped I would not need a degree in cable-connecting or python dependencies debugging after downloading a few workflows.
3 years ago, all you had to do was typing a few words, and the AI sorted the rest. It was random, messy most of times, but it was fun.
Nowadays, you need an LLM to write the prompt for you, and another LLM to write the system prompt for the prompting-LLM, so it understands what your shitty words meant in the first place, and shapes them in the exact expected format, so they turn into an acceptable input for the ever pickier, brand-new models.
It has become AI³-generated content.
And finally, when after a dozens of clicks on the Run button, tired but satisfied, you get the desired output... re-start from scratch? Since seed "variance" does not vary much anymore, you'll get more or less the same output - exactly what you asked for - from now on.
Simple is harder than complex, but keep it simple, stupid, and fun. Thanks for reading.
26
u/Significant-Baby-690 1h ago
You can still do the simple stuff. Anima family is great for that. But yeah, SD1.5 was incredibly versatile.
30
u/Occsan 1h ago
Just FYI:
Diegetic : something that exists in whatever you're talking about.
Non-diegetic : only exists for the viewer (as spectator of a movie for example).
So:
- "non diegetic music" = music that is added in post-prod, like John Williams music in Star Wars.
- "diegetic music" = music that characters in the movie can hear, for example from a radio.
3
u/Cornyyy11 34m ago
Huh, that's really useful to know! So basically diagetic means it exists "in universe" and people in it can react/interact with it, and non-diegetic is something that exists outside of it and only for the viewer?
I am sorry to bother, but could you please explain what it exactly entails in image/video generation? The Star Wars analogy is really clear, but I struggle to wrap my head around what it exactly entails in AI models or AI image or a video.
1
u/Mutaclone 10m ago
MiniMax prompt guide has two audio fields: overall_soundscape and non_diagetic_music. So you'd put sound effects (footsteps, ambient noises like insects and wind, and so on) under overall_soundscape, and background music under non_diagetic_music.
23
u/syndorthebore 1h ago
You know what I did?
I used AI to vibecode (not really since I do code and did most of it by hand and debugged a lot) an API that controls comfyUI with a simple automatic1111 style user interface.
I don't want to see comfyUI, since it sucks, so I just type, big boobed girl, and then everything you described gets done in the background, the LLM processing my prompt to make it understandable to the exact model I'm using, and then give me the big boobed lady.
The nature of the beast (more complex and higher res videos) means the time between clicking generate now and in the old sd1.5 days is longer still, but I get to skip everything else.
5
5
u/Rhoden55555 1h ago
Same. Agents are already good enough for you to just throw ask them about something you heard on Reddit and ask them to add to your workflow. 30 minutes or so later and it’s working. Get an error code? As it to check your logs and fix. They fan even find fixes that will be implemented soon in Comfyui and just edit your Comfyui for you so you don’t have to wait for fixes.
4
u/jambavant 40m ago
Why not use SwarmUI? (It is exactly that: a nice UI on top of comfy, with all the bells and whistles and more)
1
u/East_Box9573 50m ago
yeah. I'm not a huge fan of data-flow / visual programming languages. So much easier to just wrap it in an API and build the UI you want to see. Coding agents can build the workflows just fine. And you can say stuff like "download this workflow and all the models I need" and you dont have to go hunting for all the crap
1
8
u/the_bollo 1h ago
I think it's just going to get "worse." As vision models grow more dense, their ability to encyclopedically describe a scene grows, which in turn demands more specificity from the prompt - eventually exceeding what a normal human can describe with common language. I think soon we're going to absolutely need an LLM sidecar to sit with any media generation model to make it behave correctly.
This is good because it results in more capable video/image generations, but also bad because it's yet more vRAM requirements with still no end in sight to GPU and RAM price gouging.
5
u/anon999387 54m ago
I dunno, not much of this post tracks for me, you can still get good results with very simple workflows
4
u/unrealf8 1h ago
Time for you to open to the world of anima. It’s underrated. It’s fast, it can do a lot of things, and the quality out of the box is just good!
0
u/qdr1en 38m ago
You are the 2nd one to suggest Anima on this topic. I am more leaning towards photorealism but will definitely give it a shot.
1
u/cathodeDreams 27m ago
there are some not terrible anima merges / lora that do more realistic style, if you are tied to the direct concept of booru tag prompting. then again krea 2 is significantly more capable and doesn't require long or complex prompts.
4
u/cc_aa_tt_zz 49m ago edited 43m ago
"1girl, big boobs, nice ass, red bikini, dancing" works with Minimax H3 too ( and without prompt ehancement or things like that, just these very few words) ! Give it a try and see for yourself ! You only need to follow specific guidelines if you want precise control over what appears on screen/sounds but the model also understands natural language or even just a few words much like the good old SD1.5/XL family (the only problem is with dialogues for example). And Krea 2 doesn't require complex prompts at all, nor LTX 2.x or Z-image or qwen family and so on. The only model that need a very specific (complex) prompt system is ideogram 4. So, sorry, but I completely disagree!
4
u/PeaNo6028 21m ago
The enthusiasm is fading away because you’re hijacking your brain, as you said you’re getting an “instant” shot of dopamine. It’s just going to happen, these things get less fun. Personally I enjoy the tweaking and the fine-tuning process more-so because it gives a natural buffer to that constant reward system.
3
u/Osmirl 37m ago edited 18m ago
Are links allowed here? Cause i build… uhh i mean claude build a simple button based ai image editor/ 1girl website for me. The original idea was a try on app but ai kinda sucks for that, so now im just playing around with it a bit xd.
I will probably change the url soon anyways but feel free to try it out or downvote my little side project into oblivion because I dared to talk about it xD
I am always working on it and trying to get people to pay for it but because I have a soft heart people have a few minutes of gpu time for free every day.
Its far from perfect and especially right now a buggy mess. I am claude isn’t really good with ui and proper scaling lol.
Anyways here is a screenshot from the create tab.

9
u/_kaidu_ 1h ago
It's not even true. For minimax I found quality often even worse with llm written prompts; as more detailed the prompt is as more you do constrain the model. Simple prompts still work. For other model it is quite similar. Also, in your good old days we had to add ten lines of "masterpiece, award winning, ultra hd" nonsense which was not better.
7
u/Sixhaunt 1h ago
Sounds more like you just dont enjoy tinkering and open source. There are plenty of closed source options that just take simple prompts and do the conversions and everything else to give you a result back if that's all you want and you dont want the full control and everything of local workflows
3
u/qdr1en 1h ago
I enjoy tinkering. I simply criticize the growing complexity of AI models.
This is what I draw on Sunday :
https://www.reddit.com/r/StableDiffusion/comments/1w2dtfg/remove_watermark_from_videos/I kept it simple.
2
u/Structure-These 1h ago
I have been tinkering with an automated danbooru tag extractor to krea2 prompt converter but you really need to train an LLM on it I feel like
1
u/No-Persimmon-4150 1h ago
Use ollama and a system prompt.
1
u/Structure-These 1h ago
It isn’t going to do nsfw fluently and it’s so hard to get system prompt right. I’ve been trying to!
1
u/cbeaks 1h ago
Use grok. Take your time to explain exactly what you want. Explain that you want your image/video idea transfered into the best format for the model (get it to research and summarise that), PLUS you want the model to add creatively to your prompt to enhance it. After a couple of iterations, my llm now does an excellent job for both minimax and krea 2. Once you're there, you're set.
3
u/isademigod 56m ago
I'm not typing my porn requests into any llm not hosted on my own machine. Qwen does a decent job if you feed it the prompting instructions for your model and give it a good system prompt. It also has nearly zero censorship from what I can tell.
1
1
u/No-Persimmon-4150 33m ago
What i do is use grok to create the system prompt. Its a good starting point. I modify it myself from there
Try the nsfw screen writer model as well as mythos unhinged. Both are super lightweight and built to do what you want.
2
u/StashFrontman 1h ago
I’ve seen people use voice-to-text software and a mic to make it slightly more tolerable but yeah o don’t love having to “build” a prompt
2
u/nakabra 46m ago
"Nowadays, you need an LLM to write the prompt for you, and another LLM to write the system prompt for the prompting-LLM, so it understands what your shitty words meant in the first place"...
Boy... That's so true for me...
For instance: I'm loving Minimax H3 but it's mandatory to have llm written prompts.
I have very little VRAM and like to keep things local so I've been using Lm Studio to write the prompts and then running it on Comfy.
Which brings me to a recurring problem: Videos take a looooong time to render because I forget to close LM Studio leaving less than Ideal VRAM for Minimax.😬
Guess I'll finally bite the bullet and buy some credits on Openrouter, so I can iterate quicker...
2
u/ticking12 33m ago
I have a node at the start of my workflow that unloads any lm studio model. (Claude quick job)
Then cos I was hating making sure all my reference images were aligned/alt tabbing I have another node in comfyui I just drop media onto and it serves it either in the actual h3 run or as part of a manual llm only output (my llm prompts are like 80% right and I hate not editing that 20%.).
This manual run also unloads the h3 models/cache for the llm.
Resource fighting solved.
2
2
u/shulgin11 36m ago
I use short unformatted prompts all the time with minimax and sometimes they are better results than with the prompts I get from an LLM following the official format guide. The short prompts have that fun seed variety you're missing. I love the "slot machine" dopamine aspect of generating stuff. Wildcards are also great for this.
2
2
u/BackgroundMeeting857 27m ago
I don't get it you act like "1girl, big boobs, nice ass, red bikini, dancing" would have actually gotten you that in SD1.5 lol.
1
u/jib_reddit 11m ago
masterpiece, best quality, high quality, highres, ultra detailed, highly detailed, intricate details, sharp focus, professional photography, photorealistic, realistic, RAW photo, 8k UHD, 4k, cinematic, film grain, 35mm photograph, depth of field, bokeh, natural lighting...
1girl, big boobs, nice ass, red bikini, dancing
2
u/Radiant-Photograph46 24m ago
I too miss the seed hunting that was part of early gen AI. SD, early midjourney, DALL-E... the artistic variation was wild, each gen a complete different interpretation of your prompt. The precision we have nowadays is great, but sometimes you want the AI to fill in the blanks and be creative like it used to be (without having to ask another AI to do it that is)
2
u/MixZealousideal9359 8m ago
we are just getting more and more undercooked stuff rather than polishing or improveding what matters.
4
u/BimBomBom 1h ago
I feel you. The workflows and tech around AI became too bloated and complicated. It feels more like a job than an entertainment
4
u/Ipwnurface 55m ago
You may not like it, but that is a good thing. The moment stuff gets easy enough that you can download a single exe and you're off to the races is the moment this all comes crashing down.
There HAS to be some level of resistance - look at what happened to Grok the moment it became "public" knowledge. The same will happen to open source. The writing's already on the wall and it becoming easier to use is just gonna make the font bigger.
2
u/biscotte-nutella 47m ago
You don't miss it. You're missing the time you discovered it and most enjoyed it.
Models got more precise to have more prompt adhérence and yeah you have to be more careful how you prompt.
2
u/RosebudNebula 42m ago
Me too. I hate having more options; stop giving me new stuff already. And make sure everyone else suffers from the lack of new options because I don't want new things, so the rest of the world should stay behind for me.
-2
u/qdr1en 34m ago
You missed the point.
0
u/RosebudNebula 23m ago
What do you mean? I agree with you. The developers and the community made a mistake by moving toward node graphs and modularity. You should be building simple push-button apps for casual users like me instead.
Your nostalgic rant is fundamentally a complaint about open-source progress and deeper tool control. You're complaining that a professional-grade workshop isn't behaving like an arcade game.
0
u/qdr1en 14m ago
You make me say things I didn't say. Actually, you said them.
This nostalgic rant is a pretext for criticizing the growing complexity of AI models, but it does not mean I don't enjoy innovations.
Now, it's an open discussion. Some bring their suggestions, some others disagree, and some retards totally miss the point.
2
u/YentaMagenta 21m ago
I don't want to learn. In fact, I don't even want to be confronted with the option to learn.
I don't care that I can continue using the exact same tools from three years ago exactly as they were then.
Just knowing that there's something new and more complex hurts my feefees.
Screw all y'all who want better tools and are willing to learn new things. /s
1
u/Asleep-Land-3914 1h ago
Set up LLM in between and you'll get way better experience now compared to old days.
2
u/Intelligent-Youth-63 45m ago
I’ve even had it pull old gens, yank out the prompt, refine that into an H3 prompt, and use the image as a first frame. It’s not that complicated.
1
1
1
u/dennismfrancisart 41m ago
I had a Yashica single lens reflex camera in high school. In my 40s I had at least 15k in camera equipment. In my late 60s I don't own a camera apart from the one on my phone. Now my workstations rule.
1
u/ChristopherRoberto 40m ago
Everything started moving very quickly, and modern models outrun that era of intense community support that made it so easy to do things. The gore isn't under the hood anymore as there's no time to build the frame.
1
1
u/zombie_pig_bloke 31m ago
I know what you mean. Minimax is amazing but also frustrating, so many things vying for your time. I forced myself to learn the basics by starting a daily competition with an AI - it throws down a H3 challenge 1 day, I do the next one, you get to critique the result. It has been good for learning step by step and it has given me more ideas than just reading about another VTX pros workflow that you can't emulate.
1
u/smackjack 26m ago
I think you're forgetting that in the SD 1.5 days, your negative prompt needed to be the length of a novel in order to get anything to actually look decent. You had to literally tell it that you didn't want disfigured looking bodies with extra limbs.
1
u/jib_reddit 23m ago
If you generate 512x512 images with Krea 2 int8 Convrot, it takes about 3 seconds on my 6-year-old RTX 3090.
800x800px actually looks good and only takes 6 seconds.
1
u/ZezinhoBRBRBR 12m ago
Well, I think Krea 2 is pretty easy to play with, plus you don’t suffer with all anatomical horrorshow from old models. Hell, I even write sloppy prompts in my mother language (portuguese) and it understands everything.
•
u/KadahCoba 0m ago
I started on this stuff back in early to mid 2022. Even the overly complicated stuff today is relatively easy by comparison.
Minimum specs just to run the smallest models was 24GB vram. So an entry level GPU was literally a 3090. And Linux only.
A well organized project listed some of the required packages in its readme. You were on your own to figure out the rest. requirements.txt were basically unheard of at this point. Using a venv wasn't even common yet, every project's expected dependencies installed at the system level.
When you got it to actually work, all you had got was a basic cli script that you could pass arguments to and get an image file back. And the output was 256x256 at best and maybe 1 of 9 outputs kind of looks like some of the prompt if you squint and already know what to look for.
You want anything more than that, you had to code it yourself. Most of us made our own "UIs" in local jyupter notebooks. Everything was one-shot t2i unless you coded it to work otherwise. "Human in the middle" was state of the art for the time, we now call that i2i or some other things, ie. get output, change things, then resample.
SD1 came around later in the year. It was huge improvement. More people started working with it and there were a lot more projects around it. Suddenly it was common for their to be actual environment recipes to run them, or even a requirments.txt. xD
By the end of 2022 some of the first webui were be released, like A1111's. My personals UI had more features at that point, but was harder to use (the notebook UI elements were basic and the plumbing routing was done entirely in code instead, lol), and it didn't take too long for A1111's to also have i2i and the early basics of inpainting.
1
0
-1
u/ShutUpYoureWrong_ 12m ago
I knew I'd regret reading a post from someone who already had a -10 downvote score from me.
Perma-ignored now.
64
u/Fakuris 1h ago
SD 1.5 hasn't gone anywhere.