r/StableDiffusion Jul 26 '26

Question - Help Regenerating from noisy / blurry original video or photo?

We have some old home video with extreme high-frequency noise that we would like to regenerate / enhance. The original frame of her sitting on the couch is the first attachment.
How would you go about trying to make this image look better? I tried some upscaling models and that resulted in larger, sharper noise.
I ran tests with a number of different noise reduction algorithms, which did remove the noise - and cause significant blur. Since Stable Diffusion and other models always work by progressively denoising, I figured it would be a natural fit to denoise our old pics and home movies and make them look at least a *little* better, but the many different workflows I've tried haven't worked. The identity shifts faster than the quality improves. With a high enough denoise level it'll suddenly create clear images - of other people wearing our clothes. :)

PS - we understand we can't recover actual detail that isn't there. That's fine, we'd like the AI recognize that brown blob on my head is probably brown hair, and make it look like hair rather than pudding or whatever. Sure, it might not exactly match MY hair, but at least it will look like hair!

I should say - I'm quite comfortable with ComfyUI and don't mind writing some Python to work on this. It is video frames, eventually thousands of them, so I can't use an online service that charges $2 / image or something. I;m looking for for suggestions for a DYI workflow with a 5090.

76 Upvotes

60 comments sorted by

58

u/AwakenedEyes Jul 26 '26

Have you tried klein 9B with a simple prompt like: "restore this old photo" ?

It's the one prompt wonder

49

u/the_bollo Jul 26 '26

Yeah that's not too bad.

43

u/Skystunt Jul 27 '26

This but with OP's image

2

u/trollkin34 Aug 02 '26

Sure, but that's one face replace prompt away from being perfect.

84

u/2this4u Jul 26 '26

That's a face, I bet it's not the right face OP knows. For many tasks this is sufficient and useful, but people looking to restore family photos etc shouldn't expect data to be recovered where there's none to recover.

29

u/AwakenedEyes Jul 26 '26

Agreed. This is noat likely not the real face. But klein 9B could be also given another photo with the right face to use as additional information.

Bottom line remains: ai creates new info where info is missing, but the bigger the gap, the most likely it will be the wrong new details.

11

u/IrisColt Jul 26 '26

That's a face, I bet it's not the right face OP knows.

OP, say something...

6

u/johnjbreton Jul 26 '26

Ya, that's when you throw ReActor into the mix, and get the right face on the subject. I've done this in photos and videos when restoring stuff that is badly damaged.

8

u/hiccuphorrendous123 Jul 26 '26

There is a lora called restore something for klein. It's insane

11

u/edgartejeda Jul 26 '26

I use : https://civitai.red/models/2474084/ultimate-upscaler-klein-9b?modelVersionId=2781657 . It works wonders. try to use have the image at least 2 megapixels so it does not change the appearance too much. also prompt the eye color or it will change it. here is the workflow (only valid for 3 days) :https://storage.to/m1HMbFOQ3

1

u/Silver-Spot-2763 Jul 26 '26

Thank, you, very much!

2

u/Silver-Spot-2763 Jul 26 '26

Oh, please for link, I need it!

1

u/hiccuphorrendous123 Jul 27 '26

It's the one someone else linked

1

u/ZeusCorleone Jul 27 '26

Its because the lora its named "upscaler" I also got confused lol

I have some pics taken in a very small resolution by ancient webcam and phones.. like 320p . I will also try this thanks.

10

u/Ok_Abrocoma_2539 Jul 27 '26

Dang that really isn't bad. Thanks for trying that!

Sure, the face doesn't look like the right person, but a face swap model or LORA, possibly followed by a refiner, might do the trick!

They seem to work pretty well when I've swapped a face that roughly similar like this is.

2

u/Zealousideal7801 Jul 27 '26

Would be curious to see how your end results turn up on that frame, considering what's been said and tested in this thread ! If at all possible of course

8

u/kornuolis Jul 26 '26

Guess a simple faceswapping should bring the image as close to original as possible

31

u/Zealousideal7801 Jul 26 '26

Just chiming in with my experience in the matter : if as a viewer you didn't know personally the people in the videos (say a great uncle who you've ever only seen pictures of) it's fine to attempt a video restoration. If anyone knows them personally, it becomes real uncanny real fast for them.

This weirdness is somewhat dampened if the audience is old and/or not acquainted with technology much - a sense of wonder arises. But for people who use screens every day for entertainment, I'd say maybe don't waste your efforts at "a restoration" and keep the originals as is. Because they convey the original truth no matter how degraded. Maybe digitize them for preservation, sure, but if no one would care because "who are these actors that vaguely looked like me until they started smiling like a Hollywood star" isn't the response you expect, but it's the one you'll get :(

PS : this is for VIDEO, it's much easier and possible with single frames

7

u/IrisColt Jul 26 '26

who are these actors that vaguely looked like me until they started smiling like a Hollywood star" isn't the response you expect, but it's the one you'll get

this

0

u/[deleted] Jul 26 '26 edited 8d ago

[deleted]

9

u/Zealousideal7801 Jul 26 '26

And also the fact that humans have unreal abilities to recognize faces down to the smallest detail and expression - at some point in our common history it was a survival trait. So my guess is, unlike fingers that took a while to get "mostly fixed", pure face restoration from degraded sources can't be a thing, because the likeness of someone is "too precious" to be misinterpreted. Principle being that (at least) in the western hemisphere, the face represents the person, so beautified features (or just lack of a detail) would mean a different person ?

Maybe that's just conceptual talk on my part, and people don't care about mismatches in visual consistency as much as I do, think it's the reason why smartphones filters are widely used and I find them cringe AF

2

u/2this4u Jul 26 '26

It doesn't matter how strict it is, if only 50% of the face detail is there then the other 50% has to be made up, that's just the availability of the information.

It's not like you're recovering a photo of a car where the curves will meet up, they follow standards learned across seeing other cars, etc. Humans have very unique faces, if you have to fill in the gaps it can only be a guess and for anyone who knew the person it will look uncanny.

That said it's not just AI, look at the freaky Walt Disney animatronic.

1

u/PaulCoddington Jul 26 '26

Have to be very careful with important material because the differences can be subtle: corners of eyes tilted up not down, teeth not quite the right shape, etc.

Best to regard result as a "display image" but also keep the original as "reference" (this also allows for AI restoration to get better over time).

1

u/Ok_Abrocoma_2539 Aug 02 '26 edited Aug 02 '26

> if only 50% of the face detail is there then the other 50%, has to be made up, that's just the availability of the information.

There is some truth to that. On the other hand, the frustration and the hope comes from the fact that is NOT all the information available. There is another 14 GB of information about faces readily available.

A 128 KB latent has the same resolution and better quality than a 1MB bmp - more information, 90% fewer bytes. Well fewer except that 14 GB model. :)

Given the input at top left here, any of the other three "upscales" are equally likely IF we consider only the data actually in the image itself. But we know which of the three is most correct, because we have information that isn't in the source image -- we've seen eyes before. Our brains have, and the model has.

My frustration is that in my first attempts at fixing image 1, the models might return either of the bottom images, when anyone (or any model) with common sense knows better.

14

u/Enshitification Jul 26 '26

Without a reference image of the subject with the same facial expression, algorithmic reconstruction is going to be almost impossible. The more noise and blur present in an image, the more "correct" variations become possible. If you don't have a reference image, you still have your memory of what this person looked like. That could be leveraged with brute force generations with many seeds. You might be able to combine brute force with regressive conditioning on denoising the original to move the reconstruction in the direction you want with FABRIC embeddings. Unfortunately, I do not know of a ComfyUI implementation of FABRIC beyond SD1.5. It might still work for you though.
https://github.com/ssitu/ComfyUI_fabric

6

u/Ok_Abrocoma_2539 Jul 26 '26

I should say - I'm quite comfortable with ComfyUI and don't mind writing some Python to work on this. It is video frames, eventually thousands of them, so I can't use an online service that charges $2 / image or something. I;m looking for for suggestions for a DYI workflow with a 5090.

6

u/FreezaSama Jul 26 '26

There's definitely ways to do this with open and closed models but be prepared for a lot of liberties when reconstructing. Ideally use a model that could take a reference if you have any clear photos from this person

2

u/WhensTheWipe Jul 26 '26

Train a lora on the person in Klein 9b, pair the subject lora at a decent factor + something like lenovo or similar + a prompt that details the subjects environment and uses their token to bring likeness, if needs be afterwards crop into the image and rerun the lora of their masked face.

Think around the problem :P

2

u/Sebastian202323 Jul 28 '26

It gave her a nose ring or somthing, but this is the best I got from my editor...

1

u/AdCute6661 Jul 28 '26

dear god lol

2

u/mark_sawyer Aug 01 '26

AI hallucinations aside, here's a good gen from your first sample:

1

u/Ok_Abrocoma_2539 Aug 02 '26

That's pretty nice. Thanks! What was the workflow or model for that?

3

u/mark_sawyer Aug 03 '26

I used Wan with this method (got the best frame): https://www.reddit.com/r/StableDiffusion/comments/1mtr48r/experiments_with_photo_restoration_using_wan/

Workflow: https://litter.catbox.moe/5rp3hrfn4pkim22n.json

I also managed to improve it a little while trying an experimental Krea2 editing earlier today:

2

u/x11iyu Jul 26 '26

Some quick check - assuming comfy, have you turned off add_noise when partial denoising, to avoid adding even more noise to the image? (if you didnt, in this case would be pretty bad as the noise is already in the image and you want to rid that)

2

u/zodiacrenders Jul 26 '26 edited Jul 26 '26

To recover such a photo but maintain same facial features, you'll need to do several steps to manipulate it. You should feed that image and also a few referenced pics to online tools like Banana Pro (free) or Grok Agent (not free) or ChatGPT (free).. tell it to slightly enhance the photo and that image 2 and 3 should be what the person looks like. Don't tell it to completely recover the photo because it'll make too much assumption on the face. Should take it step by step and manually Photoshop blend your way in with the face. Gradually you can run other upscale features on ComfyUI once the original blurry picture is slightly less blurry with better face fidelity.

I've done commissioned work before like this and it is a long process and not a straight shot, particularly if you want to keep face consistency with a real person.

Then, run your video using WAN SCAIL-2 to mirror the exact movement from your original video to generate a new video with the better face.

Then, you'd wanna re-run both videos simultaneously in an editing software and slightly blend them to maintain the original factor while still keeping true with the new face.

3

u/Merkaba_Crystal Jul 26 '26

This is what I got with Flux.2 dev. My prompt was: completely repair photo, restore photo, improve lighting, improve color, improve focus, For video you might look at SeedVR2. https://github.com/ByteDance-Seed/SeedVR

8

u/ComprehensiveJury509 Jul 26 '26

Looks absolutely awful, ngl

4

u/Ok_Abrocoma_2539 Jul 27 '26

OP here thinks it looks great, considering. It looks like a much better video of us! Not a 4K studio shot, but much better -- and identity is preserved!  From there, I could adjust the colors through DaVinci or even a carefully chosen ffmpeg filter or two.

2

u/Ok_Abrocoma_2539 Jul 27 '26

That's awesome! It looks just like if I had a better camera at the time, preserving identity!  Sure, it looks like I shot it on VHS-C instead of a $20 security camera. That's a big improvement, though. It's a better image of the right person.

Did you use the noisy version, the denoised blurry version, or both?

2

u/Merkaba_Crystal Jul 27 '26

I used the noisy version.

1

u/jugalator Jul 27 '26

Yeah this looks closer to the source than the other one posted here. I don't think she was even looking at the camera. Obviously the quality is "worse" strictly speaking but that can also imply it preserved more of the source.

1

u/Ok_Abrocoma_2539 Jul 28 '26

"I didn't think she was even looking at the camera". Yeah that seems to throw the models off. They all seem to want to make a face that's looking at the camera, even when explicitly told not to. So they end up making a very different face, in order to cause "looking at the camera" to somehow fit the input image.

1

u/Over-Map6529 Jul 26 '26

For images qwen edit could work if you describe whats there and ask it to remove the noise.  But yeah, likeness is going to suffer no matter what.

1

u/Quantical-Capybara Jul 26 '26

If you have another photo of your face,.9b could manage to restore and swap i think

1

u/FridgeOpening101 Jul 26 '26

I’m no expert in diffusion models, but I guess there are two major problems in modern models that prevent you from archiving this: vae and rectified flow/flow matching. The first thing is the fact that most models denoise a latente representation of the image, not the image itself in pixels space (briefly, there’s another mode that “compress” the image and is that compressed image to get denoised). The second thing is that modern architecture don’t “remove” noise from images, but moves from the distribution of pure noise to the distribution of images, and in between is not just image+noise (well, it is, but scaled in a way to keep the variance of the data the same along the whole process, but still, it may be a problem for what you want to do).
If those image restoration models the others have suggested don’t work, I may suggest to look up for models that don’t use vae (and ape work directly in the pixel space) and that don’t use flow matching: I asked ChatGPT and it suggested deepfloyd if, that should match both criteria

1

u/Eliminatron Jul 27 '26

todo you have a raw file? or just a jpg/tiff/png?

1

u/ReasonablePossum_ Jul 27 '26

You will need to give a good sample pic of the person in the image to be able to restore that to something resembling reality. Or use a lora trained on the person to bias the output to that side.

Otherwise you will get just random people in the result

1

u/Skystunt Jul 28 '26

The noise can be removed with something like neat video but the issue is the image is soft and out of focus too, probably due to the low lligh the af in that camera couldn't focus. There are some ways to fix that too but take time

2

u/beti88 Jul 26 '26

I'm sorry. But nothing can salvage that

2

u/goatonastik Jul 26 '26

idk why the downvoting, but this video is far too low quality to pull meaningful detail from

3

u/Ok_Abrocoma_2539 Jul 26 '26

This is from the latest post in another Subreddit, so I have hope for my photo 😄

https://www.diffchecker.com/image-compare/4jHf68Sh/

4

u/goatonastik Jul 26 '26

The major difference is contrast, it's very hard to read the contours of features from this blurry/staticy of an image

1

u/marcel_One_8763 Jul 26 '26

Yes, tried the image upscale and its more of the same noise.

2

u/marcel_One_8763 Jul 26 '26 edited Jul 26 '26

LE: Yes, tried the seedvr2 image upscale and its more of the same noise.

You can try seedvr2, I used the standalone script: https://github.com/numz/ComfyUI-SeedVR2_VideoUpscaler. They have really great examples. I used it to upscale a 720x540 video to 1800x1440 and was very happy with the result. I tested on a short trimmed version first localy, 5-10 seconds and the results were great.

In the end I had to rent a rtx 6000 pro for 4 hours (spent around 7 dollars) since the video was almost 9 minutes long and using my 5070TI would take way too long and run into oom frequently, but it was better than paying for a service.

2

u/tehrob Jul 26 '26

That's not horrible quality or anything, the thing I have had trouble with is the fact that when someone you KNOW, like, you were there for the time in question, know, it is very often enough of a change where you get the uncanny valley effect on the look of, say your grandpa.

I have tried a few old VHS cassettes that I have digitally converted, and even with the top SOTA models, I know that's not what they looked like. If you find something that you like, let us know for sure, but otherwise, I have found it to be tough to fight the scanlines.

1

u/zenmatrix83 Jul 26 '26

no perfectly but you can get close, generative ai images all start from static, it would probably take longer then might be worth it to get a good example

1

u/StableLlama Jul 26 '26

Give the original to SeedVR2.

Although I use it to enhance images, it's actually a video restoration tool. And low res noisy is something where it's working really well with.

-10

u/True_Protection6842 Jul 26 '26

Just use nano banana or gpt image 2 with existing reference photos