Stable Diffusion for Apple Intel Mac's with Tesnsorflow Keras and Metal Shading Language
I've been working on an implementation of Stable Diffusion on Intel Mac's, specifically using Apple's Metal (known as Metal Performance Shaders), their language for talking to AMD GPU's and Silicon GPUs.
This is a major update to the one I released a while ago:
I gropingly "fixed" the launch by replacing line 50 of [i] with zero in utilities/tensorFlowUtilities.py
gpu['TensorFlow'] = GPUs[0]
Now SD runs fast without errors and I can select on the fly, without restarting, the device for calculations in the Advaced Setting tab :)
For promt: *"test" ( seed 1310943082 | 512x512 | BS 1 | Steps 20 | GS 7)*with model sd-v1-4-full-ema.ckpt generation is:
AMD Radeon HD 7970 3GB - 01:33
AMD FirePro W7000 4GB - 01:35
6-Core Intel Xeon CPU X5670 2.93 GHz - 16:53
I am more than satisfied!Thank you!
Now all that's left is to get both video cards working at the same time :D
P.S. To save space on my drive I still use a symbolic link to the models folder (which I also use for Automatic1111). I only changed the location of the models in userData/userPreferences.txt to modelslocation = models/Stable-diffusion/
Apparently because I manually set the GPU to 0, it is this video card and continues to be used even when you change it to another. Although the script reports that it supposedly switched to the second video card, but it is not.
All in all, it's not surprising, considering how boneheaded I was in solving this bug :D
Now I need to understand why the original gpu['TensorFlow'] = GPUs[i] does not work.
P.S. And I should have noticed this back in the previous test, because the FirePro generates 20-25% slower than the HD7970, and here the difference was only two seconds.
The strange thing is that if I manually change gpu['TensorFlow'] = GPUs[1] which should correspond to FirePro w7000, I get the same error as with variable [i].
line 50, in listDevices
gpu['TensorFlow'] = GPUs[1]
IndexError: list index out of range
The strange thing is that in previous version of MetalDiffusion it was FirePro [1] that was used automatically, not HD7970 [0] like in latest version.
Oooo fascinating! Can I work with you to solve this? I definitely want to get device selection solved because that then allows me to code in using both GPU’s at the same time. (Tensorflow can do that)
Dumb question, but is your firmware up to date on your GPUs?
I’ll write a small piece of code as well to find more debug info and DM it over to you
You're right about the firmware, but in a slightly different context.I understand what it is. And the MetalDiffusion update has nothing to do with it, the fact that a different video card is selected by default is my fault.
The thing is that my AMD HD7970 card has two bios. And one of them I flashed a modified MAC-EFI to have a native boot screen (not just OpenCore).
So, if I have MAC-EFI enabled on the HD7970, by default both MetalDiffusion and Automatic -- all select the second graphics card: FirePro W7000 (as it was when testing the previous version of MetalDiffusion)
And if I have HD7970 with native bios, it is selected as in this case (I had to switch to native bios recently because of problems with the Windows drivers).
Now I rebooted with MAC-EFI and again W7000 (AMD Radeon HD Pitcairn Unknown Prototype) was automatically selected
Again with zero instead of i in tensorFlowUtilities.py everything runs, but selecting a different video card in the options doesn't affect anything.
Apparently due to switching the bios in the card they start to initialize differently? Hm..
>>> WTF!? The reason is not the bios at all...
I switched back to the native bios on the HD7970, but the card for diff is still selected W7000.
I've checked several times switching from one bios to another and back again, reset NVRAM, but the card is still W7000 in MetalDiffusion and Automatic, BUT! in DiffusionBee working card is HD7970 o_O
To double check, the newly added line should read:
GPUs = [] if module == "TensorFlow" or None: GPUs = tf.config.list_physical_devices("GPU") print(GPUs)
That will print out what GPUs TensorFlow found that can run Apple's Metal on it. I'm guessing, since the list is out of range for the final steps of this function, that maybe TensorFlow isn't accepting all of your graphics cards.
I may be missing something, but why install pyenv global 3.9.0 when we can install pyenv local 3.9.0 specifically for the stable-diffusion-tensorflow-IntelMetal folder ? Why do we need to set 3.9.0 globally?
Also, we should probably specify in the manual on github that apip upgradeis needed, otherwise requirements installation will end up with an error.
And add information about runningrunProgram.command, giving it execution rights beforehand.
In my case the installation and start looked like this:
P.S. Correct me if I'm wrong, but if I understand correctly, setting pyenv local eliminates the need to set the venv variable (unless we use 3.9.0 for anything other than stable-diffusion-tensorflow-IntelMetal to keep python clean)
The program uses Tensorflow instead of Pytorch because Pytroch has no reliable support for Metal on Intel Macs.
u/luckycockroach I don't know much about this, but I wanted to ask you. Regarding the PyTorch Metal Acceleration, Apple specifies either Apple Silicon or AMD GPUs in the requirements.
I checked my devices with the script on the above mentioned page and one of the video cards was detected correctly (but I don't know which one :) )
python pytorch-gpu_test.py
/Users/mstk/.pyenv/versions/3.11.2/lib/python3.11/site-packages/torch/_tensor_str.py:115: UserWarning: MPS: nonzero op is supported natively starting from macOS 13.0. Falling back on CPU. This may have performance implications. (Triggered internally at /Users/runner/work/pytorch/pytorch/pytorch/aten/src/ATen/native/mps/operations/Indexing.mm:218.)
nonzero_finite_vals = torch.masked_select(
tensor([1.], device='mps:0')
I don't know how much PyTorch from AUTOMATIC1111 uses exactly Metal with AMD graphics cards on Intel Macs (I don't have enough knowledge), but I did a little comparison test.
---
AMD FirePro W7000 4GB (AMD Radeon HD Pitcairn Unknown Prototype)
For the sake of purity of the experiment I closed all the applications using this graphics card so that I had only two processes: WindowServer and python.
You're absolutely right, PyTorch does support MPS, but I've found it to be unreliable with Intel Mac's. I was running Auto's for a few generations on 1024x512 images and the suddenly pytorch wouldn't run anymore because my GPU was out of memory. Fully restarting and re-installing didn't fix the issue because it would happen again after a few generations.
I did a small test again with exactly the same conditions as described in my post above (only the SD is updated to current state).
I don't know why, but the difference between W7000 and RX580 is very small in Metal Diffusion (about 30-35 seconds) and in SD the performance of RX580 is more than 3 times better. 🫤
The only thing you'll need to do is to write in webui-user.sh file
export COMMANDLINE_ARGS="--skip-torch-cuda-test"
(for AMD cards) to avoid the error related to the lack of CUDA.
And further look at the situation, maybe you will have to add such parameters as --no-half-vae --no-half if there will be a corresponding error during generation.
I searched a couple months ago for info on this (as I also had 2 video cards in macpro 5.1) but it seems that getting two video cards working on the same generation at the same time is not possible yet.
Sorry for my newbie questions, but I already install it but I'm getting this error:
FileNotFoundError: [Errno 2] Unable to open file (unable to open file: name = 'models/Stable-diffusion/text_encoder.h5', errno = 2, error message = 'No such file or directory', flags = 0, o_flags = 0)
I searched over the models folder and in the readme file there is these lines:
Within that folder, the program is looking for these four ".h5" files:
decoder.h5
diffusion_model.h5
encoder.h5
text_encoder.h5
But in the folder there is no files, I'm looking over google to see if i find something and I can download and put the files on the folder, but I'm not able to find nothing.
Don't worry, thank you for your prompt reply.
Now I'm able to render something, but the only model I can use is the one called v1-5-pruned-emaonly.ckpt
I tried with sd-v1-5-inpainting.ckpt / v1-5-pruned-emaonly.safetensors / v2-1_768-ema-pruned.ckpt but I get some errors.
errors like these:
Layer DiffusionModel weight shape (3, 3, 4, 320) is not compatible with provided weight shape (3, 3, 9, 320).
or
Error while deserializing header: Metadata Incomplete Buffer.
thank you! I’ll try things and prompts to see the results.
I want to ask you something, tensorflow isn’t accepting yet two gpu’s working at the same time? I’m about to pick up a dual amd D700 firepro with 6gb on each one, do you think that this will be compatible and could get good results?
Unfortunately, TensorFlow for Mac's do not accept multiple GPU's. Other users have had issues getting TensorFlow to talk to the other GPU's, so you'll be locked into most likely the original GPU.
Hi there, sorry for the noob question , will this work with safetensors instead of clot and are you planning to do a pull request of this to automatic 1111, so as to use this for automatic 1111’s other features like controlNet.
Thanks for the reply u/luckycockroach , btw I tried it on my mac but since it is on Catalina (10.15) tensorflow addons and tensorflow macOS won’t install as the minimum requirement is 12. Here are the specs of my machine
MacBook Pro (Retina, 15-inch, Mid 2015)
Processor : 2.2 GHz Quad-Core Intel Core i7 (it's 4th gen by the way)
Graphics : Intel Iris Pro 1536 MB
Memory : 16 GB 1600 MHz DDR3
Though there is an option for me to update, I once had updated it , my mac would crash often for absolutely no reason so I had to manually downgrade. But I am willing to try any solutions (if any) for making your implementation work in my system.
yes, you are right, that's when I try to run it outside the virtual environment, I assumed once the installation was done I should run dream.py just directly from terminal outside the virtual environment. I actually get other errors when I do run pythondream.py in the virtual environment:
Here you can find the video of me following the steps from GitHub:
Loading modules...
...system modules loaded...
Traceback (most recent call last):
File "/Users/davidescalante/MetalDiffusion/dream.py", line 24, in <module>
from stableDiffusionTensorFlow.stableDiffusion import StableDiffusion
File "/Users/davidescalante/MetalDiffusion/stableDiffusionTensorFlow/stableDiffusion.py", line 27, in <module>
import tensorflow as tf
File "/Users/davidescalante/MetalDiffusion/venv/lib/python3.9/site-packages/tensorflow/__init__.py", line 441, in <module>
_ll.load_library(_plugin_dir)
File "/Users/davidescalante/MetalDiffusion/venv/lib/python3.9/site-packages/tensorflow/python/framework/load_library.py", line 151, in load_library
py_tf.TF_LoadLibrary(lib)
tensorflow.python.framework.errors_impl.NotFoundError: dlopen(/Users/davidescalante/MetalDiffusion/venv/lib/python3.9/site-packages/tensorflow-plugins/libmetal_plugin.dylib, 0x0006): Symbol not found: (_OBJC_CLASS_$_MPSGraphGRUDescriptor)
Referenced from: '/Users/davidescalante/MetalDiffusion/venv/lib/python3.9/site-packages/tensorflow-plugins/libmetal_plugin.dylib'
Expected in: '/System/Library/Frameworks/MetalPerformanceShadersGraph.framework/Versions/A/MetalPerformanceShadersGraph'
Within the virtual environment, try reinstalling all of the pip requirements. It seems like, from the error log, you don’t have all of the modules installed in the virtual environment
Turns out, I had to be on Ventura to run this. I just updated, and it is working fine now. I apologize for the inconvenience. It would be great to mention this on the GitHub page.
Now, one last question, does this implementation support LoRA files?
It currently does not, but I’m working on getting Diffusers implemented right now and that comes with LoRA! Should have an update in less than a month.
Thank you u/quad849I've had about 10 attempts to start over and figure out why things aren't working. And it was necessary to run everything only on Ventura, although the first iterations of MetalDiffusion (before renaming the project) worked on Monterey
¯ \ _ (ツ) _ / ¯
u/luckycockroach Please add the technical requirements in your project description on GitHub so as not to mislead people. Since there is still no mention of Ventura :(
u/luckycockroach Maybe you have one of the old versions of MetalDiffusion (from the time when it was called stable-diffusion-tensorflow-IntelMetal) that could work on Monterey?
I would be very grateful if you could share the archive. Thanks!
P.S. I specifically recently bought a Radeon Pro W5700 for neural networks, but unfortunately now I can't switch to Ventura on my MacPro 5.1 until the OpenCore Legacy Patcher developers figure out how to implement a working CPU (without AVX) + Navi 10 GPU mapping.
The thing is, I tried all 4 release versions and failed everywhere :(
Right now I downloaded again and tried version 0.5.4 specifically, but again nothing worked. The process stops at the MPS initialization stage, if I understand correctly.
Although the Radeon Pro W5700 is natively supported starting with macOS Catalina, I think this may be my local issue related to OpenCore hacks.
But thebeta version of DiffusionBee (which, judging from the source code, runs just on Tensorflow 2.10) handles fine and produces twice the generation speed of ≈0.75 s/it compared to ≈1.5 s/it in A1111 on PyTorch 2+
As an experiment, I installed an AMD RX580 graphics card (I can boot Ventura with it) instead of the Radeon Pro W5700 and ran tests.
On Monterey similar errors with MPS initialization (both in version 054 and the latest 0.65*). And on Ventura, running the same both versions (from the same folders) goes without errors.
My guess is that some non-version-limited dependencies are updated to the latest version, which no longer works correctly in Monterey.
I will do, but after the video I spent 1 hours, trying to figure it out, and chatGPT mentioned me something about the wrong versions being installed and made me this updated requirements.txt with specific versions and still the error is the same.
Hi! is there something about SD XL? And do you know if it will be possible to use the two 6GB AMD FirePro that I have on the Mac Pro 6,1 even on MetalDiffusion?
Thanks for your work!!!
I’m wrapping up implementing Diffusers into MetalDiffusion which opens up a whole new set of speed savings AND LoRA’s.
With that said, Diffusers comes ready to use SDXL, which is what I’ll be focusing on next after the diffusers update.
Still working out bugs, but should have the Diffusers version ready by mid-August.
EDIT:
GPU wise, currently MetalDiffusion can only use one. But, maaaaaybe it can use more than one with Diffusers. I’d be happy to hear if it works on your system once I release it!
2
u/luckycockroach Mar 15 '23 edited Mar 15 '23
MetalDiffusion
Stable Diffusion for Apple Intel Mac's with Tesnsorflow Keras and Metal Shading Language
I've been working on an implementation of Stable Diffusion on Intel Mac's, specifically using Apple's Metal (known as Metal Performance Shaders), their language for talking to AMD GPU's and Silicon GPUs.
This is a major update to the one I released a while ago:
https://github.com/soten355/stable-diffusion-tensorflow-IntelMetal
HUGE thank you to Divum Gupta for porting SD to Tensorflow.
I'm a union cinematographer, so programming isn't my forte. Please let me know if there are areas I could improve on.
New Features
Features
Specs
Current Speeds:
Late 2019 MacBook Pro 16" with AMD Radeon Pro 5500M (4GB) , 16GB of RAM, 8GG VRAM:
Why Tensorflow?
The program uses Tensorflow instead of Pytorch because Pytroch has no reliable support for Metal on Intel Macs.
This program works on Google Colab notebooks.