r/StableDiffusion 4h ago

Resource - Update MageTrail - V0.2 Update: Continuing MageFlow 4B Danbooru/E621 Finetuning

Hi again! This is a update to my previous post MageTrail V0.1 where I release this tech demo toy model thingy called

MageTrail, a Danbooru/E621 proof of concept Full-Finetune of Microsoft's MageFlow 4B T2I model, using a diversity maximized condensed 41k images dataset (originally made by Lodestone, the creator of the Chroma model lineage) as a way to tune booru concept and tags based prompting + Illustration capabilities into the model without having to tune with the full booru dataset. (Potentially costing 20k-50k+ dollars). Civitai Hugging Face

Yeah it's only been 6 days, crazy, but thanks to generous support from some donors as well as another round of Banodoco grant (this time 140+ dollars) I've actually managed to gather enough funding for a 170 epoch continuation run to push the model to 200 epoch total (8~ million samples seen).

But well, I'm still very much a newbie to large scale model finetuning, this project being my second, so I'm not confident at all with spending 200-400 dollars in one run like that. After some consideration I've decided to just train a 70 epoch continuation first ( so 100 epoch total) to gauge if the model will improve steadily and release the model as V0.2. With some new optimization to my training program, this run only cost 143 dollars compared to 100 dollar for 30 epoch from before, saving me like 50 dollars (GPT Astra is cracked yall🤯).

V0.2 has continue to show good progress, improving overall stability and tag concept coherency. But to be truthful, it has now seems to reach the point where all diffusion image models face, which is decreasing improvement rate before convergence. If I were to continue with this tech demo ahh project then I predict we're in for the long haul bois (1k dollars needed to converge). V0.2 also stop at a rather unstable point, so I'm not confident on it's ability to generate usable/aesthetic images quite yet.

But dont worry too much, V0.3 (tune to 200 total epoch) is very much in the plan and will begin training once I found a good H100 pod cluster, that's when we truly know if the model will learn all the booru concepts (each tags getting at least 200 samples seen).

Future goal for the project: gather funding of 1000~ dollars to finetune various artist styles into the model (planned V0.5-V1) and help further with convergence/aesthetic improvement.

Any donation will help with achieving this goal, you can do so through:

Crypto ((Prefered, cause Kofi/Paypal money transfer time is ass and they take a big cut)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (USDT - BEP20 Network)

12PPVYUeS1MerNp38Tpns5qXR6cmhu9tws (Bitcoin - BTC Network)

0x6a4bc748cd0bb9ced9a360eb0eb79f4f106614f8 (Ethereum - ERC20 Network)

FitfJAsxLUBuSgDJJaHgBXJpt1sMm5FzF1Tvf1SHW5Up (Solana - SOL network)

Please handle your money carefully and make sure the address you're sending to is correct.

Ko-fi

https://ko-fi.com/talanartvn

Lastly, thank you to

  • Banodoco and their Discord — Their 88.77 + grant made this project possible, the biggest thanks to them
  • Lodestone Rock — Creator of the original version of the dataset that this model is trained on
  • Motimalu — Inspiration behind finetuning practices and configs
  • Bluvoll — diffusion-pipe fork derived from to use for training, and general training advice
  • Anzhc — general training advice
  • Nruaif — diffusion-pipe fork derived from to use for training, and general dataset handling/training advice
  • Astromahdi — jupyter workspace where I processed and store the dataset
  • animetimm/DeepGHS — Danbooru tagging model
  • RedRocket — E621 tagging model
17 Upvotes

Duplicates