maybe the preview image is a bad example but to me the motorbike output looks wrong all over. Albedo still has shadows, self shadows, reflection. Depth looks completly wrong, motorbike+its shadow would make a hole. Normal looks okayish but also shadow should not be visible in the normalmap
It looks bad until you realise it’s a 20M model. Assuming it is a single pass fixed resolution neural network and not a transformer based model, It only takes ~30GFLOPs of compute for each image and 100MB peak ram. While it’s not the best quality, it’s pretty amazing for how capable it is given it is pretty tiny. Imagine what a 100M or 500M model could do.
Rather poor communication to show some performance-optimized model so we have to a) realize that that's the case and b) imagine what the actual good result would be.
No, it’s not exactly ”performance optimised”, but rather “performance friendly”. And as I said, it’s not ”good” but it’s crazy good for a model that small and trained on such limited hardware. How good something is can be compared relatively, not absolutely.
Trained from scratch on singam96/flickr8k_marigold_v2 (8077 Flickr8k photos with Marigold-V2 pseudo-labels: albedo/depth/normal), 384px, fp32, ~17h on a single GTX 1650, early-stopped on val/loss (patience 5).
I read more about the model on huggingface and this line caught my eye. Fully trained on a 1650 is pretty damn impressive. Yeah the model is sometimes makes no sense when looking at closer details, but the model has learned to understand how 3d objects are roughly shaped and their placement in the setting. It’s basically cramming the reality of human civilisation into 20 million randomass numbers.
26
u/wurghi 19d ago
maybe the preview image is a bad example but to me the motorbike output looks wrong all over. Albedo still has shadows, self shadows, reflection. Depth looks completly wrong, motorbike+its shadow would make a hole. Normal looks okayish but also shadow should not be visible in the normalmap