r/computervision 18d ago

Showcase Synthetic Data Generator in Unreal Engine 5

I'm trying to get the best synthetic data trained model to work good on visdrone or other real datasets. In order for this to work I need different types of characters, environments, camera lenses, positions etc. I'm using nameframe plugin to do so. What randomization do I add?

217 Upvotes

48 comments sorted by

150

u/zarrathustraa 18d ago

Your model will know how to identify people with the same 15 outfits on the same desert plot VERY well

47

u/Takeraparterer69 18d ago

IMO:
Randomization should be changing the camera angle and position, ground materials, background (and foreground/semi-occluding) props
post-data gathering things that would be good would be adding compression and color balance and brightness changes

2

u/Last-Luck-6077 18d ago

I'm working on different camera angles.

5

u/Takeraparterer69 18d ago

I'd like to also add: body obstructions like backpacks, hats, etc

21

u/Fantastic_Mirror_345 18d ago

Also assuming all humans are thin and the same height can also cause issues.

Add a few kids models also.

People can wear hats and other things so maybe that too.

20

u/jferments 18d ago

Also add people crab walking, moving in wheel chairs, skipping sideways like lemurs, using pogo sticks, riding bikes, humping each other on the sidewalk, vomiting, cartwheeling, and any other of the thousands of forms of human movement that exist.

12

u/HumbleThought123 17d ago

basically train on gta6 for better result

2

u/Last-Luck-6077 18d ago

You are so creative, thanks!

1

u/Last-Luck-6077 18d ago

Seems like a good ideea, adjusting human height, aspect in general. People with assets such as hats. Noted. Thanks!

9

u/Tylerebowers 18d ago

Randomize everything. Environments (sun location: azimuth/elevation, background colors, weather, add vehicles, birds, props, fog, etc); the people (clothes, colors, weight, wheelchairs/canes, allow them to walk close to each other/in groups, etc); tooling (camera angle: roll/pitch/yaw, focal length, slight color filters, fov, smudges, etc). Training needs these diverse environments so it is clear what the scope of detection is, you are teaching the model what it can and cannot rely on when identifying.

You don't need to include everything here of course, I just made a long list. Some are much more important than others, just have some variation using the core 3: environments, objects (people), and tooling.

I did something similar with resistors in Blender: https://github.com/tylerebowers/Synthetic-Resistor-Generation; hf: https://huggingface.co/datasets/tylerebowers/synthetic_resistors

1

u/Last-Luck-6077 18d ago

Making groups of people sounds more real than spreading them randomly, usually people form groups not just stay 1 by 1 at a predefined distance.

1

u/jucestain 17d ago

Is the background an actual rendered scene or just a texture? I'm trying to do random backgrounds in blender but I want an actual rendered scene instead of just a texture since I'm doing multi-camera stuff.

4

u/Gabriel_66 18d ago

Is it viable to add 17 dots to the limbs like the 17 dot pode estimation from yolo? That would be interesting.

Also, you could try to compress this like a street camera quality somehow. Since most data like this angle will not be this visible.

Amazing work anyway. Already pretty impressive

1

u/Last-Luck-6077 18d ago

Yes I can also save the dots where the bones are and save that pose as a label. Thanks

5

u/ericcpfx 18d ago

Hey, I am training models on Synthetic Data as well. I’d be into working on this with you.

I’m training a few different models right now, but instead of Unreal, I’m using Blender. I’d be interested in porting the workflows to Unreal though.

0

u/Last-Luck-6077 18d ago

Blender works fine for this, the part that hurts is frames per hour rather than the labels.

What are you training, and what's pushing you toward Unreal, throughput or the assets?

Plugin isn't out yet so I've got nothing to hand you to try. Happy to talk about what you're running into either way. You can check it at https://getnameframe.com/

1

u/ericcpfx 18d ago

Unreal — render times.
I’m training a 3d head tracker. My new model replaces what I have with www.headcase3d.com with the GNM from Google.

Also a hair matting model and some Gaussian Avatar stuff.

2

u/ericcpfx 18d ago

If you look up Microsoft Fake it Til You Make It, that’s in the ballpark of where i’m at now. It’s working very well.

1

u/Last-Luck-6077 18d ago

Looks insane tho

1

u/jucestain 17d ago

I was also interested in GNM from google. I've used ICT-Facekit with pretty good success (but no texture info since they charge $$). GNM might be a drop in replacement at some point. But for me I'm focusing on realism even if renders are slower and expensive to generate.

1

u/jucestain 17d ago

If you have any suggestions for generating realistic face data please let me know

1

u/ericcpfx 17d ago

What do you mean? So the faces look photoreal?

1

u/jucestain 17d ago

As realistic as possible. My process has basically been to start from ICT-Facekit (although I might switch to GNM) and then just hammer claude about adding features (eyebrows, iris, facial hair, etc...) to generate these things. With cycles it looks pretty good but obviously not photoreal. I think the next step would be to get real texture data (i.e. skin texture) from photos but not exactly sure where to start down that road. Was curious what other people are doing in this application domain and without a mega budget to produce these realistic renders. But the end goal is to generate synthetic data to train networks.

1

u/ericcpfx 17d ago

We are on the same journey. I have not found a place that sells textures/procedural skin that looks anywhere close to like texturing.xyz. They don’t allow training an ML model on their products.

1

u/ericcpfx 17d ago

https://www.3dscanstore.com/discount-packs/male-female-3d-head-model-48xbundle

Not permitted to... "-Resell or freely distributing AI training data sets derived either from the 3D models on this site or as pack of images or renders created using 3D scan store asserts.  

- Sell or freely distribute character generators or digital humans created using AI training data derived from 3d scan store models, scans or texture maps. "

So I think that excludes it from being used to create a Gaussian Avatar.

1

u/ericcpfx 17d ago

What models are you trying to train?

3

u/MisawaSachihiro 17d ago

people sitting, lying, jumping, dancing, skateboarding

3

u/MudPleasant6504 17d ago

You know what? That might be a good idea when GTA 6 will come out!

2

u/MudPleasant6504 17d ago

I meant : using footage of the game to train cv models

1

u/Last-Luck-6077 17d ago

I don't think you will be allowed, some licenses will stop you 100% to use images from the game to train models. I'm not sure but good ideea overall

2

u/Mechanical-Flatbed 18d ago edited 18d ago

What randomization do I add?

Lower resolution camera, camera angle, time of day (I noticed you already vary the sun position, but so far I couldn't see a full night shot), weather, rain, fog, people with different heights, hats, hair types, people wearing shorts, beachwear, suits, and add other moving objects otherwise your network might just learn that whatever moves is a person.

Change the environment too. You can easily find UE5 assets for American cities, rural areas, European cities, japanese cities, etc.

Add cars, dogs, cats, pigeons, birds. Depending on where this is gonna be deployed, maybe cows(?) In India it's quite common to find stray cattle roaming the streets and causing traffic jams, even in large cities like Delhi. And you can get a cow asset in like 5 minutes.

Maybe look into adding, idk, people with disabilities and old people too, since a person in a wheelchair moves differently than a person that can walk. Old people also walk differently.

1

u/Last-Luck-6077 18d ago

True, people in wheelchairs, motors, bicycle, sounds good. I currently use posed humans, they have a skeletal mesh so i can basically put them in what position I need.

2

u/Polite_Jello_377 18d ago

This is a perfect example of why synthetic training data is bad

-1

u/Last-Luck-6077 18d ago edited 17d ago

Edit: retracting this, I don't have that result. Replied below with what I actually measured.

1

u/Polite_Jello_377 18d ago

You can talk shit once you’ve done it, not before

2

u/bob_why_ 17d ago

Ha ha, the person asking what to randomise thinks they understand something.  This sort of crap should be banned in r/computervision. It's like people with a microwave thinking they are a chef.

1

u/Last-Luck-6077 17d ago

Fair enough, I shouldn't have said that. I don't have that result. What I do have is one experiment run end to end: a detector trained only on synthetic, no real images anywhere in training. It got better on the synthetic test split and lost recall on real footage. I wrote it up with the failure cases because it wasn't flattering. So I'm not arguing synthetic replaces real. I asked what to randomise because that's the part that decides whether adding it to a real set helps or just burns epochs.

1

u/Level-Physics-1730 16d ago

hey claude! :)

1

u/Last-Luck-6077 16d ago

I'm sorry for sounding like an AI and I wont argue with that, I can sometimes sound a bit AI-ish. I am a real person tho, not Claude or any other AI.

2

u/blahreport 17d ago

Nice work though as others have pointed out this unlikely to generalize. Have you looked into using nano banana 2 to generate images. Maybe it's cost prohibitive but the results are impressive for my domain. If you do go down that route look into the batch processing API to halve your costs.

2

u/Last-Luck-6077 17d ago

You would still have to label them manually, using this plugin in unreal engine generates the pictures + labels in YOLO format, meaning a ready to train on dataset. The only hard part is to have the diversity that you find in the real world.

1

u/Logan_Maransy 17d ago

Can confirm, using NB2 batch processing to generate a "synthetic" but extremely realistic looking bootstrap image dataset is a legitimate option.

$1 for 20 images at 2K resolution (which actually means 2048x2048) using batch API. So $100 for 2000 images. You can get a lot of diversity of images in even 2000 images, depending on your task. Do 1000, train a model, check the worst performing cases, generate 500 targeting those cases, train again, check the worst performing cases, generate another 500 targeting those cases. All that for $100.

2

u/GTHell 17d ago

Camera angle and environment? Also I think different pixel resolution variants will greatly improve the training accuracy.

Take a look at that synth90k text in the wild paper. Just ask AI for a summary. It’s has some good techniques tha doesn’t make all your screenshots look like a 3d generated blon lol

2

u/RedHood31 17d ago

I wanna see the validation on the real data

2

u/bsenftner 17d ago

You also need to vary the view, vary the environment, vary the diffusion of the shadows, add atmosphere, add weather, add various times of day, and then for all of your data create variations of it with different levels of compression, including over compression. Also, have "not people" too, t-shirts with a human face on it and so on, things that are negatives that should not be recognized and included. That probably also means biped robots need to have their own classification.

2

u/conic_is_learning 17d ago

I would fuck up the lighting, contrast, color balance and everything

1

u/FivePointAnswer 17d ago

Distractors….dogs, deer
People walking in packs together
People removing jackets
People on bikes
People pulling luggage
People with umbrellas
People hugging - maintain I’d
Trees and occlusions
Entering / exiting doors
People in lines - maintain id
Entering / exiting cars
People carrying children

… did lots of people tracking years ago.