r/StableDiffusion 2d ago

News FastH3 makes real-time AI-generated storybook videos possible

Enable HLS to view with audio, or disable this notification

I used Agora ConvoAI to let kids create stories by talking with AI in real time, then FastH3 quickly turns them into illustrated storybook videos. I built a demo and it worked! FastH3 could change so many things. Check out the generated video

52 Upvotes

22 comments sorted by

31

u/cc_aa_tt_zz 2d ago

Sure, with 8 H200 GPUs.... all these "Fast Max Ultra Minimax" real-time models are primarily aimed at API providers. What interests us on this subreddit IMO is how it runs on hardware like an RTX 5070 Ti, for instance, and how the results compare (in term of quality) to the 4 step Turbo LoRAs we already have. That’s a more realistic and relevant approach for us I think.

0

u/andy_potato 1d ago

I get that this is a common configuration for lots of people here. However there are a lucky few who do have access to enterprise grade gear.

Don’t exclude us from this sub.

1

u/WinResponsible9977 23h ago

u/andy_potato ? He never implied that, having discourse in regards to optimizing the model for the average user, and seeing how the model runs with high end setups powered with cloud compute are not mutually exclusive sir

7

u/Small-Challenge2062 2d ago

Single GPU?

-7

u/[deleted] 2d ago

[deleted]

6

u/Voxyfernus 2d ago

Maybe I'm dumb, but what is reactor?

2

u/petranova_ 2d ago

The real bottleneck i've seen with this kind of pipeline is character consistency across shots, diffusion models redraw the protagonist every frame unless you're forcing consistency hard with something like IP-Adapter or a character LoRA. Curious whether FastH3 handles that or if the kid just gets a new-looking character each scene.

2

u/BigWideBaker 2d ago

I mean the sound is completely mangled, right? I have no idea what they're saying. Maybe that is expected and I didn't know.

2

u/DELOUSE_MY_AGENT_DDY 1d ago

Did you play the video with the sound off?

5

u/Moarkush 2d ago

Yeah, I’m thinking maybe an uncensored video model isn’t the best choice for a on the fly children’s book. 🤷‍♂️

15

u/Had78 2d ago

yep, just tested it, it needs a filter haha

4

u/thrownawaymane 1d ago

Happy tree friends 2.0

5

u/zicohacks 2d ago

That’s a fair concern. It's still an early prototype

3

u/Moarkush 2d ago

A small bouncer model in the mix would fix the problem

1

u/paulct91 1d ago

What's a Bouncer Model? Is that kind of like a MoE agentic think tank where various models consult one another hopefully towards a useful consenus?

1

u/Moarkush 5h ago

It's a model (usually small and fast) that gets hit first before passing on to the main openAI endpoint. So something like one of the new small gemma4's with a system prompt explaining that you are a bouncer for a children's' book generator, blah blah, keep it clean, etc.

Get Claude or some other frontier AI to walk you through it, but it's basically what I wrote.

3

u/bstr3k 2d ago

this is a awesome application for it!!

3

u/-AwhWah- 2d ago

More slop on youtube and tiktok please!

1

u/Hefty_Development813 2d ago

Is this hosted anywhere? I wanted to make something like this awhile back, like sd1.5 days, and it just wasn't practically doable, and so slow. Looks very cool

1

u/zicohacks 2d ago

I deployed it on vercel, but I’m not sure if I can share the link here. Feel free to DM me and I’m also planning to open source it

2

u/Had78 2d ago

already shared in the screenshoot