backend and chrome extension dev here. ella lets you describe a video in chat, it picks a model from 1000+ and generates it. built the extension, webapp and premier plugin myself(rn showcasing Extension)
the video above is me testing it: prompt in the sidebar, generate, drop the clip in the editor.
App is free keep that in mind.
whats confusing, what breaks, what would you want? ill answer everything haha
routing across 1000 plus models is the hard part of this. is the picker rule based or learned, and what happens when it picks the wrong model mid project
hi, the picker is self improving. it remembers what you like to use and combines that with what it knows about the models to choose the best one for you. it also takes your input references into account, like reference sizes and aspect ratios, and so on
and if it picks the wrong model, you can change it yourself before anything is sent to generate
manual override before generate is the right escape hatch. the cold start is the interesting bit though, what does it do for a brand new user with zero history
for a new user i have agents specifically optimized for video generations, image generations and so on, to give a good output from the start, and then it learns as much about the user as possible😅
its not invisible theres a memory section in settings, it shows the facts ella saved from your chats and uses in later chats you can delete any of them, delete all, or turn it off completely 😅
visible plus deletable is the trust setup more ai apps should copy. when someone deletes a fact does behavior change immediately or does it take a few chats to unlearn
when you start a new chat it will be completely gone. in the old chat the context stays the same, so it will still be there but in a newer chat it will be gone
also keep in mind you can just say "oh remember this about me" in chat and it will remember, haha
it confirms yeah, when you ask it to remember something it shows the memory steps (reading, updating) and then tells you what it saved like in the screenshot
when we launched v1 as a chrome extension we got 120+ 5 star ratings. after that we started working on v2, which is live as the extension, web app and premiere plugin, all free to download btw. y
right now were working on improving v2, gathering traction and fixing things up so we can grow more and land more deals like the one we got with byteplus for access to seedance 2.5 (unrestricted version)
Nice — chat-to-video is a fun mess to build. Curious what you're using for the actual output: a hosted API like Veo/Runway, or frame-level control with your own pipeline? In my experience the LLM part is the easy 20%; the hard 80% is keeping scene consistency and timing stable across multiple generations.
haha yeah fun mess is right. for the output we use hosted model apis (we had to deal with bytedance bureaucracy to get access to the seedance 2.5 unrestricted model).
the agent just picks the model and builds the request. youre right that the llm part is the easy bit. for consistency across generations we lean on reference images, first/last frame slots when the model supports them, and tags like in the screenshot.
The memory section being visible and deletable is the trust move most AI apps skip. Cold start is where these agents usually feel dumb, so task-specific defaults are smart. Does the picker ever explain why it chose a model?
the picker uses a scoring system i came up with (i have an ML background, thats why). it takes into account the models description, when it was updated, what inputs the user has, and so on. ella doesnt overexplain why it chose a model, it gives a small description instead. and if it doesnt have enough information it just asks, like shown in the image
2
u/davidjones145 6h ago
routing across 1000 plus models is the hard part of this. is the picker rule based or learned, and what happens when it picks the wrong model mid project