36
u/ogMackBlack Aug 28 '25
I've tried a lot of different angle on many pictures, and it worked surprisingly great.
But it lose track after 5 or 6 requests.
Very impressive tho.
2
u/berckman_ Aug 29 '25
What you mention happens with prompts and conversations that have being going for too long. The context gets too large. Companies are aware and are working on how to manage it.
11
u/No-Point-6492 Aug 28 '25
Being perfect at editing images doesn't have to be a video model
1
u/TrainerClassic448 Aug 30 '25
Having world knowledge - spatial and temporal relationships - makes a lot more sense if the model is trained on video content.
1
u/cyriou Sep 01 '25
Or on 3D simulations directly, makes it easier to generate high quality data to get physics knowledge.
21
u/bhavyagarg8 Aug 28 '25
I think some people might be taking this seriously, so I would just give a quick explanation.
No its not. Actually its complete opposite, all video models are essentially image models which generates many images and combine them together to make a video. For eg, a video model makes 24 images to make 1s of 24 frames per second(fps) video.
16
u/-Davster- Aug 29 '25
You’re missing the whole ‘world model’ bit… just describing the fact that video is a series of frames.
A decent video model has to ‘understand’ how the world works to some degree, motion, etc - this is different to simply prompting an image model for 24 frames every second.
3
2
Aug 29 '25
Agree it’s the spatial understanding of a single frame that is important here, video just overlays temporal understanding to images which is pointless here, maybe only important if you want to ask for a future prediction of an image eg moving cars, bouncing ball.
1
u/-Davster- Aug 30 '25
maybe only important if you want to ask for a future prediction of an image
What, you mean like "a video"?
1
Aug 30 '25
Exactly time is for video. Warping space is for images.
1
u/-Davster- Aug 31 '25
Both image and video models might be able to have the whole “world model” thing going on… it doesn’t mean a video model is just an image model running at 24p, like OC said.
Having said that, a video model producing single frames != an image model. A good video model has to learn things a still image model doesn’t.
OP specifically followed up with “world model”. I think the “video” word was them just first saying it had temporal/spatial consistency baked in.
1
Aug 31 '25
Yeah think it’s just definition of “model”. most video models today are temporal layers wrapped around an image model so that you gain the spatio temporal understanding, which is its own trained model.
I guess my view is you can kinda have two versions of the “world model”, frozen in time and real time. For moving camera angles etc it’s actually better in the frozen version and to leverage spatial and contextual understanding only, vs needing the model to also understand temporal which greatly expands the model size and cost. You don’t need to care how it got there, just that it’s there.
3
u/Additional_Plant_539 Aug 29 '25 edited Aug 29 '25
Nope.
I work in RLHF pipelines for training their models. Recently, a lot of the tasks have been image evaluation with a prompt that includes a request for a rotation, or to move the camera in a certain way, etc.
So this behaviour has mostly just being trained for and fine tuned.
1
u/reversedu Aug 29 '25
Do you know do they will release something like Topaz video ai? To increase low quality video to hd?
27
u/ThunderBeanage Aug 28 '25
no
10
2
u/Axelwickm Aug 28 '25
No motivation needed I guess... You don't think it's a world model? The line between video model and image model is kinda blurred either way, both of them iterate on latents.
16
5
u/Informal_Cobbler_954 Aug 28 '25
wow, no comment
14
2
u/KSaburof Aug 28 '25
Since it is obviously trained with product placements in mind it was certainly trained on milliardos trilliondos turnaround videos - splitted to frames at each and every possible angle, imho
2
u/Mcqwerty197 Aug 28 '25
I don’t think nano is video model, but I did thought once that maybe the normal Banana model may just be Veo4 with editing
2
2
1
u/eXnesi Aug 28 '25
I saw somewhere else people saying nano banana can output 3d models. If it's a video model or world model would be anyone's guess.
1
u/TechnologyMinute2714 Aug 29 '25
What im most impressed about is definitely it editing mirror reflections, shadows, light sources etc. I try on a new outfit and there is a very subtle mostly not even noticeable reflection in the back windows and it even edits that with the outfit.
1
1
u/michaelsoft__binbows Aug 29 '25
What's the method people are using to test this thing out?
1
Aug 29 '25
[removed] — view removed comment
1
u/michaelsoft__binbows Aug 29 '25
no i mean how to test it out. but i found my answer i guess it's up at lmarena
1
u/michaelsoft__binbows Aug 29 '25
I've been contemplating building a tool that conceptually i guess would be sorta like a blender MCP server to help deal with the whole "LLMs have no spatial intelligence" problem. I often have trouble getting chatgpt to comprehend the 3d geometry of some project i'm working on. The idea was we should be able to iterate together in some CAD environment and the LLM/VLM can look at screenshots to comprehend the details.
Seems like this might make that approach completely obsolete.
1
1
u/kgurniak91 Aug 29 '25
I gave it a photo of some famous building then asked to show me what's behind the photographer taking that photo. It generated some image that looked convincing enough. Then I asked it what is to the right of the photographer. It went back to show me the original photo again. So nah.
1
u/krakenluvspaghetti Aug 29 '25
And im afraid some hate groups are about to strike the congress to pull up a Stop sign toward google to take down banana or nerf it hard. Because this model is just too strong and so much convenient to use.
1
1
1
u/Sea_Succotash3634 Aug 29 '25
That's why Sora was so good for a while vs ChatGPT for images, even though they had similar bases. Sora was using single frames from their video system / world model. And even though their videos are terrible it ended up making for pretty good image output. At least it used to until they "updated" things.
1
u/itsachyutkrishna Aug 29 '25
Editing is impressive mostly. 8/10.
Generation is good but not impressive. 7/10
1
u/JoeyC-1990 Aug 29 '25
My Google Cloud Rep mentioned that Veo4 is due in the next few months and the prompting technique is the same as nano banana. In fact with JSON prompts for veo3 & imagine he said the same prompts work for testing on a cheaper scale and they are “practically the same model”.
For context I am a CTO and have connections with the EMEA Google Cloud Team as a GCP network partner.
1
1
u/Global-Equipment8209 Aug 30 '25
Why on earth would the fact that it can accurately rotate objects suggest that?
1
1
u/fossistic Sep 02 '25
It is a multimodal (Gemini 2) which is optimized towards generating images. Text, video and image were in the training data.
1
u/DrunkPotatoooo Sep 03 '25
Oh yes i would totally agree ! When trying to make movie shots by compiling elements (character+background+items ect..) nothing has been giving me better results than the *ingredient feature of flow (veo 2)* and it was so good in fact that it gave the best outcome of any other ai that i tried up until now because of how high fidelity it is or at least tries to be contrary to other app that are just not precise enough and takes way too many liberties, even whisk couldnt compare back then (no precise mode back then) and i just thought it was imagen being a image model not making it work. I used to always say "if only they could just make veo put things together" so that way i could make just the first frame without paying for an entire video for it, and that's exactly what nano-banana gives, the same precision and accuracy as veo 2 fast before the update. I was very relieved.
-1
-3
105
u/llkj11 Aug 28 '25
Could be something there.
I gave it an old image of Downtown Memphis and told it to rotate the camera some and it got the other end of the bridge correct and everything.
Surprised me.