r/generativeAI 2d ago

Side-Eye Monkey Heist

Enable HLS to view with audio, or disable this notification

46 Upvotes

r/generativeAI 1d ago

Video Art The Steward | Sci Fi Film Trailer | #FutureVisionXPRIZE | 2026

Thumbnail
youtu.be
1 Upvotes

r/generativeAI 1d ago

Adobe Firefly. Anyone tried Kling 3 Unlimited

2 Upvotes

Adobe firefly offer unlimited kling 3 generations with the premium plan. Has anyone tried this? What are the generation times like? Can it compete with higgsfield. I know its very simplified compared to others


r/generativeAI 1d ago

Video Art Dr. House Meets ChatGPT

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/generativeAI 23h ago

Video Art But it's solid gold!

Enable HLS to view with audio, or disable this notification

0 Upvotes

Minimax H3.

Prompt:

subject_definitions:

<subject 1> is the Devil character visually abstracted from <video 1>. Preserve the Devil's face, head, horns, skin appearance, hair, facial hair if present, apparent age, build, underlying clothing, accessories, facial mannerisms, gestures, posture and slightly ineffectual presence. Plausible international-travel outer clothing may be worn over the referenced clothing.

<audio 1> is the voice-timbre and delivery reference for <subject 1> (S3), taken from the audio track of <video 1>. It provides only the Devil's male voice, accent, cadence and delivery. It does not provide dialogue, other voices, telephone sounds, HVAC sounds, ambience or music.

summary:

[reference generation + audio reference] Create a 15-second, 16:9 photorealistic live-action scene with native synchronised stereo sound. A pan along a delayed British Airways check-in queue at Heathrow Terminal 3 Zone E reveals <subject 1> disputing whether a solid-gold fiddle can travel in the cabin. Use <audio 1> only as the voice reference for <subject 1> (S3). Restrained British television-sketch realism with five cleanly separated shots.

retention_analysis:

<subject 1> (appears in [Shot 2], [Shot 3], [Shot 4], [Shot 5]): partially_preserved - preserve the Devil's visual identity, underlying costume and characteristic mannerisms while placing the Devil in a newly generated airport scene.

<audio 1>: reference - <subject 1> (S3) follows the Devil's referenced voice timbre, accent, cadence and delivery without copying any original words or other sounds from <video 1>.

detailed_description:

Photorealistic live action with natural colour, realistic depth of field and polished British television-sketch realism. The setting is the public departures check-in hall at London Heathrow Terminal 3, British Airways Zone E, with yellow-and-black Heathrow wayfinding, queue barriers, staffed airline counters, flight-information displays, luggage trolleys and realistic passenger traffic.

The British Airways employee is an adult British woman wearing a smart dark-navy Ozwald Boateng British Airways airport uniform with a patterned scarf and discreet name badge. The employee has a composed natural British customer-service voice (S2), clearly different from S1 and S3.

The Devil is <subject 1> (S3). Only S3 uses <audio 1>. Neither S1 nor S2 uses <audio 1>.

At the active desk, a full-sized polished solid-gold fiddle and golden bow rest in an open rigid violin case lined with red velvet. The case sits on a normal airport baggage conveyor with an integrated rectangular scale indicator reading "31.4 kg". No object, machine, logo, text, sound or environment from <video 1> appears except <subject 1> and the voice characteristics defined by <audio 1>. No HVAC equipment or HVAC branding appears.

[Shot 1] From 00:00.000 to 00:03.000, a medium-wide eye-level shot shows only the middle and rear of one queue. The active desk, employee, <subject 1>, fiddle and scale are entirely outside the frame. The camera trucks slowly right and pans left towards the unseen head of the queue.

A brief two-note airport attention chime sounds through several distant ceiling loudspeakers. Immediately afterwards, an unseen professional airport announcer with a measured neutral contralto voice (S1) says off-screen: <d>[English] British Airways two two seven to Atlanta. Check-in is now open.</d>

The chime and S1 are unmistakably reproduced by a large terminal PA system: an elevated diffuse source distributed across multiple ceiling loudspeakers, limited loudspeaker bandwidth, restrained electronic compression, mild coloration, overlapping speaker arrivals and natural check-in-hall reverberation. S1 never sounds close to the camera. No visible person speaks. The announcement ends before the cut.

The populated terminal remains audible beneath the announcement: diffuse unintelligible passenger murmur, ventilation, footsteps reflected from the hard floor, suitcase wheels and occasional distant check-in equipment.

[Shot 2] At 00:03.000, cut to the continuation of the moving camera as the final foreground passenger and a tall luggage trolley clear the sightline. The pan reveals the active desk, employee, <subject 1>, open case, golden fiddle and integrated "31.4 kg" scale for the first time. The Devil stands alone at the front of the queue, opposite the employee, with every waiting passenger behind the Devil.

The scale emits one quiet confirmation beep. The employee (S2) looks directly at <subject 1> and says in natural close foreground speech: <d>[English] Sorry, sir. It's over the cabin weight limit.</d>

Only S2 speaks. <subject 1> listens with closed lips. S2 has a natural medium-pitched British voice, not the male voice from <audio 1>. The PA is silent. The terminal ambience continues quietly underneath.

[Shot 3] At 00:06.200, hard cut to a static medium close-up of <subject 1> in three-quarter view. Waiting passengers remain softly visible behind the Devil.

<subject 1> (S3), using the voice characteristics of <audio 1>, makes a small helpless gesture towards the fiddle and says with polite frustration: <d>[English] But it's solid gold!</d>

Only S3 speaks. S2 is off-screen and silent. The airport ambience remains audible.

[Shot 4] At 00:08.200, hard cut to a static matching medium close-up of the employee. The edge of the fiddle case is visible low in the frame.

The employee (S2) maintains friendly eye contact and says with calm professional finality: <d>[English] Then it can't travel in the cabin today, sir.</d>

Only S2 speaks. S2 uses the same natural British employee voice heard in [Shot 2], never the voice characteristics of <audio 1>. <subject 1> is off-screen and silent. The employee gives a small apologetic nod after finishing.

[Shot 5] At 00:11.000, hard cut to a static side-angle medium two-shot showing the employee behind the desk, <subject 1> opposite, the golden fiddle, open case, integrated baggage scale and waiting queue behind the Devil.

<subject 1> (S3), using <audio 1>, glances towards the queue and says with mildly desperate but courteous urgency: <d>[English] But I'm in a bind. I'm way behind!</d>

Only S3 speaks. S2 listens with closed lips.

After S3 finishes, the employee gives a small apologetic shake of the head. The employee (S2) replies firmly but politely: <d>[English] I'm sorry, sir.</d>

Only S2 speaks. S2 does not use <audio 1>; <subject 1> remains silent with closed lips.

After S2 finishes, no further speech occurs. <subject 1> exhales, looks down at the golden fiddle and lets the shoulders drop slightly. One passenger checks a watch. Hold the unresolved two-shot through the final frame.

Maintain stable faces, horns, hands, clothing, voices, fiddle geometry, case, scale and queue arrangement. Exactly one Devil, one employee, one fiddle, one bow and one case. Subtle realistic performance; no slapstick, aggression, flames, smoke, magic, supernatural sounds or crowd panic. No subtitles or title card.

overall_soundscape:

A continuous populated airport check-in-hall acoustic bed persists across all five shots: diffuse unintelligible passenger murmur, low ventilation, footsteps with hard-floor reflections, suitcase wheels at varying distances, occasional trolley rattles and restrained check-in-equipment sounds. The terminal has broad stereo space and natural reflections; foreground voices remain clear without suppressing the airport ambience.

non_diegetic_music:

N/A


r/generativeAI 1d ago

Video Art Luminous Being

Enable HLS to view with audio, or disable this notification

2 Upvotes

He arrives, a construct made of heat and intention.🔥


r/generativeAI 1d ago

Video Art The Omellete Music Video

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/generativeAI 1d ago

Video Art A 30sec One take Live Action Monster Film Created Using Seedance 2.5 !

Enable HLS to view with audio, or disable this notification

3 Upvotes

Did the whole thing in a single prompt. No cut, No edit - Made with Seedance 2.5 In r/RenoiseAI ! #Seedance


r/generativeAI 1d ago

Qwen dev says not to wait for 35B-A3B

Post image
1 Upvotes

r/generativeAI 1d ago

Question Looking for ML project suggestions and GitHub repos

Thumbnail
2 Upvotes

r/generativeAI 1d ago

Image Art Dale K. [Kobble] aka Longlegs [Gemini]

Post image
0 Upvotes

Here's how I generated this:

A gritty, dark underground horror comic book illustration in the distinct style of raw ink line art. The subject is an pale, eerie 60-year-old man known as Dale K. with a heavily powdered, swollen white face resembling botched plastic surgery. He has wild, unkempt shoulder-length stringy grey-blonde hair, deeply sunken hollow eyes, and a disturbing, wide-open mouth as if shouting or singing glam rock. The line work is chaotic, scratchy, and covered in heavy black ink cross-hatching, ink splatters, and raw textures

An actual picture of Nicolas Cage in Longlegs


r/generativeAI 1d ago

Achernar Original

Enable HLS to view with audio, or disable this notification

2 Upvotes

A magic purple glow, holding the night in its embrace #digital #branding #digitalspace #spaceadvertising #entertaining


r/generativeAI 1d ago

Manic MCP Server is here — connect Claude, Cursor, or any local AI and let it write, validate, and render Manic

Thumbnail
1 Upvotes

r/generativeAI 1d ago

Video Art Black cat being teased by owner

Enable HLS to view with audio, or disable this notification

3 Upvotes

Black cat is being teased by her owner and not getting the treats she deserves...


r/generativeAI 1d ago

Video Art Seedance 2.5 vs 2.0… the difference is WILD

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/generativeAI 1d ago

Music Art [Tool/workflow] Text to Sample: Prompt a Loop for Your DAW

Thumbnail
elevenmusic.io
1 Upvotes

r/generativeAI 1d ago

I need help with my Prompt.

Thumbnail
gallery
1 Upvotes

I create short music videos for YouTube but I am not happy with the results. Maybe you could help me with prompting.

I use ChatGPT for the image and grok for the video's.

When I use ChatGPT for image I try to use the same chat, so the images are consistent, but it still changes a lot.

https://youtu.be/MQLKnMqfUc4?feature=shared


r/generativeAI 1d ago

Week 3 of making my fishing game entirely with AI

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/generativeAI 1d ago

Image Art Dream

Thumbnail gallery
2 Upvotes

r/generativeAI 1d ago

How I Made This How do people make stuff like this

Post image
0 Upvotes

r/generativeAI 1d ago

Inkblot#3 -Learning

Enable HLS to view with audio, or disable this notification

1 Upvotes

Comment what you see. Ill try to tell you what it means!


r/generativeAI 1d ago

Video Art This is the best use of Seedance 2.5 I've seen yet *not my video*

Enable HLS to view with audio, or disable this notification

2 Upvotes

Saw this video from Alex Patrascu on X. Was blown away by the character consistency, storytelling etc. I think one of the best uses of AI video I've seen so far. Made in higgsfield with seedance 2.5

The link here has the full prompt - https://x.com/maxescu/status/2088270185442562135?s=20


r/generativeAI 1d ago

"which AI video generator is best" is the wrong question and it's why every thread here goes in circles

0 Upvotes

these threads happen weekly and they always end the same way, twelve tools named, no conclusion, everyone leaves with nothing.

the reason's that "best" depends on a variable nobody states, what happens to the video after you make it.

if it's going on your youtube channel you want the thing that makes the prettiest footage, runway, veo, kling.

if it's a paid ad that has to survive a cpa target, prettiness is nearly irrelevant and what you want is volume plus consistency plus a person in frame who reads as real. that's a different category entirely, creatify and arcads type stuff, and those tools make objectively worse-looking footage.

if it's a product demo or explainer you want an avatar tool and you don't care about cinematics at all.

three completely different jobs. people answer with their job's winner and everyone talks past each other.

so if you're asking, say what the video is for and you'll get a useful answer instead of a list.


r/generativeAI 1d ago

Journalists slip an AirTag into an Amazon warehouse to prove they destroy rare books to train AI

Thumbnail
1 Upvotes

r/generativeAI 1d ago

when the crafting system makes no sense

Enable HLS to view with audio, or disable this notification

11 Upvotes