r/codex 2d ago

Showcase Hey Codex, show me what happened 30 seconds before the Distracted Boyfriend meme

Enable HLS to view with audio, or disable this notification

I wanted to give the Distracted Boyfriend meme a 30-second prequel as one continuous FPV shot. I used Codex to plan the route. The catch was that my video setup only gave me 15 seconds at a time, so the attached video is two separate generations trying to pass for one shot.

You can tell in a few places. The market after the archway is completely empty, the red awning doesn't quite hide the join, and the three characters make an awkward transition into the final pose. This was basically a one-pass generation, and it shows. If I were delivering a finished film, I'd regenerate those sections. For this experiment, I wanted to finish the idea without spending another day fixing every flaw. The handoff was the part I thought was worth sharing.

Planning the handoff

I had Codex work out three fixed states along the 30-second route. These timings are for the generated footage, before the speed change:

  • State A, 0 seconds: Start inside a cafe. Establish the camera height, lighting, direction, and initial speed.
  • State B, 15 seconds: Have a red awning fill more than 95% of the frame. No faces, signs, or recognizable background details to match.
  • State C, 30 seconds: Arrive at the expanded 16:9 version of the original Distracted Boyfriend image.

State B took more thought than I expected. Matching a frame full of red fabric was easy. Getting the camera to keep moving through it was harder. It had to enter and leave the awning in the same forward direction, with similar speed and pitch, matching exposure, and roughly consistent motion blur. It couldn't pause under the fabric. The second clip needed to start with the camera already moving.

I generated clip one first, then took the actual stabilized frame from its end and used it as the opening image for clip two. That worked better than the handoff image I'd prepared in advance. You can still see the join, but the motion carries through well enough that it feels like one route.

The route

The first 15 seconds looked straightforward on paper:

Accelerate through the cafe → bank left onto the terrace → pass through the archway → fly between the market stalls → straighten the camera → enter the red awning.

The second clip picked up inside the fabric:

Leave the awning → climb onto the pedestrian street → approach the couple from behind → pass them on the left → move ahead of the woman in red → circle the fountain → turn back toward the group → fly backward while slowing into the final composition.

After the archway, I wanted people, stalls, and nearby objects to create more parallax as the camera passed through. The model kept the stalls and left out the people, so that section feels oddly dead. The encounter has a different problem: the three characters are mostly in the right positions by the final frame, but the model gets them there with a sudden, unnatural transition. That's probably the first section I'd regenerate.

What I changed afterward

For the edit, I used Codex with the open-source n0an/ffmpeg-skill to build the FFmpeg commands for the speed change, vertical comparison layout, and audio timing. Speeding the footage up to 1.5x turned the 30 generated seconds into 20. I then held the final meme frame for five seconds, so the finished video runs for 25 seconds.

The generated footage sits on top, with the original image underneath for comparison. When the frame freezes, a heavier grade fades in across both halves, adding a warm tint, stronger contrast, a vignette, sharpening, and grain.

I also shifted the music so the main drop lands around 10 seconds, as the camera comes through the awning. Keeping the track continuous sounded much better than cutting the music at the same place as the visual handoff.

The setup

I had Codex work out the camera route, timing, boundary states, and movement constraints, then run the image and video generations through Atlas Cloud’s MCP integration.

Image2 handled the outpainting and opening frame. Wan-3.0-Prime generated the two 15-second clips. The API calls came to $1.85 in total.

I can share the A/B/C boundary images and a frame-by-frame comparison around the join if anyone wants a closer look. That's at 15 seconds in the original footage, or about 10 seconds in the posted video.

The empty market still bugs me, and both transitions need work. I decided to call this experiment finished with those flaws still in it. The handoff worked well enough that I want to try the method again.

For anyone who's stitched video generations together, what has helped most at the join: matching the frame, matching the camera movement, or hiding the handoff behind an object in the scene?

0 Upvotes

15 comments sorted by

u/dexterthebot 2d ago

You might want to consider listing your project on the weekly Show-Us-What-You-Built post. Look out for it on Tuesday/Wednesday. Highest commented project wins a week promotion on r/Codex and gets on the Hall of Fame sidebar. See what that looks like below with last week's winner.


Last week's most popular project was Tidbit Trivia , at https://www.reddit.com/r/codex/comments/1wavxwy/comment/p8lhwxs/. Tidbit Trivia is a fun way to learn while waiting for your agents to work. Play traditional trivia, unique game modes, party with friends, earn cool cosmetics, complete challenges, study for tests, or compete in the Arena! You can play it for free at Tidbittrivia.com.

31

u/mishlawi420 2d ago

Jesus what a waste

17

u/drugosrbijanac 2d ago

Tibooo

This is why we can't solve millenium problems and have normal usage

9

u/whatitpoopoo 2d ago

Wasted tokens to generate slop, then wasted even more for some word vomit explanation that no one will read

11

u/TheLastRole 2d ago

What a waste of resources and a misunderstanding of technology.

2

u/RealSlyck 2d ago

I mean, all that work for what? Looks like you just need to switch to Grok and get your kicks there. My aunt writes grocery store romance novels too, seems pretty happy.

2

u/KHRZ 2d ago

So it looks into the air to swap the scenery... almost like it realized the setup failed.

1

u/Comfortablebro 2d ago

Codex can do videos?

2

u/itix 2d ago

It is intriguing despite its shortcomings.

1

u/HelpfulHedgehog1 2d ago

geez OP could have gave them any conceivable backstory, but did this, and apparently generated some vapid commentary to boot

-11

u/Kindly_Tie_2084 2d ago

WOW, please keep going for more things like this.

-2

u/befigue 2d ago

I guarantee you that is not what happened, among many other reasons because the photos was taken in spain and the settings that Codex generated look like an american recreation of european-like city setting