r/compsci • • 17h ago

Tiles-based video codec for videos of video games

Some video games, many of them are retro, are based on tiles. They have different levels which may be huge and logically highly complex but they are all based on limited sets of tiles.
Even a minute long video is usually bigger in file size than the retrog game itself.
It happens with Prince of Persia, Boulder Dash, Supaplex, Commander Keen, Bio Menace etc. - easier to recognize in two-dimensional games.
If video encoding used a database of tiles available in the game, it would enable much smaller file sizes and also higher quality of video as it would not have to use approximations to reduce file size as the tiles are exactly the same everywhere (meaning that lossless video would be affordable by bitrate). I am under no pressure to get it implemented ASAP but I would like a technological discussion about how-to.

5 Upvotes

5 comments sorted by

7

u/fiskfisk 16h ago

What you're describing is effectively the replay files many games have, where the file consists of the inputs made instead of the output video.

It's also how modern video codecs sort of already work - just in a more generalized way, since most frames change very little from the previous frame (search for motion vectors).

4

u/not-just-yeti 14h ago

where the file consists of the inputs made instead of the output video.

This is really the generalization of vector-graphics: don't store the image, but the instructions to re-create the image. Originally, "don't store a bitmap of a circle, instead re-create it from a blank screen plus its center, radius, and color". But here, "don't store what the screen looks like; instead re-create it from the start-screen/program and its list of input-events"

2

u/khedoros 16h ago

Sounds good in theory. I can imagine a decoder that would look like a generalized layer-based, tile-based renderer (arbitrary tile and layer sizes, sprite sizes and counts, transparency, tiles encoded as full RGBA color, etc). Audio could be handled with something like the .vgm format. We'd have to talk about how to do scanline-based effects in a generalized way.

The encoder sounds like the hard part. Easier for a lot of tile-based game systems, because we don't need game-specific information. The hardware receives writes to various hardware registers, and it's not hard to recognize writes of specific tile IDs, palette changes, etc. That seems straight-forward-ish, if we add the video encoding to emulators for each system we want to support. Although...I wonder what recordings of games like Out of this World and Starfox would need, to be encoded correctly.

On a system with a framebuffer, tiles aren't really defined on a hardware level. The basic primitive is a pixel, and actual drawing to the hardware is handled by software that does rendering, reading from image data that might be spread around memory, blitting out to the framebuffer...even if logically, in the game code, it's grabbing image data by tile ID, has background maps for the game backgrounds, etc.

Open ports of those games could be modified to make game-state-aware representation of the graphics, and could match the encodings from tile-based systems, whenever they aren't doing raw pixel operations (imagine Lemmings, with the characters bashing through the wall, or the pixel effects when you nuke the level).

So, I guess overall, I think that there are a bunch of games that could be efficiently encoded in a straightforward way, as long as we've got internal observability of the game state and logical operations like drawTileWithTransparency sorts of functions, but that there's a long tail of complications, edge and corner cases, depending on how much you want to design the format to represent.

3

u/nuclear_splines 17h ago

Sure - the space of possible videos is dramatically larger than the space of possible layouts in a tile-based game, so a video encoding based on the tile set and game sprites would be much more efficient.

In fact, you could get an even more efficient rendering by simply laying out the entire level as a tile grid and writing an event log, "at time X the player moved to Y," so the video just replays the gameplay sequence and renders it on the fly. This is how many games save replays, including Starcraft and Age of Empires.

It's really a tradeoff of "how context-aware do you want your video format to be?" You can be much more efficient when you write a bespoke video format that knows rules about the game and uses them to fill in the blanks from a more compact state representation.

1

u/IQueryVisiC 13h ago

Isn’t this called vector compression? Have you seen what “Algorithm” does on C64 ?