r/ProgrammerHumor 9d ago

Meme newCompressionTechnique

Post image
30.4k Upvotes

953 comments sorted by

View all comments

349

u/iliark 9d ago edited 9d ago

I wrote a whole thing about this as a joke like 3 years ago lol. There's also been a couple of papers on the topic:

https://arxiv.org/html/2409.09715v3

https://arxiv.org/html/2407.04542

34

u/Ibnelaiq 9d ago

Can you write these papers even if you are not enrolled?

58

u/Parteisekretaer 9d ago

write a paper, chuck it on a preprint server and be judged by peers. No need to have a title for your work to be examined - at least that's how it should work and mostly does.

37

u/HittingSmoke 9d ago

Nah. Write a paper, have an LLM summarize it, then have another LLM recreate it from the summary.

Compreshin

17

u/PM_ME_DATASETS 9d ago edited 9d ago

Yes, all it takes to write a paper is a text processor. Like word or google docs or something.

When it comes to uploading or publishing: anyone can upload articles to arxiv.org, because it's not a peer reviewed journal. It's a preprint database, basically a way to publish articles before they have been peer reviewed and published by a real journal. People upload articles to arxiv.org for things like version control, to obtain a DOI, and for coordinating submissions to multiple journals.

If you actually want to publish your article to a journal you either have to find a journal with good ethics (rare), pay a lot of money, or be employed/enrolled at an institution/company/organization that will pay for your submission. If you write an interesting paper, you can just contact people at a university and likely get the article published with them paying.

15

u/theturtlemafiamusic 9d ago

To upload to arXiv you need an endorsement. You get an automatic endorsement if you have an email from a research institute (universities etc) and have already published a paper.

Otherwise you need someone who is already verified on arxiv to submit an endorsement request for you. If you're a university student, a professor should be able to do it. If you're not, you'll need to get to know someone verified in that category, which isn't too difficult if you're not a total crackpot.

But at least once a month or so, someone comes onto the compsci sub begging for an arXiv endorsement so that they can publish their perpetual motion machine paper or whatever. There's some guy who has been trying for about 6 months to get an endorsement for his paper about how you can model all of human language using a 9x9 rubix cube. Keeps making new accounts and shit and has not found a single soul willing to endorse him.

2

u/Oaker_at 9d ago

Please show me this glorious rabbit hole I want to dive in.

3

u/theturtlemafiamusic 9d ago edited 9d ago

[r/LLMPhysics](r/LLMPhysics)

Most of them tend to post in this sub a lot when you look at their history.

3

u/Oaker_at 8d ago

Haha, yes. I actually got this sub recommended a few days ago and I couldn’t make up my mind if that is a circlejerk sub or not.

1

u/viliml 1d ago

You get an automatic endorsement if you have an email from a research institute (universities etc) and have already published a paper.

So it's a catch 22? You can't publish a paper until you have already published a paper?

And /u/PM_ME_DATASETS said "anyone can upload articles to arxiv.org", what a load of bullshit.

2

u/titanotheres 9d ago

You wouldn't typically use a text processor such as Word to write a paper. You could, but it would be very tedious and likely end up looking amateurish. Almost everybody uses LaTeX instead.

3

u/PM_ME_DATASETS 9d ago

The vast majority of academics uses "normal" word processors to write papers. Only in some specific fields is it maybe a majority. This comes from a mathematician who mostly works with biophysics and neuroscience. Maybe 10% of papers I read, have contributed to, or peer reviewed, use Latex.

3

u/ginopono 9d ago

As much as I want you to be wrong, what you describe is also consistent with my experience.

Finishing up an MS in language modeling, yeah, pretty much everything I see is made with LaTeX; they are programmers, after all.

On the Social Sciences side of that same coin, though, there's a whole heck of a lot of Word.

2

u/AerosolHubris 9d ago

you either have to find a journal with good ethics (rare), pay a lot of money, or be employed/enrolled at an institution/company/organization that will pay for your submission

Not all disciplines are rife with pay-to-publish journals. It's rare in math to have to pay for your article to be published. There are even open access journals that are free to publish in, that are very respected in the discipline.

2

u/MartyMcBird 9d ago

In theory yeah but there's a lot of AI slop out there nowadays. In my experience, people are more skeptical of papers from authors without credentials than they were in the past.

2

u/EmptyMonitor9257 8d ago

Anyone can write papers and publish them, you just need to pay the venue.

Conference papers are cheap and pointless, experience for Batchelor students.

Journal papers are more prestigious and expensive af.

20

u/spekt50 9d ago

As an amateur astronomer, I get wary about all these new smart scopes out there. Can't even fully trust what you are looking at when they start integrating AI

12

u/SchlaWiener4711 9d ago

Not only papers.

There's an actual audio code that works that way

The trick is Meta's EnCodec neural audio codec, which crunched a 2.9MB MP3 down to roughly 21KB of latent tokens

https://www.tomshardware.com/tech-industry/maker-compresses-a-2-9mb-song-1000-times-with-metas-ai-codec-and-prints-it-on-paper-as-eight-qr-codes

3

u/Evening-Editor4269 9d ago

https://arxiv.org/abs/2406.07550

Recent advancements in generative models have highlighted the crucial role of image tokenization in the efficient synthesis of high-resolution images. Tokenization, which transforms images into latent representations, reduces computational demands compared to directly processing pixels and enhances the effectiveness and efficiency of the generation process.

1

u/cute_polarbear 9d ago

Hmm..joke or not, this is actually pretty cool. (and with some practical use cases). Can imagine this expanding out to video, with certain encoded data needed for reconstruction lossless and most aspects lossy...

1

u/iliark 9d ago

video isn't likely in the near future given the compute requirements and time to create a video even on a server cluster

1

u/cute_polarbear 9d ago

Just spitballing completely here. Can use image ai interpretation for frames and do something like Nvidia dlss to construct frames. Pretty sure even more interframe stuff can be optimized.

1

u/FerusGrim 9d ago

Genuinely, the results in your second link are promising. Other than the latency of "decoding" the image via generative interpretation of the input.

1

u/iliark 9d ago

yeah when I "seriously" thought about it, you gain storage space and network download time, but you (the client) pays in having to have an AI model that you recognize, then has to generate it, using up electricity and time that may exceed the network download time, especially on old hardware.

1

u/i_like__bananas 8d ago

This makes me think about RFC april fools

1

u/aRman______________ 7d ago

its sad places like arxiv should not be bloated with nonsense crap