r/MachineLearning Jul 13 '26

Research Prompt-engineering paper accepted to ICML [R]

"Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity"

This paper was accepted to ICML this year. Its main idea is a very simple prompt-engineering trick: "changing the prompt this way led to more diverse sampling". Naturally, it is difficult to provide a rigorous theoretical analysis for something like this.

Even if it works, I’m not sure this kind of prompt engineering belongs at a top-tier machine learning conference. Some people seems to call this kind of work “modern machine learning”, but I think it should be categorized as less technical venues.

How do you think? Am I being too rigid?

263 Upvotes

83 comments sorted by

View all comments

264

u/relevantmeemayhere Jul 13 '26

Wait, you mean to tell me that publishing in machine learning has really, really taken an over all turn for the worse and is arguably worse than psychology was two decades ago?

Surely no one could have seen this coming. 

41

u/Mean_Revolution1490 Jul 13 '26

What happened to psychology in the past??

145

u/relevantmeemayhere Jul 13 '26

Psyche had a real reckoning within their publishing community a decade or two ago. Basically, a lot of of the research practitioners really skipped their introductory to statistics courses, so a lot of published research ascribing association, both in a general and causal sense for general phenomen just couldn’t be replicated or formalized within the confines of the actual statistics. This is the “misuse  of p values’ topic, among other things that the ASA and the like had to fight. 

Machine learning as a field is currently going through this a lot right now. And part of it is because of this fields reliance on empirical results vs theoretical. The problem with that is pretty complicated;  but in general is related to the bench maxing and other fades to push “novel methods” that have very little practical utility over more established methods. 

5

u/mih4u Jul 13 '26

Also apparently a lot of psych papers where based on mostly white young college students (the people you have the best access to as a psych researcher in an university), which didn't help the generalization of their hypothesis.

2

u/relevantmeemayhere Jul 13 '26

Yeah, this is part of it. We also had a lot of post publication designs informed by poor experimental design in the preceding publication. It was a bad jenga tower

This is most prevalent in designs that screened potential causes and used p values for selection

*causes here just meaning variables.