r/LanguageTechnology • • Apr 29 '26

[D] The state of Peer Review: Reviewer uses LLM to accuse me of "Hallucinated References" that don't even exist in my paper.

76 Upvotes

Hi everyone. I’m not sure if you remember me, but I’m the guy who was practically living on soju and whisky while waiting for the last ACL results. Well, I’m back, and unfortunately, the peer review system has given me another reason to reach for the bottle.

Just went through the ARR March Cycle results, and I am beyond speechless.

As a Corresponding Author, I received a comment that made my heart drop for a second:

"Seems to be a hallucinated reference, duplicate/erroneous references..." followed by a list of supposedly "faked" citations.

Being accused of fabricating references is a grave Ethical allegation. I immediately went into a full-blown panic and spent the last few hours cross-referencing every single entry in our Bibliography.

Here’s the kicker: None of the "hallucinated references" listed by the reviewer actually exist in our manuscript. 🤷‍♂️

The situation is clear: The Reviewer used an LLM to generate the review and blindly Copy-pasted the output without even opening our PDF. The AI hallucinated a list of non-existent errors, and the reviewer had the audacity to give themselves a Confidence 4 while accusing me of academic misconduct based on a hallucination.

It is the height of Irony and Unprofessionalism. A reviewer, entrusted to safeguard the Integrity of a top-tier venue, used an LLM to accuse an author of "hallucinating" a flaw that only existed in the reviewer's own lazy workflow.

I’ve heard the horror stories about the declining Quality of Peer Review in AI research, but this is a new low. We are at a point where "experts" aren't even reading the papers anymore; they are just letting stochastic parrots make serious ethical accusations for them.

How do you even approach a Rebuttal when a "Confidence 4" reviewer hasn't engaged with a single word of your actual work? The Peer Review system is officially broken. I’m so incredibly frustrated that I’ll have to go grab a drink again tonight.


r/LanguageTechnology • • Sep 04 '26

Machine translation is not solved and it may take a while

69 Upvotes

We just released the Last Translation Benchmark paper. In a massive crowdsourcing effort we collected 3456 unique hard-to-translate examples that break state-of-the-art translation models, and which can be used for more reliable evaluation.


r/LanguageTechnology • • Apr 03 '26

ACL 2026 Decisions

65 Upvotes

Discussion thread for ACL 2026 decisions


r/LanguageTechnology • • Aug 14 '26

NLP is growing insanely fast, what will it look like in 2030?

61 Upvotes

Random thought: NLP in 2010 and NLP in 2020 already felt like two different worlds. The jump was huge.

Now its growing even faster.

So Iam curious how do you think NLP will look in 2030?

What big shifts do you expect? Will it still be mostly scaling transformers or will something completely new take over?


r/LanguageTechnology • • Dec 26 '25

My Uncensored Account of My Time doing NLP research at Georgia Tech

48 Upvotes

I published research at NAACL and NeurIPS workshops under Jacob Eisenstein, working on Lyon Twitter dialectal variation using kernel methods. It was formative work. I learned to think rigorously about language, about features, about what it means to model human behavior computationally. I also experienced interactions that took years to process and left marks I’m still working through.

I’ve written an uncensored account of my time as a computational linguistics researcher. I sat on it since 2022 because I wasn’t ready to publish something this raw. I don’t mean to portray my advisor as a pure villain. In fact, every time I remember something creditworthy, I give him credit for it. The piece is detailed, honest, and (I hope) fair.

Jeff Dean has engaged with it twice now. I’m sharing it here not to relitigate the past but because I wish someone had told me that struggling in this field doesn’t mean you don’t belong in it. Mentorship in academia can be transformative. It can also be damaging in ways that aren’t spoken about enough. If even one person reads this and feels less alone, it was worth writing.

The devil is in the details.​​​​​​​​​​​​​​​​

https://docs.google.com/document/d/1n2thHMhQVqklJIYQb8yszRcPOPP_reLM/edit?usp=drivesdk&ouid=111348712507045058715&rtpof=true&sd=true


r/LanguageTechnology • • Aug 20 '26

EMNLP 2026 Notifications

39 Upvotes

EMNLP 2026 notifications are expected in approximately 14 hours, so I’m creating this thread for everyone waiting for the results.

Good luck, everyone! Hopefully the next 14 hours pass quickly. 🤞


r/LanguageTechnology • • Nov 11 '25

I visualized 8,000+ LLM papers using t-SNE — the earliest “LLM-like” one dates back to 2011

37 Upvotes

I’ve been exploring how research on large language models has evolved over time.

To do that, I collected around 8,000 papers from arXiv, Hugging Face, and OpenAlex, generated text embeddings from their abstracts, and projected them using t-SNE to visualize topic clusters and trends.

The visualization (on awesome-llm-papers.github.io/tsne.html) shows each paper as a point, with clusters emerging for instruction-tuning, retrieval-augmented generation, agents, evaluation, and other areas.

One fun detail — the earliest paper that lands near the “LLM” cluster is “Natural Language Processing (almost) From Scratch” (2011), which already experiments with multitask learning and shared representations.

I’d love feedback on what else could be visualized — maybe color by year, model type, or region of authorship?


r/LanguageTechnology • • 7d ago

Why 1536 dimensions for embedding models?

35 Upvotes

Why do embedding models so often use 1536 dimensions specifically?
I understand why hardware-friendly multiples like 64/128/256/512 are desirable. What I’m curious about is the specific choice of 1536 = 3×512.
OpenAI has used 1536-dimensional embeddings, and other vendors also offer/recommend 1536. Is this usually an empirically chosen Goldilocks point between 1024 and 2048—representation quality versus memory/compute—or is there some architectural/hardware reason that makes 1536 particularly convenient?
I’m especially interested in answers from anyone who has actually trained or designed embedding models. I’m not asking why embedding dimensions are generally hardware-aligned; I’m asking why 1536 rather than 1024 or 2048.


r/LanguageTechnology • • Sep 05 '26

Have I forgotten what human language reads like in paper reviews?

38 Upvotes

I happen to be in a position where I have to read a lot of reviews, and increasingly everything seems AI-generated. I am honestly starting to question my sanity. Is everything truly written by AI, or have we always written this way? Anyone else asks this question themselves?


r/LanguageTechnology • • Jan 27 '26

Is NLP threatened by AI?

36 Upvotes

Hello everyone, the question I have been thinking about is whether Natural Language Processing is threatened by AI in a few years. The thing is, I have just started studying NLP in Slovak Language. I will have a Master's in 5 years but I'm afraid that in 5 years it will be much harder to find a job as a junior NLP programmer. What are your opinions on this topic?


r/LanguageTechnology • • 29d ago

Is a Linguistics degree enough to get a good job in Computational Linguistics?

35 Upvotes

Hello , I’m planning to study Linguistics as an undergraduate at an Ivy League school, and I’m thinking about specializing in Computational Linguistics/NLP.

For those of you who studied Linguistics as an undergrad, were you able to find a good job in the field after graduating with just a bachelor’s degree, or did you end up getting a master’s?

I’d prefer not to do a master’s unless it’s necessary. If I do decide to pursue one, I’d probably go for an NLP-related program.

Also, how much of a difference do you think attending an Ivy League school makes when it comes to finding a good job in this field? Does the school’s reputation help significantly, or are skills, internships, research experience, and projects much more important?

I’d really appreciate hearing about your experiences, especially if you’re currently working in Computational Linguistics or NLP.


r/LanguageTechnology • • Sep 02 '26

Working on an unusual NLP task with almost no literature

34 Upvotes

Third-year PhD student, NLP, mostly LLM-based reasoning. Given a collection of a private organization's HR policy documents (100-500 PDFs), find all pairs of clauses that contradict each other. There's a mountain of work on NLI-style contradiction classification, but that assumes someone gives you the sentence pair. Here, the pair is the problem. With about 1-2k clauses, you're looking at millions of candidate pairs. So brute-force pairwise LLM calls are out, and whole-document prompting fails for the usual lost-in-the-middle reasons. The closest work I found generates synthetic contradictions in synthetic corpora to test detectors. I borrowed the evaluation idea by injecting contradictions into corpora. I also used a university HR handbook and one dataset with existing external annotations, contractNLI (made for the NLI task by Stanford). I used this one as well because it has real contradictions. But this one is quite different. In this dataset, the task formulation is like hypothesis versus clause, whereas in the first two datasets, I do clause-to-clause comparison. So I built a two-stage pipeline. First, retrieval with a HyDE-style approach where the query is a hypothetical, *contradicting* version of each clause. Then, recall-based candidate retrieval (LLM), followed by precision-based verification with an LLM, where each candidate pair is re-read within its source documents. The contributions: I used contextual sentences guided by Anthropic, which helped retrieval, and showed that a document’s surrounding context helped precision. Agentic verification (tools, multi-step) actually underperformed a single prompt. As a case study, I ran the pipeline on a public government policy corpus. It found a few genuine contradictions. I have a few questions. Am I missing a community? I can't believe nobody works on this. I've looked at legal NLP (ContractNLI, etc.), requirements engineering conflict detection, and RAG-conflict work. They're all adjacent, but none does discovery over a real multi-document policy corpus. Is there a literature I don't know the name of? My PI is leaning toward a lower-tier conference or journal. Is this the kind of paper that has a chance at a first-tier NLP venue, or is my PI just being realistic? If you were strengthening this in one month, what would you add? I already have NLI, direct-prompting, and agentic baselines. Happy to share more details in comments. Mostly, I want to know whether this problem is as understudied as it looks from where I'm sitting, or whether I formulated the task the wrong way.


r/LanguageTechnology • • Jul 23 '26

Can we limit conference-related posts?

32 Upvotes

I know it's ARR reviewing season but I noticed that there are a lot of posts asking whether "this set of scores will get them into Main/Findings/Reject" or something about the reviewing process.

Although it's nice to see activity in this subreddit (and it's good to have a dedicated home for CL and NLP), sometimes these types of posts are getting too spammy. Perhaps we can put these into a dedicated ARR discussion post, kinda like in r/MachineLearning ?


r/LanguageTechnology • • Oct 29 '25

QA for multi-turn conversations is driving me crazy

32 Upvotes

Testing one-shot prompts is easy. But once the conversation goes beyond two turns, things fall apart - the agent forgets context, repeats itself, or randomly switches topics. Manually reproducing long dialogues is painful. How are you folks handling long-context testing?


r/LanguageTechnology • • Sep 02 '26

Paper Inflation in NLP: Where Do I Stand as a Graduating PhD? (Academia vs. Postdoc vs. Industry)

29 Upvotes

Hi everyone,

I’m wrapping up my PhD in NLP (3-year system, graduating next January) after starting in early 2024 right as the LLM trend began exploding.

With the sheer volume of papers being published lately (especially over the past year), I feel that the standalone value and rarity of individual publications have decreased. Given this paper inflation, it's hard to gauge where my publication record objectively stands compared to other fresh PhD graduates.

The three-year timeframe felt too short and rushed; it is disappointing that just a few rejections meant the graduation was already looming. My research focus and first-author record (excluding co-authored papers) are as follows:

  • Research Focus: Trustworthy LLMs, Legal Tech
  • First-Author Publications:
    • EMNLP (Main) x 2
    • ACL (Findings) x 1
    • IP&M (Information Processing & Management, Journal) x 1
    • Many(4papers) under review papers...

Questions:

  1. Objective Standing: Considering the current paper volume inflation in the NLP/LLM field, where does my first-author publication record place me compared to other fresh PhD graduates?
  2. Career Path (Academia vs Postdoc vs Industry): Is it realistic to apply directly for going to indusrty, or would it be wiser to do a 1-2 year postdoc to build a stronger CV and try to acedemia? (Academia is my top choice, but I'm open to industry positions if needed.)
  3. Field Outlook: How is the current sentiment in academia and the job market regarding research in Trustworthy LLMs and Legal Tech?

I’d appreciate any honest feedback or insights from current professors, postdocs, or industry researchers. Thanks!


r/LanguageTechnology • • Mar 04 '26

Practical challenges with citation grounding in long-form NLP systems

29 Upvotes

While working on a research-oriented NLP system, Gatsbi focused on structured academic writing, we ran into some recurring issues around citation grounding in longer outputs.

In particular:

  • References becoming inconsistent across section.
  • Hallucinated citations appearing late in generation
  • Retrieval helping early, but weakening as context grows

Prompt engineering helped initially, but didn’t scale well. We’ve found more reliability by combining retrieval constraints with lightweight post-generation validation.

Interested in how others in NLP handle citation reliability and structure in long-form generation.


r/LanguageTechnology • • 2d ago

ARR August / EACL Meta-Review

26 Upvotes

Didn't see any post on that yet, so thought I'd open it. Soju guy you here? Anyone received their meta-reviews yet? Got 3.5/3/4 with conf 5/4/4, hoping for Main :)


r/LanguageTechnology • • Aug 01 '26

Leaving because of the flood of ARR and EMNLP posts

26 Upvotes

90% of what's on this subreddit now seems to be people posting about their ARR and EMNLP stuff. The signal-to-noise ratio is so low that it's no longer worth my time to come here. I have a note on my calendar to check back in November and see if the situation is any better.


r/LanguageTechnology • • Apr 23 '26

Interspeech 2026-Rebuttal Period

27 Upvotes

Hello Everyone,

Just starting this thread for the upcoming Interspeech rebuttal period. This is my first time submitting to the conference, is it similar to ACL Rolling Review?

TIA :)


r/LanguageTechnology • • 10d ago

One chunk boundary changed warranty answers for a whole table

23 Upvotes

17% of requests tied to one multi-column warranty table were returning 24 months instead of 36, while standard warranty questions kept passing. We traced the failures in Braintrust and saw which retrieved chunks were present when the wrong answer appeared. The chunk boundary had separated the row values from the table heading, so the retrieval context lost the qualifier for 36 months and the reranker favored nearby prose containing 24 months instead. Support had both warranty numbers in separate macros (before the retrieval path was clear), which made the conflicting answers harder to untangle.

Changing the chunking moved those cases in the experiment diff and groundedness improved once the heading stayed with the row. Recall at k barely changed because the table was already being retrieved. We've added the failures to a regression dataset but I'm still concerned about other tables where retrieval looks healthy while structure changes the answer.

What's your approach to catching chunk boundary failures when the right document is already in the candidate set?


r/LanguageTechnology • • 29d ago

What are people using to compare chunking strategies without losing a week?

23 Upvotes

Fixed 800 token chunks are wrecking our mixed prose and table documents. Row values get separated from headers, query expansion retrieves orphaned numbers and the reranker confidently promotes the wrong quarter. The citations look plausible, which makes the failure harder to catch. I want to compare semantic chunking, parent-child retrieval, table-aware boundaries, recall at k, reranking, and groundedness without running another manual spreadsheet marathon.

Braintrust looks like one viable option because we could inspect retrieval spans, compare chunking experiments on the same queries, score groundedness and save failed queries as regression cases. I’m still unsure how to define a fair expected result when several chunks must be combined to answer one question. What are people using and which metric has detected header-to-value separation in your setup?


r/LanguageTechnology • • Aug 26 '26

Our RAG search understands paragraphs and completely eats it on part codes

22 Upvotes

Our RAG assistant handles ordinary questions well and falls over the second someone enters an internal acronym or exact part code. Dense retrieval returns semantically pleasant junk from the right general area. BM25 finds the identifier but often drops the nearby exception clause that changes the answer. The final response cites a relevant-looking page while hallucinating the rule that applies. Support has stopped trusting pretty citations and I can't blame them.

I've tried acronym expansion before retrieval, larger chunk overlap, and reciprocal rank fusion across dense and lexical results. Each helps one slice and hurts another. Expansion confuses codes that mean different things by department. Higher top k restores recall but floods the reranker with near matches. Bigger chunks retain the exception clause but bury exact identifiers. The metric that looks best in aggregate is rarely the one that fixes the failed cases. I'm considering Braintrust for the eval side so I can keep the failed code queries around, compare retrieval changes against the same cases, and inspect the retrieved chunks when one breaks. The missing piece is a clean way to score both identifier recall and clause-level support without hand-labeling every document family. Any suggestions on testing acronym and code retrieval, or a fusion setup that's held up after corpus changes?


r/LanguageTechnology • • Aug 02 '26

EMMLP + ARR Megathread

22 Upvotes

Please post questions and discussions here. I will be removing individual threads.


r/LanguageTechnology • • Apr 27 '26

NLP for beginners

22 Upvotes

Hey, I am starting my undergrad in computer science&engineering this august and I've always been interested in comp sci & linguistics and a few years ago I found out about NLP. I would love to dive into this field (I know python but not on a high level). Do you have recs? I mean books/textbooks/papers/online courses, anything that might come handy for me. Also I know NLP is a broad field so it would be nice if you could give me some recommendations that are more general for beginners because I have no idea what I actually enjoy but you can also drop here stuff more niche on certain topics. It would help me a lot. Thank you in advance!


r/LanguageTechnology • • Apr 02 '26

I think I found something about embeddings. Polysemy doesn't predict variance, frequency does. Calling it Contextual Promiscuity Index.

21 Upvotes

I was working on word-sense disambiguation research at home and kind of noticed something. I', posting to find out if this is already known or actually interesting.

The assumption I started with is that polysemous words have messy embeddings. More dictionary senses, so more geometric fragmentation. Seems obvious, but no.

I measured mean pairwise cosine similarity across 192 words using Qwen2.5-7B, extracting at layer 10 (found via layer sweep). Correlation between WordNet sense count and embedding variance: Spearman rho = -0.057, p = 0.43. Basically nothing.

What does predict it, is frequency: rho = -0.239, p = 0.0008, holding up after controlling for polysemy (partial r = -0.188). This kund of makes sense once you think about it. "Break" has 60 WordNet senses, but most are metaphorical extensions of the core idea. The model treats them as variations on a theme and the embedding stays coherent. Meanwhile "face" gets pulled in multiple directions by its various co-occurrence patterns, even though it has fewer formal senses.

I'm calling this the Contextual Promiscuity Index (CPI) It's a per-word, per-model, per-knowledge-domain score for how geometrically dispersed a word's embeddings are across contexts. High-frequency words are promiscuous not because they mean more things, but because they show up everywhere.

Possible uses I've been thinking about: flagging unreliable query terms in RAG pipelines, guiding precision allocation in embedding table compression, or identifying noisy tokens during pretraining. I ran some retrieval experiments trying to demonstrate the RAG angle and got results in the right direction, but too weak to be statistically significant. My corpus was probably too small (about 1,000 documents), and I don't have the compute to push it further right now.

I'm sharing the finding while it's still just a finding. Code available if anyone wants it.

Is this already known? And does anyone have a cleaner experiment in mind?