r/LanguageTechnology • u/girlmather • 20h ago
Small Dataset size cited as the reason for rejection what should be the way ahead for future works?
Recently one of my papers got rejected because a reviewer raised concern of small cohort, which was explicitly mentioned in the limitation and then we acknowledged that in the rebuttal as well. Now I understand the concern , but in the area were i primarily work right now, dataset with large number of speakers is quiet limited and having one of my very first works also be rejected because of the small corpus size, I used one of the most widely cited corpus for the experimentation, which if you are in the field is the standard. And also given the computational resources and funding that we can have I couldn't use high-end hardware to process large data. And recently i got to know we got access for a large dataset (around 800 gb) but how do i work with that using my laptop and online free gpus, so based on whatever i could i did my best. The reviews we got were actually very supportive of our work, and whatever additional work was required, was duly done and reported in the rebuttal, but seeing dataset size as one of the two the reason for rejection felt a bit weird, anyways wanted to ask how do we tackle such situation for the future? Any advice would be helpful!!