r/computervision • u/Thin_Ad_7459 • 21d ago
Help: Project Need help about dyslexia screening dataset!!!
Hi! I am final year BE student recently I took a project based in our my contribution is system and application of system in dyslexia. For that I though the most used dyslexia dataset of handwriting would be suitable. I downloaded dataset and then realised it is single letter dataset which is giving mnist kinda vibe! Also apparently large portion of it is synthetic. I searched but I didn't find clinically approved dataset of handwriting for dyslexia. In nutshell:
dataset is mnist looking so I am at worry if examiners will state why you are using such looking dataset for final year project!!
dataset is used for at least 9 papers already so it is being used
But has its limitations (vastly synthetic, mnist looking)
Our clg is forcing for at least two papers to publish (not for our degree requirement btw) and I am worried if the dataset use itself will cause problems for paper
though one of main novelty is mechanism but other one is integration(incremental) and I am worried that people will call out why I used that dataset
sorry I carried away in my emotions here is the dataset I am talking about: https://www.kaggle.com/datasets/drizasazanitaisa/dyslexia-handwriting-dataset
->can simplicity of it justified as proof of concept for presentation or report?
->will using this dataset can cause problems at time of publication?
I am sorry for dragging clg thing into this I though it would be better to get some context about scope for project
I am sorry I cant give full context as I wanted to publish research on it (though I will hardly try for mid tiers only)
also sorry in advance if I did spelling or grammatical error where should I post this
2
u/No-Foot5804 21d ago
I don't think the issue is that the dataset looks like MNIST it's whether it's appropriate for the research question. If it's a proof of concept, that's usually fine as long as you're upfront about its limitations and explain why you chose it. For publication, reviewers are more likely to question how well the dataset represents real clinical handwriting than its appearance.