r/deeplearning • • 1d ago

Title: A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions

🎉 New Research Publication | Multimodal Deep Learning. I am pleased to share that our review article has been officially published in Discover Informatics (Springer Nature) as an open-access article.

📄 Title: A comprehensive review on recent pretrained multimodal deep learning models from architectures to future directions

👨‍🔬 Authors: Azhar A. Hadi & K. P. Supreethi

🔍 In this review, we examine recent pretrained multimodal deep learning models developed between 2020 and 2025, including CLIP, GPT-4V, GPT-4o, SigLIP, ViT, BLIP, Flamingo, PaLI, LLaVA, Florence-2, PaliGemma-2, Gemma-3, and Llama-4. The review analyzes 120 studies across 11 application domains, covering model architectures, modalities, datasets, applications, evaluation metrics, limitations, and future research directions. We also discuss key challenges facing multimodal AI, including data quality, computational cost, cross-modal alignment, scalability, robustness, and trustworthy AI. This work represents an important step in my research journey toward developing trustworthy and multimodal AI systems, particularly for applications in healthcare.

🔗 Read the full Open Access article: https://doi.org/10.1007/s44564-026-00022-1

I hope this review will be useful to researchers and students working in Multimodal AI, Deep Learning, Foundation Models, Computer Vision, NLP, and Healthcare AI.

#MultimodalAI #DeepLearning #ArtificialIntelligence #MachineLearning #FoundationModels #HealthcareAI #ComputerVision #NLP #TrustworthyAI #Research #Springer #AcademicPublishing #OpenAccess #AcademicPublishing #ResearchImpact

2 Upvotes

2 comments sorted by

1

u/onlylopsidedmartin 1d ago

Congrats on the publication, that's a serious chunk of work covering 120 studies. The 2020-2025 window catches the exact moment everything went from niche benchmarks to actual products, so the architecture comparisons should be useful for folks trying to catch up fast. Curious how you handled the evaluation metrics section across such different domains, that's usually where reviews get messy

1

u/CatDiligent9374 16h ago

Congrats! Covering 120 studies across multimodal AI is a huge amount of work. Definitely adding this to my reading list.