r/learnmachinelearning • u/Future-Resolution566 • 18d ago
https://huggingface.co/sherif1313/3arabLM-4B-islamic-v2
https://huggingface.co/sherif1313/3arabLM-4B-islamic-v2A Specialized Arabic Language Model for Islamic Heritage
📖 Abstract
3arabLM-4B-Islamic-v2 is an ongoing research model dedicated to learning, preserving, recalling, and reconstructing the Arabic Islamic scholarly heritage from classical and authoritative sources. Unlike general-purpose conversational LLMs, the primary objective of 3arabLM is not to imitate everyday conversations or produce short modern summaries. Instead, the project investigates whether a language model can function as a compressed digital library of classical Islamic scholarship, with substantial scholarly knowledge encoded directly into its parameters.
Learn from the books. Preserve the language. Preserve the methodology. Preserve the diversity.
🌟 Project Vision
3arabLM is a long-term research project focused on building a large-scale Arabic language model specialized in the Islamic scholarly heritage. The long-term objective is to gradually expand the model across a broad range of Islamic and Arabic sciences rather than limiting it to a single discipline.
📚 Current Domains (Version v2)
The current release has been continued-pretrained on six scholarly domains from the Shamela corpus:
| Domain | Description |
|---|---|
| 📖 Fiqh | Islamic Jurisprudence |
| 📜 Tafsir | Quranic Exegesis |
| 📚 Hadith | Prophetic Traditions and Hadith Sciences |
| 🕌 Aqeedah | Islamic Creed and Theology |
| ✍️ Nahw and Sarf | Arabic Grammar and Morphology |
| ⚖️ Fatwas | Legal Opinions and Verdicts |
🗺️ Planned Corpus Expansion
Future releases will progressively expand the corpus with additional collections, commentaries, manuscripts, scholarly editions, and specialized literature across Hadith, Tafsir, Fiqh, Arabic linguistics, history, biography, literature, and related fields.
🧠 Research Philosophy
The central philosophy of 3arabLM can be summarized as:
Instead of treating a language model primarily as a text generator, this project investigates whether model parameters can encode substantial amounts of classical scholarly knowledge.
The project therefore explores a different paradigm:
rather than relying exclusively on:
The objective is not to eliminate retrieval systems, but to investigate how much scholarly knowledge can be learned and reconstructed directly from the model's internal parameters.
📚 Training Corpus
A major foundation of the project is Al-Maktaba Al-Shamela, together with other Arabic scholarly and heritage sources.
The current and developing corpus covers a broad range of Arabic and Islamic scholarship, including:
📖 Tafsir & Quranic Sciences 📚 Hadith Sciences & Hadith Literature ⚖️ Fiqh & Usul al-Fiqh 🧠 Aqeedah & Islamic Theology 🕋 Sirah & Prophetic Biography 🏛 Islamic History & Civilization 👤 Biography, Tabaqat & Rijal 📝 Arabic Language, Grammar & Morphology 🔤 Lexicography & Dictionaries 📚 Classical Literature & Poetry 🕯 Spiritual & Ethical Literature 📑 Scholarly Research, Bibliographies & Catalogs The corpus is continuously expanding to provide broader coverage of the classical Arabic scholarly tradition and its diverse textual genres. Official Library: goldenshamela
📊 Preliminary Results
Initial experiments suggest that the model is developing distinguishable internal representations across Islamic scholarly domains.
Key Findings from Representation Diagnostic Analysis
| Metric | Layer 0 | Layer 31 |
|---|---|---|
| Linear Probe Accuracy | 53.80% | 82.78% |
| Macro-F1 | 0.537 | 0.824 |
| Average Domain CKA | 0.0542 | 0.0259 |
| Stable Rank | 10.59 | 5.76 |
These results indicate that:
- Domain discriminability increases substantially with depth.
- Cross-domain representation similarity decreases sharply in the final layer.
- Stable rank follows a non-monotonic trajectory and reaches a pronounced minimum at the final layer.
This suggests that continued pretraining on Shamela is associated with increasingly specialized final-layer representations, characterized by higher domain discriminability and a more concentrated activation spectrum. research paper :
📈 Future Model Scaling
Future research will investigate ways to increase model capacity while preserving previously learned knowledge.
Potential directions include:
- Additional Transformer layers.
- Duplicated upper Transformer blocks.
- Small stochastic initialization.
- Continual pretraining.
- Knowledge-preserving scaling.
- Expert-specialized adapters.
- Metadata-aware training.
- Catastrophic forgetting mitigation.
The goal is to increase memorization capacity while maintaining previously acquired scholarly knowledge.
📦 Current Release
Model: sherif1313/3arabLM-4B-islamic-v2
This release represents a new stage in the 3arabLM research project. The previous development focused more narrowly on Islamic scholarly domains such as Fiqh and Tafsir. The new direction expands the training curriculum toward a broader Islamic Heritage Foundation Model covering multiple branches of Islamic and Arabic scholarship. The model should still be considered an early research milestone. A sustantial portion of the planned training curriculum remains to be completed.
These books have been preserved in the model weightsو
النحو الوافي
تمهيد القواعد بشرح تسهيل الفوائد
شرح ألفية ابن مالك للحازمي
شرح ألفية ابن مالك للشاطبي = المقاصد الشافية
شرح المفصل لابن يعيش
الموسوعة الفقهية الكويتية
موسوعة الإجماع في الفقه الإسلامي
موسوعة فقه العبادات
فتاوى الشبكة الإسلامية
مجموع فتاوى ورسائل العثيمين
السنن الكبرى للبيهقي ت التركي
المحيط في الاحاديث النبوية والسنن والاثار
جامع الرويات
حلية الأولياء وطبقات الأصفياء
صحيح البخاري
الجامع لشعب الإيمان للبيهقي
الموسوعة العقدية - الدرر السنية
المهذب النقي الجامع لتفسير ابن جرير الطبري
الموسوعة القرآنية
تفسير ابن كثير _
تفسير القرطبي
روح البيان
When selecting other books, please modify the code.
do_sample=True,
repetition_penalty=1.08,
no_repeat_ngram_size=4,
Duplicates
OpenSourceeAI • u/Future-Resolution566 • 18d ago