- ๐ A Specialized Arabic Language Model for Islamic Heritage
- Research Preview
- ๐ Vision
- ๐งญ Project Philosophy
- โจ What Makes This Model Unique?
- ๐๏ธ Long-Term Architecture
- ๐ Scalability
- ๐ Preliminary Results
- ๐ฌ Research Philosophy
- ๐ฏ Design Goal
- Why Another Arabic LLM?
- Current Capabilities
- ๐ฆ Current Release
- ๐ Training Corpus
- ๐ง Knowledge Extraction Pipeline
- ๐ Why Not Train on the General Internet?
- ๐บ๏ธ Long-Term Research Strategy
- ๐ฎ Future Dataset Evolution
- ๐งช Future Research Directions
- ๐ฏ Project Goal
- ๐๏ธ Future Roadmap
- ๐ Towards a Comprehensive Digital Library
- ๐ Usage Example The code must be used because it is compatible with the training method.
- Current Limitations
- This project explores several research directions:
- Why Memorization?
๐ Github | ๐ค Hugging Face | ๐ Cookbooks
๐ฅ๏ธ Demo
๐ A Specialized Arabic Language Model for Islamic Heritage
*3arabLM is an ongoing research project dedicated to building a large-scale Arabic language model that preserves, memorizes, and reconstructs the classical Islamic scholarly heritage directly from its original sources.
Unlike general-purpose LLMs, this project is not designed to imitate conversations. Its primary objective is the faithful reconstruction of scholarly knowledge while preserving the language, methodology, and diversity of the classical Islamic tradition..*
Research Preview
*This model is an early research and should not be used as a source for religious rulings or legal verdicts. Its purpose is to explore large-scale memorization of the classical Islamic corpus and faithful reconstruction of scholarly texts.
5๏ธโฃ Domain Specialization
The model is optimized for:
- ๐ Fiqh (Islamic Jurisprudence)
- ๐ Tafsir (Exegesis)
- The current model represents less than 2% of the planned continual pretraining schedule.
- The full project is expected to expand over multiple stages covering nearly the complete Al-Maktaba Al-Shamela ecosystem.
๐ Vision
This project aims to build a large-scale Arabic language model primarily trained on Al-Maktaba Al-Shamela and other authoritative Islamic heritage sources.
The objective is to make scholarly knowledge itself part of the model's parameters, rather than relying on external retrieval or internet-scale mixed corpora.
The philosophy behind the project is simple:
- ๐ Learn directly from the original books.
- โ๏ธ Preserve the language of classical scholars.
- ๐ฏ Preserve each author's methodology.
- ๐ Preserve scholarly terminology.
- โ๏ธ Preserve differences between schools and commentators.
The model is therefore gradually moving toward a paradigm of:
Retrieval from Weights
rather than:
Generative Summarization
Instead of producing heavily paraphrased modern summaries, the model attempts to recall scholarly knowledge from its internal parameters using language close to the original sources.
๐งญ Project Philosophy
This is not a general conversational model.
Its objective is not creative writing.
Its objective is not producing modern short-form answers.
Instead, the project focuses on:
Faithful reconstruction of classical Islamic knowledge while preserving the original scholarly language and style.
Accordingly, this model can be viewed as a:
- ๐ Knowledge Recall Model
- ๐ง Memorization-Oriented Language Model
rather than simply a:
- ๐ฌ Generative Chat Model
โจ What Makes This Model Unique?
1๏ธโฃ Training on Original Sources
Training relies primarily on classical Islamic literature from trusted scholarly sources rather than general web data.
This allows responses to be grounded in the scholarly material the model actually learned instead of relying on broad internet correlations.
2๏ธโฃ Preserving Classical Scholarly Style
The model is trained to maintain traditional Arabic scholarly language.
Typical responses naturally include expressions such as:
- Imam Al-Nawawi said...
- Ibn Kathir said...
- Al-Qurtubi said...
- The majority of scholars held...
- The preferred opinion is...
Instead of modern conversational wording.
3๏ธโฃ Distinguishing Between Scholarly Methodologies
One of the central goals of this project is preventing the collapse of Islamic scholarship into a single homogeneous representation.
The model learns that:
- Tafsir Ibn Kathir is distinct from Tafsir Al-Qurtubi.
- Tafsir Al-Tabari follows its own methodology.
- Every scholar has a unique style of evidence, reasoning, and preference.
Likewise, it learns that:
- A single verse may have multiple valid interpretations.
- A single jurisprudential issue may contain several legitimate opinions.
- Scholarly disagreement is part of the knowledgeโnot noise to be removed.
4๏ธโฃ Faithful Quotations
Unlike many instruction-tuned language models that aggressively summarize or paraphrase,
this model attemptsโwhenever possibleโto reproduce scholarly explanations in wording close to their original form while preserving technical terminology and classical expressions.
๐๏ธ Long-Term Architecture
The long-term vision is to maintain a stable foundational model and build specialized expert models on top of it.
Examples include:
| Expert | Domain |
|---|---|
| ๐ Fiqh Expert | Islamic Jurisprudence |
| ๐ Hadith Expert | Prophetic Traditions |
| ๐ Tafsir Expert | Quranic Exegesis |
| ๐ Arabic Language Expert | Grammar & Morphology |
| ๐ Aqeedah Expert | Islamic Creed |
| ๐ Islamic History Expert | History & Biography |
This architecture allows future domains to be added without retraining the entire base model.
๐ Scalability
Future research directions include:
- Copying the final transformer layers with small random perturbations.
- Increasing model capacity through additional layers.
- Preserving previous knowledge while creating space for specialized learning.
The expected advantages are:
- โ Better knowledge preservation.
- โ Increased memorization capacity.
- โ Improved continual training stability.
๐ Preliminary Results
Initial experiments indicate that the model is capable of:
- ๐ง Distinguishing between Islamic sciences internally.
- โ๏ธ Preserving author-specific writing styles.
- ๐ Reproducing scholarly passages with wording close to the originals.
- ๐ฏ Separating Fiqh, Tafsir, Hadith, and Arabic Language representations during training.
Internal Evaluation
| Metric | Value |
|---|---|
| Linear Probe Accuracy | 96.26% |
This suggests that the model forms distinct internal representations for different branches of Islamic knowledge rather than treating them as an undifferentiated corpus.
๐ฌ Research Philosophy
The entire project may be summarized by the following statement:
Large Language Models as Compressed Digital Libraries
Rather than viewing language models merely as text generators,
this project explores the possibility of treating them as compressed scholarly libraries capable of faithfully reconstructing classical knowledge from their parameters.
Related research directions include:
- Large Language Models as Compressed Digital Libraries: Full-Parameter Memorization of Classical Arabic Knowledge
- Beyond Retrieval-Augmented Generation: End-to-End Memorization of a Large Arabic Classical Corpus
๐ฏ Design Goal
Unlike conventional instruction-tuned LLMs that tend to summarize or paraphrase,
our objective is faithful knowledge reconstruction from the classical Islamic corpus.
The model is intentionally optimized to reproduce scholarly explanations in their original style rather than generating concise summaries.
Why Another Arabic LLM?
Modern large language models are remarkably capable across many domains, yet they are trained on extremely broad internet-scale corpora.
For highly specialized scholarly fields such as classical Islamic sciences, this often introduces several limitations:
Classical references represent only a tiny fraction of the training data. Distinct scholarly methodologies become blended together. Responses are frequently paraphrased instead of preserving original scholarly wording. Source attribution becomes difficult because the model no longer distinguishes where knowledge originated.
3arabLM follows a different philosophy.
Instead of maximizing domain diversity, it maximizes fidelity to a carefully curated corpus of classical Islamic l
Current Capabilities
The current model already demonstrates:
Preservation of classical Arabic scholarly writing. Book-aware text generation. Author-aware stylistic reconstruction. Recognition of differences between jurisprudential schools. Differentiation between major Tafsir methodologies. Long-form scholarly quotations. Reconstruction of passages closely matching the original literature.
๐ฆ Current Release
This release represents the first public milestone of a long-term research project.
The model should be viewed as:
An early research preview, not the final objective.
Although the model already demonstrates promising capabilities in preserving classical Arabic scholarly language and recalling information from Islamic heritage, it has been trained on less than 2% of the planned training curriculum.
The long-term vision extends far beyond this initial release.
๐ Training Corpus
The foundation of this project is Al-Maktaba Al-Shamela, one of the largest digital collections of classical Islamic literature.
Official Library: goldenshamela
The original Shamela library contains approximately 35,000 classical Arabic books covering nearly every branch of Islamic scholarship, including:
- ๐ Tafsir
- ๐ Hadith
- โ๏ธ Fiqh
- ๐ Aqeedah
- ๐ Usul al-Fiqh
- ๐ Arabic Grammar
- ๐ History
- ๐ค Biography
- ๐ Literature
- ๐ Lexicons
- And many other disciplines.
Rather than using these books directly as plain text, the entire corpus is reconstructed through a dedicated extraction pipeline designed specifically for language-model training.
๐ง Knowledge Extraction Pipeline
The corpus is transformed into a structured representation through multiple stages.
Shamela
โ
โโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโ
โ โ
Books Indices
โ โ
โผ โผ
Text Parser Keyword Parser
โ โ
โโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโ
โผ
Unified Page Builder
โ
โโโโโโโโโโโโโโฌโโโโโโโโโโโโโฌโโโโโโโโโโโโโ
โผ โผ โผ
Metadata TOC Keywords
โ
โผ
pages.parquet
โ
โโโโโโโโโโโโโโฌโโโโโโโโโโโโโฌโโโโโโโโโโโโโ
โผ โผ โผ
CPT SFT RAG
โ
Book โ Section โ Paragraph โ Sentence
This pipeline preserves the hierarchical structure of the original books instead of flattening them into raw text.
Consequently, the model learns not only the words, but also the relationships between books, chapters, sections, and scholarly discussions.
๐ Why Not Train on the General Internet?
Most modern Large Language Models are trained on extremely broad internet-scale corpora.
While this makes them highly versatile, it also introduces significant challenges for specialized scholarly domains.
For example:
- Classical Islamic sources become a very small fraction of the overall training data.
- The model often loses awareness of the original source of a statement.
- Distinct scholarly methodologies become mixed together.
- Modern paraphrases may dominate over the original classical wording.
This project follows a different philosophy.
Instead of maximizing domain diversity, it maximizes scholarly consistency within a carefully curated corpus.
๐บ๏ธ Long-Term Research Strategy
The long-term architecture is based on a single foundational model followed by multiple specialized expert models.
Instead of independently training several large models from scratch, the project follows a staged approach.
Stage 1 โ Foundation Model
The base model is continually pretrained on a balanced mixture of Islamic sciences to develop a unified understanding of classical Arabic.
This foundation becomes the permanent reference model for the project.
Stage 2 โ Specialized Experts
Each expert inherits the linguistic and scholarly knowledge of the foundation model while specializing in its own discipline.
This architecture allows new experts to be added without retraining the entire system.
๐ฎ Future Dataset Evolution
Future releases will gradually enrich the training data with structured metadata instead of relying solely on plain text.
Planned metadata includes:
| Field | Description |
|---|---|
| Book ID | Book identifier |
| Author ID | Author identifier |
| Century | Hijri century |
| Madhhab | School of thought |
| Book Hierarchy | Book โ Chapter โ Section |
| Semantic Tags | Domain-specific tags |
The objective is to strengthen the model's internal understanding of scholarly context rather than treating every passage as isolated text.
๐งช Future Research Directions
The project is currently exploring several research ideas, including:
- ๐ง Metadata-aware continual pretraining.
- ๐ Knowledge-oriented transformer scaling.
- ๐ง Layer expansion through duplicated upper transformer blocks with small stochastic initialization.
- ๐พ Higher-capacity memorization without catastrophic forgetting.
- ๐ฏ Expert-specialized adapters built upon a single foundational model.
These ideas aim to transform the model from a general language model into a faithful digital representation of the classical Islamic scholarly tradition.
๐ฏ Project Goal
The ultimate objective is not to replace scholars.
Nor is it simply to build another conversational assistant.
The objective is to develop a language model capable of:
Preserving, recalling, and reconstructing the classical Islamic heritage with the highest possible degree of fidelity while maintaining the language, methodology, and scholarly diversity found in the original sources.
๐๏ธ Future Roadmap
This release represents only the beginning of the project.
The model has been trained on a small fraction of the planned corpus, and future releases will substantially expand both the breadth and depth of the model's knowledge.
The long-term objective is to build the most comprehensive Arabic language model dedicated to the Islamic scholarly heritage.
๐ Planned Corpus Expansion
| Discipline | Planned Additions |
|---|---|
| Hadith Sciences | สฟIlal, Suสพฤlฤt, Commentaries, Mustalah, Takhrij, Zawฤสพid, Mutลซn, Manuscripts, Works of al-Albani |
| Tafsir | Classical & Contemporary Works, Broader Methodologies |
| Fiqh | Four Schools, Comparative Fiqh, Fatwas, Legal Maxims, Contemporary Research |
| Biography & History | Sirah, Shamฤ'il, Chronicles, Tabaqฤt, Genealogy, Civilization |
| Creed | Theology, Sects, Refutations, Debates |
| Arabic Language | Grammar, Morphology, Lexicography, Dictionaries, Rhetoric, Literature, Poetry |
| Spiritual Literature | Adab, Ruqaq, Dhikr, Manners, Purification |
| General | Research Papers, Discussions, Indexes, Bibliographies, Catalogs |
๐ Towards a Comprehensive Digital Library
The ultimate objective is to transform the model into a compressed digital representation of the Islamic scholarly tradition.
Rather than specializing in only one branch of knowledge, the long-term model will gradually encompass nearly every major discipline preserved within Al-Maktaba Al-Shamela.
Each training stage will improve:
- ๐ Breadth of scholarly coverage.
- ๐ Depth of domain knowledge.
- โ๏ธ Preservation of author-specific writing styles.
- ๐ฏ Recognition of methodological differences between scholars.
- ๐ Faithful reconstruction of classical texts.
As the corpus grows, the model will increasingly function as an internal scholarly library whose knowledge is stored directly within its parameters rather than reconstructed from fragmented internet sources.
This roadmap represents the long-term vision of the project: a continually evolving foundation model upon which increasingly specialized expert models can be built while preserving a unified understanding of the classical Islamic heritage.
๐ Usage Example The code must be used because it is compatible with the training method.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
BASE_MODEL = "sherif1313/3arabLM-4B-Fiqh-v1"
print("โณ ุชุญู
ูู ุงููู
ูุฐุฌ...")
tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL,
dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
model.eval()
def generate_text(prompt, book="ุงูุชูุณูุฑ", author="ุงุจู ูุซูุฑ", category="ุงูุชูุงุณูุฑ"):
# ุงูุงูุชุฒุงู
ุจุชูุฒูุน ุงูุชุฏุฑูุจ
full_prompt = f"book {book}\nauthor {author}\ncategory {category}\n{prompt}"
print(f"\n๐ ุงูุณุคุงู/ุงููุต: {prompt}")
print("-" * 50)
inputs = tokenizer(full_prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
do_sample=True,
temperature=0.15, # ุฅุนุฏุงุฏุงุชู ุงูู
ู
ุชุงุฒุฉ
top_p=0.85,
top_k=40,
repetition_penalty=1.08,
no_repeat_ngram_size=5,
max_new_tokens=300, # ุชู
ุฑูุน ุงูุญุฏ ููููุงู ูุชูุงุณุจ ุงูุฃุณุฆูุฉ ุงูููููุฉ
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.eos_token_id,
)
answer = tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
print(f"๐ค ุฅุฌุงุจุฉ ุงููู
ูุฐุฌ:\n{answer}")
print("=" * 80)
# ูุงุฆู
ุฉ ุจุงูุฃุณุฆูุฉ ุงูู
ุชููุนุฉ ูุงุฎุชุจุงุฑ ุงููู
ูุฐุฌ
test_cases = test_cases = [
# ==================== 1. ุงูุชูุณูุฑ ====================
{
"prompt": "ููููู ุชุนุงูู: {ุฅููููุงูู ููุนูุจูุฏู ููุฅููููุงูู ููุณูุชูุนูููู} ุฃู:",
"book": "ุชูุณูุฑ ุงุจู ูุซูุฑ",
"author": "ุงุจู ูุซูุฑ",
"category": "ุงูุชูุงุณูุฑ"
},
{
"prompt": "ูุณุฑ ูููู ุชุนุงูู: {ุงูููุฐูููู ุขู
ููููุง ููููู
ู ููููุจูุณููุง ุฅููู
ูุงููููู
ู ุจูุธูููู
ู} [ุงูุฃูุนุงู
: 82]",
"book": "ุชูุณูุฑ ุงูุทุจุฑู",
"author": "ุงูุทุจุฑู",
"category": "ุงูุชูุงุณูุฑ"
},
{
"prompt": "ู
ุง ูู ุณุจุจ ูุฒูู ุณูุฑุฉ ุงููุงุชุญุฉุ",
"book": "ุชูุณูุฑ ุงููุฑุทุจู",
"author": "ุงููุฑุทุจู",
"category": "ุงูุชูุงุณูุฑ"
},
{
"prompt": "ู
ุงุฐุง ูุงู ุงูู
ูุณุฑูู ูู ูููู ุชุนุงูู: {ููุฅููู ุชูุนูุฏูููุง ููุนูู
ูุฉู ุงูููููู ูุง ุชูุญูุตููููุง} [ุงููุญู: 18]ุ",
"book": "ุชูุณูุฑ ุงูุฑุงุฒู",
"author": "ุงูุฑุงุฒู",
"category": "ุงูุชูุงุณูุฑ"
},
{
"prompt": "ู
ุง ู
ุนูู ุงูุงุณุชุนุงุฐุฉ ูู ูููู ุชุนุงูู: {ููุงุณูุชูุนูุฐู ุจูุงูููููู ุฅูููููู ูููู ุงูุณููู
ููุนู ุงููุจูุตููุฑู} [ุบุงูุฑ: 56]ุ",
"book": "ุชูุณูุฑ ุงูุจุบูู",
"author": "ุงูุจุบูู",
"category": "ุงูุชูุงุณูุฑ"
},
# ==================== 2. ุงูููู (ุงูุนุจุงุฏุงุช) ====================
{
"prompt": "ู
ุง ุญูู
ุงูู
ุณุญ ุนูู ุงูุฎูููุ",
"book": "ุงูุฃู
",
"author": "ุงูุดุงูุนู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ูู ุดุฑูุท ุตุญุฉ ุงูุตูุงุฉุ",
"book": "ุงูู
ุฌู
ูุน",
"author": "ุงููููู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ุญูู
ุชุฑู ุงูุตูุงุฉุ",
"book": "ุงูู
ุบูู",
"author": "ุงุจู ูุฏุงู
ุฉ",
"category": "ููู ุนุงู
"
},
{
"prompt": "ูู
ุนุฏุฏ ุฑูุนุงุช ุตูุงุฉ ุงูุฌู
ุนุฉุ",
"book": "ุจุฏุงุฆุน ุงูุตูุงุฆุน",
"author": "ุงููุงุณุงูู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ูู ู
ุจุทูุงุช ุงููุถูุกุ",
"book": "ุงููุฏุงูุฉ",
"author": "ุงูู
ุฑุบููุงูู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ุญูู
ุตูุงุฉ ุงูู
ุณุงูุฑุ",
"book": "ุงูู
ุฌู
ูุน",
"author": "ุงููููู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ูู ูุฌูุฒ ุงูุฌู
ุน ุจูู ุงูุตูุงุชูู ูู ุงูุณูุฑุ",
"book": "ุงูู
ุบูู",
"author": "ุงุจู ูุฏุงู
ุฉ",
"category": "ููู ุนุงู
"
},
# ==================== 3. ุงูููู (ุงูู
ุนุงู
ูุงุช) ====================
{
"prompt": "ู
ุง ุญูู
ุจูุน ุงูุฑุทุจ ุจุงูุชู
ุฑุ",
"book": "ุงูุฃู
",
"author": "ุงูุดุงูุนู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ุญูู
ุจูุน ุงูุฐูุจ ุจุงูุฐูุจุ",
"book": "ุงูู
ุจุณูุท",
"author": "ุงูุณุฑุฎุณู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ูู ุดุฑูุท ุตุญุฉ ุนูุฏ ุงูุจูุนุ",
"book": "ุงูู
ุฌู
ูุน",
"author": "ุงููููู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ุญูู
ุงูุฑุจุง ูู ุงูุฅุณูุงู
ุ",
"book": "ุงูู
ุบูู",
"author": "ุงุจู ูุฏุงู
ุฉ",
"category": "ููู ุนุงู
"
},
{
"prompt": "ูู ูุฌูุฒ ุจูุน ุงูุฏูู ุจุงูุฏููุ",
"book": "ุจุฏุงุฆุน ุงูุตูุงุฆุน",
"author": "ุงููุงุณุงูู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ุญูู
ุงูุบุด ูู ุงูุจูุนุ",
"book": "ุงูุฃู
",
"author": "ุงูุดุงูุนู",
"category": "ููู ุนุงู
"
},
# ==================== 4. ุงูููู (ุงูุฒูุงุฌ ูุงูุทูุงู) ====================
{
"prompt": "ู
ุง ูู ุดุฑูุท ุตุญุฉ ุนูุฏ ุงูููุงุญุ",
"book": "ุงูู
ุฌู
ูุน",
"author": "ุงููููู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ุญูู
ุงูุทูุงู ูู ุงูุฅุณูุงู
ุ",
"book": "ุงูู
ุบูู",
"author": "ุงุจู ูุฏุงู
ุฉ",
"category": "ููู ุนุงู
"
},
{
"prompt": "ูู
ุนุฏุฏ ุงูุทููุงุช ุงูุชู ูู
ูููุง ุงูุฒูุฌุ",
"book": "ุงููุฏุงูุฉ",
"author": "ุงูู
ุฑุบููุงูู",
"category": "ููู ุนุงู
"
},
# ==================== 5. ุงูููู (ุงูุฌูุงุฆุฒ) ====================
{
"prompt": "ู
ุง ูู ุฃุญูุงู
ุบุณู ุงูู
ูุชุ",
"book": "ุงูู
ุฌู
ูุน",
"author": "ุงููููู",
"category": "ููู ุนุงู
"
},
{
"prompt": "ู
ุง ุญูู
ุงูุตูุงุฉ ุนูู ุงูู
ูุชุ",
"book": "ุงูู
ุบูู",
"author": "ุงุจู ูุฏุงู
ุฉ",
"category": "ููู ุนุงู
"
},
# ==================== 6. ุงูููู (ุงูุฌูุงุฏ) ====================
{
"prompt": "ู
ุง ุญูู
ุงูุฌูุงุฏ ูู ุงูุฅุณูุงู
ุ",
"book": "ุงูู
ุจุณูุท",
"author": "ุงูุณุฑุฎุณู",
"category": "ููู ุนุงู
"
}
]
print("\n๐ ุจุฏุก ุงุฎุชุจุงุฑ ุงููู
ูุฐุฌ ุนูู ุฃุณุฆูุฉ ู
ุชููุนุฉ...\n")
for test in test_cases:
generate_text(test["prompt"], test["book"], test["author"], test["category"])
Current Limitations
This release should be considered an early research .
Current limitations include:
Less than 2% of the planned continual pretraining has been completed. Some domains remain significantly underrepresented. The model may occasionally mix neighboring passages. Citation grounding is still under active development. Long-context reconstruction will improve in future releases.
This project explores several research directions:
Full-parameter memorization of a classical corpus. Metadata-aware continual pretraining. Hierarchical book representation. Weight-based knowledge retrieval. Capacity scaling through layer expansion. Domain-specialized expert models.
Why Memorization?
Most recent research focuses on Retrieval-Augmented Generation (RAG), where knowledge remains outside the model and is retrieved at inference time.
This project investigates the opposite direction:
Can a large language model become a compressed scholarly library whose knowledge is stored directly in its parameters?
Our hypothesis is that sufficiently large continual pretraining on a carefully curated corpus allows faithful reconstruction of classical scholarly knowledge without relying on external retrieval for every query.
๐ค Final Note
This project seeks to preserve the scientific heritage of the Islamic library inside a language modelโfaithfully, transparently, and with respect for the diversity of classical scholarship.
If you have any questions, would like to contribute, or have ideas for future research directions, please feel free to reach out.
The journey is just beginning.
๐ Citation
If you use this model in your research, please cite:
@misc{shamela-arabic-foundation, author = {Sherif1313}, title = {3arabLM-4B-Fiqh-v1: A Specialized Arabic Language Model for Islamic Heritage}, year = {2026}, publisher = {Hugging Face}, url = { https://huggingface.co/sherif1313/3arabLM-4B-Fiqh-v1 } }
- Downloads last month
- 13