๐Ÿ’œ Github   |   ๐Ÿค— Hugging Face   |   ๐Ÿ“š Cookbooks  
๐Ÿ–ฅ๏ธ Demo  

๐Ÿ•Œ A Specialized Arabic Language Model for Islamic Heritage

*3arabLM is an ongoing research project dedicated to building a large-scale Arabic language model that preserves, memorizes, and reconstructs the classical Islamic scholarly heritage directly from its original sources.

Unlike general-purpose LLMs, this project is not designed to imitate conversations. Its primary objective is the faithful reconstruction of scholarly knowledge while preserving the language, methodology, and diversity of the classical Islamic tradition..*

Research Preview

*This model is an early research and should not be used as a source for religious rulings or legal verdicts. Its purpose is to explore large-scale memorization of the classical Islamic corpus and faithful reconstruction of scholarly texts.


5๏ธโƒฃ Domain Specialization

The model is optimized for:

  • ๐Ÿ“š Fiqh (Islamic Jurisprudence)
  • ๐Ÿ“– Tafsir (Exegesis)
  • The current model represents less than 2% of the planned continual pretraining schedule.
  • The full project is expected to expand over multiple stages covering nearly the complete Al-Maktaba Al-Shamela ecosystem.

๐ŸŒŸ Vision

This project aims to build a large-scale Arabic language model primarily trained on Al-Maktaba Al-Shamela and other authoritative Islamic heritage sources.

The objective is to make scholarly knowledge itself part of the model's parameters, rather than relying on external retrieval or internet-scale mixed corpora.

The philosophy behind the project is simple:

  • ๐Ÿ“– Learn directly from the original books.
  • โœ๏ธ Preserve the language of classical scholars.
  • ๐ŸŽฏ Preserve each author's methodology.
  • ๐Ÿ“š Preserve scholarly terminology.
  • โš–๏ธ Preserve differences between schools and commentators.

The model is therefore gradually moving toward a paradigm of:

Retrieval from Weights

rather than:

Generative Summarization

Instead of producing heavily paraphrased modern summaries, the model attempts to recall scholarly knowledge from its internal parameters using language close to the original sources.


๐Ÿงญ Project Philosophy

This is not a general conversational model.

Its objective is not creative writing.

Its objective is not producing modern short-form answers.

Instead, the project focuses on:

Faithful reconstruction of classical Islamic knowledge while preserving the original scholarly language and style.

Accordingly, this model can be viewed as a:

  • ๐Ÿ“– Knowledge Recall Model
  • ๐Ÿง  Memorization-Oriented Language Model

rather than simply a:

  • ๐Ÿ’ฌ Generative Chat Model

โœจ What Makes This Model Unique?

1๏ธโƒฃ Training on Original Sources

Training relies primarily on classical Islamic literature from trusted scholarly sources rather than general web data.

This allows responses to be grounded in the scholarly material the model actually learned instead of relying on broad internet correlations.

2๏ธโƒฃ Preserving Classical Scholarly Style

The model is trained to maintain traditional Arabic scholarly language.

Typical responses naturally include expressions such as:

  • Imam Al-Nawawi said...
  • Ibn Kathir said...
  • Al-Qurtubi said...
  • The majority of scholars held...
  • The preferred opinion is...

Instead of modern conversational wording.

3๏ธโƒฃ Distinguishing Between Scholarly Methodologies

One of the central goals of this project is preventing the collapse of Islamic scholarship into a single homogeneous representation.

The model learns that:

  • Tafsir Ibn Kathir is distinct from Tafsir Al-Qurtubi.
  • Tafsir Al-Tabari follows its own methodology.
  • Every scholar has a unique style of evidence, reasoning, and preference.

Likewise, it learns that:

  • A single verse may have multiple valid interpretations.
  • A single jurisprudential issue may contain several legitimate opinions.
  • Scholarly disagreement is part of the knowledgeโ€”not noise to be removed.

4๏ธโƒฃ Faithful Quotations

Unlike many instruction-tuned language models that aggressively summarize or paraphrase,

this model attemptsโ€”whenever possibleโ€”to reproduce scholarly explanations in wording close to their original form while preserving technical terminology and classical expressions.


๐Ÿ—๏ธ Long-Term Architecture

The long-term vision is to maintain a stable foundational model and build specialized expert models on top of it.

Examples include:

Expert Domain
๐Ÿ“š Fiqh Expert Islamic Jurisprudence
๐Ÿ“– Hadith Expert Prophetic Traditions
๐Ÿ“œ Tafsir Expert Quranic Exegesis
๐Ÿ“ Arabic Language Expert Grammar & Morphology
๐Ÿ•Œ Aqeedah Expert Islamic Creed
๐Ÿ› Islamic History Expert History & Biography

This architecture allows future domains to be added without retraining the entire base model.


๐Ÿ“ˆ Scalability

Future research directions include:

  • Copying the final transformer layers with small random perturbations.
  • Increasing model capacity through additional layers.
  • Preserving previous knowledge while creating space for specialized learning.

The expected advantages are:

  • โœ… Better knowledge preservation.
  • โœ… Increased memorization capacity.
  • โœ… Improved continual training stability.

๐Ÿ“Š Preliminary Results

Initial experiments indicate that the model is capable of:

  • ๐Ÿง  Distinguishing between Islamic sciences internally.
  • โœ๏ธ Preserving author-specific writing styles.
  • ๐Ÿ“– Reproducing scholarly passages with wording close to the originals.
  • ๐ŸŽฏ Separating Fiqh, Tafsir, Hadith, and Arabic Language representations during training.

Internal Evaluation

Metric Value
Linear Probe Accuracy 96.26%

This suggests that the model forms distinct internal representations for different branches of Islamic knowledge rather than treating them as an undifferentiated corpus.


๐Ÿ”ฌ Research Philosophy

The entire project may be summarized by the following statement:

Large Language Models as Compressed Digital Libraries

Rather than viewing language models merely as text generators,

this project explores the possibility of treating them as compressed scholarly libraries capable of faithfully reconstructing classical knowledge from their parameters.

Related research directions include:

  • Large Language Models as Compressed Digital Libraries: Full-Parameter Memorization of Classical Arabic Knowledge
  • Beyond Retrieval-Augmented Generation: End-to-End Memorization of a Large Arabic Classical Corpus

๐ŸŽฏ Design Goal

Unlike conventional instruction-tuned LLMs that tend to summarize or paraphrase,

our objective is faithful knowledge reconstruction from the classical Islamic corpus.

The model is intentionally optimized to reproduce scholarly explanations in their original style rather than generating concise summaries.


Why Another Arabic LLM?

Modern large language models are remarkably capable across many domains, yet they are trained on extremely broad internet-scale corpora.

For highly specialized scholarly fields such as classical Islamic sciences, this often introduces several limitations:

Classical references represent only a tiny fraction of the training data. Distinct scholarly methodologies become blended together. Responses are frequently paraphrased instead of preserving original scholarly wording. Source attribution becomes difficult because the model no longer distinguishes where knowledge originated.

3arabLM follows a different philosophy.

Instead of maximizing domain diversity, it maximizes fidelity to a carefully curated corpus of classical Islamic l

Current Capabilities

The current model already demonstrates:

Preservation of classical Arabic scholarly writing. Book-aware text generation. Author-aware stylistic reconstruction. Recognition of differences between jurisprudential schools. Differentiation between major Tafsir methodologies. Long-form scholarly quotations. Reconstruction of passages closely matching the original literature.


๐Ÿ“ฆ Current Release

This release represents the first public milestone of a long-term research project.

The model should be viewed as:

An early research preview, not the final objective.

Although the model already demonstrates promising capabilities in preserving classical Arabic scholarly language and recalling information from Islamic heritage, it has been trained on less than 2% of the planned training curriculum.

The long-term vision extends far beyond this initial release.


๐Ÿ“š Training Corpus

The foundation of this project is Al-Maktaba Al-Shamela, one of the largest digital collections of classical Islamic literature.

Official Library: goldenshamela

The original Shamela library contains approximately 35,000 classical Arabic books covering nearly every branch of Islamic scholarship, including:

  • ๐Ÿ“– Tafsir
  • ๐Ÿ“œ Hadith
  • โš–๏ธ Fiqh
  • ๐Ÿ•Œ Aqeedah
  • ๐Ÿ“ Usul al-Fiqh
  • ๐Ÿ“š Arabic Grammar
  • ๐Ÿ› History
  • ๐Ÿ‘ค Biography
  • ๐Ÿ“ Literature
  • ๐Ÿ“– Lexicons
  • And many other disciplines.

Rather than using these books directly as plain text, the entire corpus is reconstructed through a dedicated extraction pipeline designed specifically for language-model training.


๐Ÿ”ง Knowledge Extraction Pipeline

The corpus is transformed into a structured representation through multiple stages.

                    Shamela
                       โ”‚
        โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
        โ”‚                             โ”‚
      Books                       Indices
        โ”‚                             โ”‚
        โ–ผ                             โ–ผ
   Text Parser                  Keyword Parser
        โ”‚                             โ”‚
        โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                       โ–ผ
            Unified Page Builder
                       โ”‚
      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
      โ–ผ            โ–ผ            โ–ผ
   Metadata        TOC       Keywords
                       โ”‚
                       โ–ผ
                 pages.parquet
                       โ”‚
      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
      โ–ผ            โ–ผ            โ–ผ
      CPT          SFT          RAG
                       โ”‚
        Book โ†’ Section โ†’ Paragraph โ†’ Sentence

This pipeline preserves the hierarchical structure of the original books instead of flattening them into raw text.

Consequently, the model learns not only the words, but also the relationships between books, chapters, sections, and scholarly discussions.


๐ŸŒ Why Not Train on the General Internet?

Most modern Large Language Models are trained on extremely broad internet-scale corpora.

While this makes them highly versatile, it also introduces significant challenges for specialized scholarly domains.

For example:

  • Classical Islamic sources become a very small fraction of the overall training data.
  • The model often loses awareness of the original source of a statement.
  • Distinct scholarly methodologies become mixed together.
  • Modern paraphrases may dominate over the original classical wording.

This project follows a different philosophy.

Instead of maximizing domain diversity, it maximizes scholarly consistency within a carefully curated corpus.


๐Ÿ—บ๏ธ Long-Term Research Strategy

The long-term architecture is based on a single foundational model followed by multiple specialized expert models.

Instead of independently training several large models from scratch, the project follows a staged approach.

Stage 1 โ€” Foundation Model

The base model is continually pretrained on a balanced mixture of Islamic sciences to develop a unified understanding of classical Arabic.

This foundation becomes the permanent reference model for the project.

Stage 2 โ€” Specialized Experts

Each expert inherits the linguistic and scholarly knowledge of the foundation model while specializing in its own discipline.

This architecture allows new experts to be added without retraining the entire system.


๐Ÿ”ฎ Future Dataset Evolution

Future releases will gradually enrich the training data with structured metadata instead of relying solely on plain text.

Planned metadata includes:

Field Description
Book ID Book identifier
Author ID Author identifier
Century Hijri century
Madhhab School of thought
Book Hierarchy Book โ†’ Chapter โ†’ Section
Semantic Tags Domain-specific tags

The objective is to strengthen the model's internal understanding of scholarly context rather than treating every passage as isolated text.


๐Ÿงช Future Research Directions

The project is currently exploring several research ideas, including:

  • ๐Ÿง  Metadata-aware continual pretraining.
  • ๐Ÿ“ˆ Knowledge-oriented transformer scaling.
  • ๐Ÿ”ง Layer expansion through duplicated upper transformer blocks with small stochastic initialization.
  • ๐Ÿ’พ Higher-capacity memorization without catastrophic forgetting.
  • ๐ŸŽฏ Expert-specialized adapters built upon a single foundational model.

These ideas aim to transform the model from a general language model into a faithful digital representation of the classical Islamic scholarly tradition.


๐ŸŽฏ Project Goal

The ultimate objective is not to replace scholars.

Nor is it simply to build another conversational assistant.

The objective is to develop a language model capable of:

Preserving, recalling, and reconstructing the classical Islamic heritage with the highest possible degree of fidelity while maintaining the language, methodology, and scholarly diversity found in the original sources.


๐Ÿ—“๏ธ Future Roadmap

This release represents only the beginning of the project.

The model has been trained on a small fraction of the planned corpus, and future releases will substantially expand both the breadth and depth of the model's knowledge.

The long-term objective is to build the most comprehensive Arabic language model dedicated to the Islamic scholarly heritage.


๐Ÿ“– Planned Corpus Expansion

Discipline Planned Additions
Hadith Sciences สฟIlal, Suสพฤlฤt, Commentaries, Mustalah, Takhrij, Zawฤสพid, Mutลซn, Manuscripts, Works of al-Albani
Tafsir Classical & Contemporary Works, Broader Methodologies
Fiqh Four Schools, Comparative Fiqh, Fatwas, Legal Maxims, Contemporary Research
Biography & History Sirah, Shamฤ'il, Chronicles, Tabaqฤt, Genealogy, Civilization
Creed Theology, Sects, Refutations, Debates
Arabic Language Grammar, Morphology, Lexicography, Dictionaries, Rhetoric, Literature, Poetry
Spiritual Literature Adab, Ruqaq, Dhikr, Manners, Purification
General Research Papers, Discussions, Indexes, Bibliographies, Catalogs

๐Ÿ“š Towards a Comprehensive Digital Library

The ultimate objective is to transform the model into a compressed digital representation of the Islamic scholarly tradition.

Rather than specializing in only one branch of knowledge, the long-term model will gradually encompass nearly every major discipline preserved within Al-Maktaba Al-Shamela.

Each training stage will improve:

  • ๐Ÿ“ˆ Breadth of scholarly coverage.
  • ๐Ÿ“Š Depth of domain knowledge.
  • โœ๏ธ Preservation of author-specific writing styles.
  • ๐ŸŽฏ Recognition of methodological differences between scholars.
  • ๐Ÿ“– Faithful reconstruction of classical texts.

As the corpus grows, the model will increasingly function as an internal scholarly library whose knowledge is stored directly within its parameters rather than reconstructed from fragmented internet sources.

This roadmap represents the long-term vision of the project: a continually evolving foundation model upon which increasingly specialized expert models can be built while preserving a unified understanding of the classical Islamic heritage.


๐Ÿš€ Usage Example The code must be used because it is compatible with the training method.

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

BASE_MODEL = "sherif1313/3arabLM-4B-Fiqh-v1"

print("โณ ุชุญู…ูŠู„ ุงู„ู†ู…ูˆุฐุฌ...")
tokenizer = AutoTokenizer.from_pretrained(BASE_MODEL, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    BASE_MODEL, 
    dtype=torch.bfloat16, 
    device_map="auto", 
    trust_remote_code=True
)
model.eval()

def generate_text(prompt, book="ุงู„ุชูุณูŠุฑ", author="ุงุจู† ูƒุซูŠุฑ", category="ุงู„ุชูุงุณูŠุฑ"):
    # ุงู„ุงู„ุชุฒุงู… ุจุชูˆุฒูŠุน ุงู„ุชุฏุฑูŠุจ
    full_prompt = f"book {book}\nauthor {author}\ncategory {category}\n{prompt}"
    
    print(f"\n๐Ÿ“ ุงู„ุณุคุงู„/ุงู„ู†ุต: {prompt}")
    print("-" * 50)
    
    inputs = tokenizer(full_prompt, return_tensors="pt").to(model.device)
    
    with torch.no_grad():
        outputs = model.generate(
            **inputs,
            do_sample=True,
            temperature=0.15,          # ุฅุนุฏุงุฏุงุชูƒ ุงู„ู…ู…ุชุงุฒุฉ
            top_p=0.85,
            top_k=40,
            repetition_penalty=1.08,
            no_repeat_ngram_size=5,
            max_new_tokens=300,        # ุชู… ุฑูุน ุงู„ุญุฏ ู‚ู„ูŠู„ุงู‹ ู„ุชู†ุงุณุจ ุงู„ุฃุณุฆู„ุฉ ุงู„ูู‚ู‡ูŠุฉ
            eos_token_id=tokenizer.eos_token_id,
            pad_token_id=tokenizer.eos_token_id,
        )
    
    answer = tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
    print(f"๐Ÿค– ุฅุฌุงุจุฉ ุงู„ู†ู…ูˆุฐุฌ:\n{answer}")
    print("=" * 80)

# ู‚ุงุฆู…ุฉ ุจุงู„ุฃุณุฆู„ุฉ ุงู„ู…ุชู†ูˆุนุฉ ู„ุงุฎุชุจุงุฑ ุงู„ู†ู…ูˆุฐุฌ
test_cases = test_cases = [
    # ==================== 1. ุงู„ุชูุณูŠุฑ ====================
    {
        "prompt": "ูˆู‚ูˆู„ู‡ ุชุนุงู„ู‰: {ุฅููŠูŽู‘ุงูƒูŽ ู†ูŽุนู’ุจูุฏู ูˆูŽุฅููŠูŽู‘ุงูƒูŽ ู†ูŽุณู’ุชูŽุนููŠู†ู} ุฃูŠ:",
        "book": "ุชูุณูŠุฑ ุงุจู† ูƒุซูŠุฑ",
        "author": "ุงุจู† ูƒุซูŠุฑ",
        "category": "ุงู„ุชูุงุณูŠุฑ"
    },
    {
        "prompt": "ูุณุฑ ู‚ูˆู„ู‡ ุชุนุงู„ู‰: {ุงู„ูŽู‘ุฐููŠู†ูŽ ุขู…ูŽู†ููˆุง ูˆูŽู„ูŽู…ู’ ูŠูŽู„ู’ุจูุณููˆุง ุฅููŠู…ูŽุงู†ูŽู‡ูู…ู’ ุจูุธูู„ู’ู…ู} [ุงู„ุฃู†ุนุงู…: 82]",
        "book": "ุชูุณูŠุฑ ุงู„ุทุจุฑูŠ",
        "author": "ุงู„ุทุจุฑูŠ",
        "category": "ุงู„ุชูุงุณูŠุฑ"
    },
    {
        "prompt": "ู…ุง ู‡ูˆ ุณุจุจ ู†ุฒูˆู„ ุณูˆุฑุฉ ุงู„ูุงุชุญุฉุŸ",
        "book": "ุชูุณูŠุฑ ุงู„ู‚ุฑุทุจูŠ",
        "author": "ุงู„ู‚ุฑุทุจูŠ",
        "category": "ุงู„ุชูุงุณูŠุฑ"
    },
    {
        "prompt": "ู…ุงุฐุง ู‚ุงู„ ุงู„ู…ูุณุฑูˆู† ููŠ ู‚ูˆู„ู‡ ุชุนุงู„ู‰: {ูˆูŽุฅูู†ู’ ุชูŽุนูุฏูู‘ูˆุง ู†ูุนู’ู…ูŽุฉูŽ ุงู„ู„ูŽู‘ู‡ู ู„ุง ุชูุญู’ุตููˆู‡ูŽุง} [ุงู„ู†ุญู„: 18]ุŸ",
        "book": "ุชูุณูŠุฑ ุงู„ุฑุงุฒูŠ",
        "author": "ุงู„ุฑุงุฒูŠ",
        "category": "ุงู„ุชูุงุณูŠุฑ"
    },
    {
        "prompt": "ู…ุง ู…ุนู†ู‰ ุงู„ุงุณุชุนุงุฐุฉ ููŠ ู‚ูˆู„ู‡ ุชุนุงู„ู‰: {ููŽุงุณู’ุชูŽุนูุฐู’ ุจูุงู„ู„ูŽู‘ู‡ู ุฅูู†ูŽู‘ู‡ู ู‡ููˆูŽ ุงู„ุณูŽู‘ู…ููŠุนู ุงู„ู’ุจูŽุตููŠุฑู} [ุบุงูุฑ: 56]ุŸ",
        "book": "ุชูุณูŠุฑ ุงู„ุจุบูˆูŠ",
        "author": "ุงู„ุจุบูˆูŠ",
        "category": "ุงู„ุชูุงุณูŠุฑ"
    },

    # ==================== 2. ุงู„ูู‚ู‡ (ุงู„ุนุจุงุฏุงุช) ====================
    {
        "prompt": "ู…ุง ุญูƒู… ุงู„ู…ุณุญ ุนู„ู‰ ุงู„ุฎููŠู†ุŸ",
        "book": "ุงู„ุฃู…",
        "author": "ุงู„ุดุงูุนูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ู‡ูŠ ุดุฑูˆุท ุตุญุฉ ุงู„ุตู„ุงุฉุŸ",
        "book": "ุงู„ู…ุฌู…ูˆุน",
        "author": "ุงู„ู†ูˆูˆูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ุญูƒู… ุชุฑูƒ ุงู„ุตู„ุงุฉุŸ",
        "book": "ุงู„ู…ุบู†ูŠ",
        "author": "ุงุจู† ู‚ุฏุงู…ุฉ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ูƒู… ุนุฏุฏ ุฑูƒุนุงุช ุตู„ุงุฉ ุงู„ุฌู…ุนุฉุŸ",
        "book": "ุจุฏุงุฆุน ุงู„ุตู†ุงุฆุน",
        "author": "ุงู„ูƒุงุณุงู†ูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ู‡ูŠ ู…ุจุทู„ุงุช ุงู„ูˆุถูˆุกุŸ",
        "book": "ุงู„ู‡ุฏุงูŠุฉ",
        "author": "ุงู„ู…ุฑุบูŠู†ุงู†ูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ุญูƒู… ุตู„ุงุฉ ุงู„ู…ุณุงูุฑุŸ",
        "book": "ุงู„ู…ุฌู…ูˆุน",
        "author": "ุงู„ู†ูˆูˆูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู‡ู„ ูŠุฌูˆุฒ ุงู„ุฌู…ุน ุจูŠู† ุงู„ุตู„ุงุชูŠู† ููŠ ุงู„ุณูุฑุŸ",
        "book": "ุงู„ู…ุบู†ูŠ",
        "author": "ุงุจู† ู‚ุฏุงู…ุฉ",
        "category": "ูู‚ู‡ ุนุงู…"
    },

    # ==================== 3. ุงู„ูู‚ู‡ (ุงู„ู…ุนุงู…ู„ุงุช) ====================
    {
        "prompt": "ู…ุง ุญูƒู… ุจูŠุน ุงู„ุฑุทุจ ุจุงู„ุชู…ุฑุŸ",
        "book": "ุงู„ุฃู…",
        "author": "ุงู„ุดุงูุนูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ุญูƒู… ุจูŠุน ุงู„ุฐู‡ุจ ุจุงู„ุฐู‡ุจุŸ",
        "book": "ุงู„ู…ุจุณูˆุท",
        "author": "ุงู„ุณุฑุฎุณูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ู‡ูŠ ุดุฑูˆุท ุตุญุฉ ุนู‚ุฏ ุงู„ุจูŠุนุŸ",
        "book": "ุงู„ู…ุฌู…ูˆุน",
        "author": "ุงู„ู†ูˆูˆูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ุญูƒู… ุงู„ุฑุจุง ููŠ ุงู„ุฅุณู„ุงู…ุŸ",
        "book": "ุงู„ู…ุบู†ูŠ",
        "author": "ุงุจู† ู‚ุฏุงู…ุฉ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู‡ู„ ูŠุฌูˆุฒ ุจูŠุน ุงู„ุฏูŠู† ุจุงู„ุฏูŠู†ุŸ",
        "book": "ุจุฏุงุฆุน ุงู„ุตู†ุงุฆุน",
        "author": "ุงู„ูƒุงุณุงู†ูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ุญูƒู… ุงู„ุบุด ููŠ ุงู„ุจูŠุนุŸ",
        "book": "ุงู„ุฃู…",
        "author": "ุงู„ุดุงูุนูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },

    # ==================== 4. ุงู„ูู‚ู‡ (ุงู„ุฒูˆุงุฌ ูˆุงู„ุทู„ุงู‚) ====================
    {
        "prompt": "ู…ุง ู‡ูŠ ุดุฑูˆุท ุตุญุฉ ุนู‚ุฏ ุงู„ู†ูƒุงุญุŸ",
        "book": "ุงู„ู…ุฌู…ูˆุน",
        "author": "ุงู„ู†ูˆูˆูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ุญูƒู… ุงู„ุทู„ุงู‚ ููŠ ุงู„ุฅุณู„ุงู…ุŸ",
        "book": "ุงู„ู…ุบู†ูŠ",
        "author": "ุงุจู† ู‚ุฏุงู…ุฉ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ูƒู… ุนุฏุฏ ุงู„ุทู„ู‚ุงุช ุงู„ุชูŠ ูŠู…ู„ูƒู‡ุง ุงู„ุฒูˆุฌุŸ",
        "book": "ุงู„ู‡ุฏุงูŠุฉ",
        "author": "ุงู„ู…ุฑุบูŠู†ุงู†ูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },

    # ==================== 5. ุงู„ูู‚ู‡ (ุงู„ุฌู†ุงุฆุฒ) ====================
    {
        "prompt": "ู…ุง ู‡ูŠ ุฃุญูƒุงู… ุบุณู„ ุงู„ู…ูŠุชุŸ",
        "book": "ุงู„ู…ุฌู…ูˆุน",
        "author": "ุงู„ู†ูˆูˆูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    },
    {
        "prompt": "ู…ุง ุญูƒู… ุงู„ุตู„ุงุฉ ุนู„ู‰ ุงู„ู…ูŠุชุŸ",
        "book": "ุงู„ู…ุบู†ูŠ",
        "author": "ุงุจู† ู‚ุฏุงู…ุฉ",
        "category": "ูู‚ู‡ ุนุงู…"
    },

    # ==================== 6. ุงู„ูู‚ู‡ (ุงู„ุฌู‡ุงุฏ) ====================
    {
        "prompt": "ู…ุง ุญูƒู… ุงู„ุฌู‡ุงุฏ ููŠ ุงู„ุฅุณู„ุงู…ุŸ",
        "book": "ุงู„ู…ุจุณูˆุท",
        "author": "ุงู„ุณุฑุฎุณูŠ",
        "category": "ูู‚ู‡ ุนุงู…"
    }
]

print("\n๐Ÿš€ ุจุฏุก ุงุฎุชุจุงุฑ ุงู„ู†ู…ูˆุฐุฌ ุนู„ู‰ ุฃุณุฆู„ุฉ ู…ุชู†ูˆุนุฉ...\n")

for test in test_cases:
    generate_text(test["prompt"], test["book"], test["author"], test["category"])

Current Limitations

This release should be considered an early research .

Current limitations include:

Less than 2% of the planned continual pretraining has been completed. Some domains remain significantly underrepresented. The model may occasionally mix neighboring passages. Citation grounding is still under active development. Long-context reconstruction will improve in future releases.

This project explores several research directions:

Full-parameter memorization of a classical corpus. Metadata-aware continual pretraining. Hierarchical book representation. Weight-based knowledge retrieval. Capacity scaling through layer expansion. Domain-specialized expert models.

Why Memorization?

Most recent research focuses on Retrieval-Augmented Generation (RAG), where knowledge remains outside the model and is retrieved at inference time.

This project investigates the opposite direction:

Can a large language model become a compressed scholarly library whose knowledge is stored directly in its parameters?

Our hypothesis is that sufficiently large continual pretraining on a carefully curated corpus allows faithful reconstruction of classical scholarly knowledge without relying on external retrieval for every query.

๐Ÿค Final Note

This project seeks to preserve the scientific heritage of the Islamic library inside a language modelโ€”faithfully, transparently, and with respect for the diversity of classical scholarship.

If you have any questions, would like to contribute, or have ideas for future research directions, please feel free to reach out.

The journey is just beginning.

๐Ÿ“Ž Citation

If you use this model in your research, please cite:

@misc{shamela-arabic-foundation, author = {Sherif1313}, title = {3arabLM-4B-Fiqh-v1: A Specialized Arabic Language Model for Islamic Heritage}, year = {2026}, publisher = {Hugging Face}, url = { https://huggingface.co/sherif1313/3arabLM-4B-Fiqh-v1 } }

Downloads last month
13
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ 1 Ask for provider support

Model tree for sherif1313/3arabLM-4B-Fiqh-v1

Finetunes
1 model

Collection including sherif1313/3arabLM-4B-Fiqh-v1