CoolFace
Datasetpublic

nizarun/FineWeb-Edu-Arabic-24M

English العربية FineWeb-Edu Arabic 24M An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics. Highlight Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.

sourceHugging Faceodc-byupdated 25d agoView on Hugging Face
0likes464downloads
Dataset Card

<a id="english"></a> English

العربية

FineWeb-Edu Arabic 24M

An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the `sample-350BT` configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics.

Highlight

Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local materials and design to climate, comfort, privacy, and beauty. doc_id: <urn:uuid:f00fd3ac-a6e2-4c3b-8287-3152496dabf4>

The 71-row featured split is a viewer-only showcase duplicated from train; all corpus statistics refer only to train.

Methodology

1. Build an in-domain translation set

We first used the stronger DeepSeek-v4-flash teacher model to translate authentic educational, scientific, and reasoning material, including FineWeb-Edu. After quality filtering, this produced the 106,114-pair nizarun/English_Arabic_Translation_Pairs dataset. This was not a generic translation benchmark: its source distribution was selected to resemble the educational and technical documents the production models would later translate.

2. Fine-tune the Gemma 4 translators

All three production translators belong to the Gemma 4 family and were fine-tuned on the same English–Arabic dataset above.

Public productFine-tuning methodDocuments
Model B1Full fine-tuning16,435,776
Model B2Full fine-tuning1,648,715
Model ALoRA6,709,934

The public labels distinguish translation products while avoiding unnecessary deployment-specific metadata.

3. Select the FineWeb-Edu source documents

The main target was documents with a FineWeb-Edu source score of 3.5 or higher. We also intentionally retained 1,313,640 documents scoring 3.0–<3.5 in Model A. A hard cutoff can change the mix of subjects and writing styles, so this lower-score band broadens coverage and helps reduce subject-selection bias introduced by relying on one automated threshold.

The score column describes the original English document, not the Arabic translation and not translation quality.

Original FineWeb-Edu scoreDocumentsShare
3.0–<3.51,313,6405.30%
3.5–<4.019,741,19179.62%
4.0–<4.53,561,68714.36%
≥4.5177,9070.72%
Total24,794,425100.00%

Overall, 23,480,785 documents (94.70%) have a source score of at least 3.5.

4. Filter during translation and after reassembly

Documents were translated in chunks and reassembled under their original IDs. Quality control was applied both while generating chunks and again after complete documents were rebuilt.

CheckWhat it was designed to catch
Empty output and generation-cap hitsMissing translations, cut-off output, and the looping failures that often consume the maximum token allowance
Chunk and document length ratiosSuspiciously short, dropped, duplicated, or runaway translations
Arabic-script shareOutput that did not contain enough Arabic while still allowing legitimate equations, citations, URLs, and technical terms
Source-aware 4-token repetitionNew repetitive loops in Arabic without penalizing lists or repetition already present in the English source
Adjacent-word repetitionDegenerate sequences repeating the same word
Compression/diversity signalLow-information or mechanically repetitive output
Unsupported-script charactersUnexpected third scripts such as Cyrillic, Greek, or kana
Document-ID checksDuplicate documents within a product or shared documents across products

In this release, foreign_chars does not mean every non-Arabic character or every foreign-language word. Latin characters were permitted because correct translations often need URLs, DOI references, email addresses, acronyms, chemical names, code, and citations. The signal counts characters from unsupported third scripts that the translation pipeline was not expected to produce.

The product-specific thresholds and measured signals remain in the public rows so users can build stricter or looser subsets. Exact ID checks found no repeated doc_id within a product and no shared IDs across the three products.

Scale and length

ProductDocumentsSource tokensArabic tokensArabic/source token ratio
Model B116,435,77616,653,225,45623,113,631,9251.3879
Model B21,648,7151,627,190,6562,252,746,9171.3844
Model A6,709,9346,809,303,4649,497,777,2091.3948
Total24,794,42525,089,719,57634,864,156,0511.3896

These are tokenizer-derived counts, not whitespace-delimited word counts. We report the exact reproducible measure available in the corpus rather than estimate “words” from tokens. An exact Arabic-to-English word ratio would require another full pass over the upstream English text, which is not redistributed in this Arabic-only release.

Mean Arabic-script ratio is 0.9769, median Arabic-script ratio is 0.9892, and the median Arabic/source token-length ratio is 1.3976.

Usage

python
from datasets import load_dataset

dataset = load_dataset(
    "nizarun/FineWeb-Edu-Arabic-24M",
    split="train",
    streaming=True,
)

first = next(iter(dataset))
print(first["ar_text"][:500])

Filter by translation product or any documented signal:

python
model_a = dataset.filter(lambda row: row["translation_model"] == "Model A")

Schema

ColumnDescription
doc_idOriginal FineWeb-Edu id; use it to locate the corresponding English record in sample-350BT
ar_textComplete reassembled Arabic document
scoreAutomated educational-quality score of the original English document; not a translation score
n_chunksNumber of translated chunks rejoined into the document
en_tokens, ar_tokensTokenizer-derived source and Arabic token counts
arabic_ratioArabic-script character ratio
len_ratioArabic/source token-length ratio
diac_ratioArabic diacritic ratio
foreign_charsCount of characters from unexpected third scripts; ordinary Latin technical content is allowed
rep4_max, rep4_en_maxArabic and corresponding source four-token repetition signals
chunk_len_ratio_min, chunk_len_ratio_maxMinimum and maximum chunk-level length ratios
chunk_ar_tokens_max, chunk_ar_ratio_minMaximum Arabic tokens and minimum Arabic ratio across chunks
chunk_zr_minMinimum chunk compression/diversity signal
adj_repeat_maxMaximum adjacent-word repetition signal
corpus_versionSanitized public corpus version
shardOriginal numeric source shard
translation_modelPublic label: Model A, Model B1, or Model B2

Translation-quality audit

A deterministic, blinded, full-document audit reviewed 400 documents: 120 score-and-length-matched triplets across the three products, plus 40 Model A documents from the 3.0–<3.5 band. Every selected document was reviewed from first to last character, 20% received repeat annotation, and all initially cited errors received a conservative full-context adjudication.

In the matched sample, at least one core linguistic issue was identified in 68.3% of Model B1, 65.0% of Model B2, and 65.8% of Model A documents. This demanding category counts even an isolated agreement, grammar, terminology, literal-language, consistency, or intelligibility issue; it should not be read as the percentage of unusable documents. Parser-assisted estimates placed gender-agreement accuracy above 99.5% for all three products, but these are not hand-counted gold measurements.

Observed errors were generally sparse and inconsistent across documents and products. At massive pretraining scale they may average out, but this is an informed expectation, not a guarantee: translationese, source-corpus bias, and recurring model tendencies can still be learned.

Recommended training use

Use this corpus for large-scale pretraining or mid-training, mixed with substantial native-Arabic data. We do not recommend making it a dominant source in the final annealing stage. Near the end of training, taper or exclude it in favor of carefully curated native Arabic, human-reviewed material, and high-confidence task-specific examples, because late-stage data can disproportionately shape final style and language quality.

Limitations

  • Every Arabic document is machine translated.
  • Chunk boundaries can weaken continuation, pronoun reference, or terminology consistency.
  • Original web documents may contain fragments, navigation, markup, lists, tables, or repeated material.
  • Grammatical-agreement errors, literal renderings, terminology errors, incomplete continuations, lost equations, and markup artifacts can remain.
  • Source scores and translation diagnostics are automated signals, not calibrated human judgments.
  • The finite AI-assisted audit is not a substitute for blinded review by multiple independent native-Arabic specialists.
  • Human review is necessary for educational, medical, legal, safety-critical, or publication-facing use.

Source, license, and integrity

The source is FineWeb-Edu `sample-350BT`, a random sample of approximately 350 billion GPT-2 tokens from the full corpus. Each doc_id is the original upstream id, allowing the English source to be retrieved by filtering or indexing that configuration.

FineWeb-Edu is distributed under ODC-By 1.0, and Common Crawl terms also apply. This translated collection is conveyed as a Derivative Database under the same ODC-By 1.0 terms. Preserve the required notice and attribution, consult NOTICE.md, and independently assess source-level rights for the intended use.

release_manifest.json records the row count, compressed size, and SHA-256 digest of every public Parquet file. Public Parquet metadata contains only release identifiers, source attribution, language, license, row count, and public product labels; internal filesystem paths and unnecessary deployment metadata were removed.


<a id="arabic"></a> English

العربية

فاين ويب التعليمي بالعربية — ٢٤ مليون وثيقة

هذه مجموعة عربية واسعة للتدريب المسبق، تضم ٢٤٬٧٩٤٬٤٢٥ وثيقة كاملة وقرابة ٣٤٫٨٦ مليار رمز عربي. جاءت الوثائق من العينة التعليمية الضخمة في المجموعة الأصلية، واحتفظنا بمعرّف كل وثيقة ودرجتها ومؤشرات تساعد الباحث على اختيار المستوى الذي يناسبه.

FineWeb-Edu sample-350BT

مثال من إحدى الوثائق

عمارة سعودية شكّلها المكان. تأخذنا الوثيقة في جولة بين نجد وساحل الخليج والحجاز وعسير، وتبين كيف انسجمت مواد البناء والتصميم مع المناخ والراحة والخصوصية والجمال.

<urn:uuid:f00fd3ac-a6e2-4c3b-8287-3152496dabf4>

أما قسم العرض فيضم واحدًا وسبعين مثالًا مختارًا من بيانات التدريب نفسها، ولا يدخل مرة أخرى في حساب الإحصاءات.

المنهجية

١. إعداد بيانات ترجمة قريبة من مجال العمل

لم نبدأ بترجمة ملايين الوثائق مباشرة. أنشأنا أولًا مجموعة تدريب أصغر وأكثر إحكامًا باستخدام نموذج معلّم أقوى:

DeepSeek-v4-flash

ترجم النموذج نصوصًا تعليمية وعلمية واستدلالية حقيقية، من بينها نصوص من المجموعة الأصلية. وبعد المراجعة والتنقية بقي ١٠٦٬١١٤ زوجًا من الإنجليزية والعربية. وبهذا تعلّمت نماذجنا من مادة تشبه، في موضوعاتها وأسلوبها، ما ستراه لاحقًا أثناء الإنتاج.

nizarun/English_Arabic_Translation_Pairs

٢. تدريب نماذج الترجمة

استخدمنا ثلاثة نماذج من عائلة جيما، ودربناها جميعًا على مجموعة الأزواج السابقة. خضع النموذجان ب١ وب٢ لضبط دقيق كامل، بينما دُرّب النموذج أ بطريقة لورا.

Gemma 4

المنتجأسلوب التدريبعدد الوثائق
النموذج ب١ضبط دقيق كامل١٦٬٤٣٥٬٧٧٦
النموذج ب٢ضبط دقيق كامل١٬٦٤٨٬٧١٥
النموذج ألورا٦٬٧٠٩٬٩٣٤

وتقابلها في الملفات المنشورة الأسماء الآتية، بالترتيب نفسه:

Model B1
Model B2
Model A

٣. اختيار الوثائق الأصلية

كان هدفنا الرئيس هو ترجمة الوثائق التي حصلت على ثلاث درجات ونصف فأكثر في التقييم التعليمي للمصدر. لكن الاعتماد على حد واحد قد يستبعد موضوعات أو أساليب كتابة معينة، لذلك احتفظنا عمدًا بـ ١٬٣١٣٬٦٤٠ وثيقة تقع درجاتها بين ثلاث وثلاث درجات ونصف. جاءت هذه الإضافة كلها ضمن النموذج أ، والغاية منها توسيع التنوع وتقليل أثر الانحياز الموضوعي الذي قد تسببه عتبة آلية واحدة.

من المهم هنا أن الدرجة تخص الوثيقة الإنجليزية الأصلية؛ فهي لا تقيس جودة الترجمة العربية.

score

درجة الوثيقة الأصليةعدد الوثائقالنسبة
من ٣ إلى أقل من ٣٫٥١٬٣١٣٬٦٤٠٥٫٣٠٪
من ٣٫٥ إلى أقل من ٤١٩٬٧٤١٬١٩١٧٩٫٦٢٪
من ٤ إلى أقل من ٤٫٥٣٬٥٦١٬٦٨٧١٤٫٣٦٪
٤٫٥ فأعلى١٧٧٬٩٠٧٠٫٧٢٪
الإجمالي٢٤٬٧٩٤٬٤٢٥١٠٠٪

وبذلك تبلغ نسبة الوثائق ذات الدرجة ثلاث ونصف فأعلى ٩٤٫٧٠٪ من المجموعة.

٤. التنقية في أثناء الترجمة وبعدها

قُسّمت الوثائق الطويلة إلى مقاطع ثم أُعيد جمعها تحت معرّفاتها الأصلية. ولم نعتمد على اختبار واحد؛ فبعض العيوب لا يظهر إلا لحظة التوليد، وبعضها لا يُكتشف إلا بعد عودة الوثيقة إلى صورتها الكاملة.

ما فحصناهالغرض من الفحص
الفراغ وبلوغ الحد الأقصى للتوليدكشف النصوص المفقودة أو المقطوعة، ولا سيما الحلقات التكرارية التي تستهلك كامل المساحة المتاحة
طول المقطع والوثيقةكشف ما سقط من الترجمة، أو تضاعف، أو طال وقصر على نحو غير طبيعي
حضور الكتابة العربيةالتأكد من أن الناتج عربي فعلًا، مع عدم حذف المعادلات والروابط والمراجع والمصطلحات التقنية
التكرار مقارنة بالنص الأصليالتفريق بين التكرار الموجود أصلًا في القوائم والجداول وبين حلقة جديدة صنعها النموذج
تتابع الكلمة نفسهاكشف الانهيار الذي يعيد الكلمة مرارًا
الضغط والتنوعكشف النصوص قليلة المعلومات أو المتكررة آليًا
أنظمة الكتابة غير المتوقعةرصد حروف سيريلية أو يونانية أو يابانية ظهرت من غير سبب
المعرّفاتمنع تكرار الوثائق داخل المنتج أو اشتراكها بين المنتجات

لدينا عمود اسمه:

foreign_chars

ولا يقصد به كل ما ليس عربيًا. أبقينا الحروف اللاتينية حين تكون جزءًا صحيحًا من رابط أو مرجع أو بريد أو اختصار أو اسم كيميائي أو شيفرة. ما يرصد هذا العمود هو ظهور نظام كتابة ثالث لم يكن متوقعًا في الترجمة.

أبقينا مؤشرات الفحص في البيانات بدل إخفائها، حتى يستطيع الباحث تشديد شروطه أو تخفيفها. كما فحصنا المعرّفات كاملة، ولم نجد وثيقة مكررة داخل أي منتج أو مشتركة بين منتجين.

الحجم والطول

المنتجالوثائقرموز المصدرالرموز العربيةالنسبة
النموذج ب١١٦٬٤٣٥٬٧٧٦١٦٬٦٥٣٬٢٢٥٬٤٥٦٢٣٬١١٣٬٦٣١٬٩٢٥١٫٣٨٧٩
النموذج ب٢١٬٦٤٨٬٧١٥١٬٦٢٧٬١٩٠٬٦٥٦٢٬٢٥٢٬٧٤٦٬٩١٧١٫٣٨٤٤
النموذج أ٦٬٧٠٩٬٩٣٤٦٬٨٠٩٬٣٠٣٬٤٦٤٩٬٤٩٧٬٧٧٧٬٢٠٩١٫٣٩٤٨
الإجمالي٢٤٬٧٩٤٬٤٢٥٢٥٬٠٨٩٬٧١٩٬٥٧٦٣٤٬٨٦٤٬١٥٦٬٠٥١١٫٣٨٩٦

هذه أعداد رموز دقيقة بحسب أداة الترميز، وليست تقديرًا لعدد الكلمات. لم نحولها إلى «كلمات» لأن ذلك سيكون رقمًا مضللًا. وحساب نسبة الكلمات بدقة يحتاج إلى مرور جديد على جميع النصوص الإنجليزية الأصلية، وهي غير مكررة داخل هذه النسخة العربية.

يبلغ متوسط حضور الحروف العربية ٠٫٩٧٦٩، ووسيطه ٠٫٩٨٩٢، بينما يبلغ وسيط نسبة طول النص العربي إلى المصدر ١٫٣٩٧٦ بحسب أداة الترميز.

الاستخدام

python
from datasets import load_dataset

dataset = load_dataset(
    "nizarun/FineWeb-Edu-Arabic-24M",
    split="train",
    streaming=True,
)

first = next(iter(dataset))
print(first["ar_text"][:500])

ويمكن اختيار ناتج نموذج بعينه أو التصفية بأي مؤشر منشور:

python
model_a = dataset.filter(lambda row: row["translation_model"] == "Model A")

محتوى كل صف

يحتوي كل صف على النص العربي كاملًا، ومعرّف الوثيقة الإنجليزية الأصلية، ودرجتها التعليمية، واسم منتج الترجمة. وتأتي معها أعداد الرموز ونسب الطول والكتابة العربية والتشكيل، إضافة إلى مؤشرات التكرار والتنوع وأطوال المقاطع. وجود هذه التفاصيل يتيح بناء نسخة تناسب تجربة الباحث بدل فرض اختيار واحد على الجميع.

أسماء الأعمدة كما تظهر في الملفات:

doc_id
ar_text
score
n_chunks
en_tokens
ar_tokens
arabic_ratio
len_ratio
diac_ratio
foreign_chars
rep4_max
rep4_en_max
chunk_len_ratio_min
chunk_len_ratio_max
chunk_ar_tokens_max
chunk_ar_ratio_min
chunk_zr_min
adj_repeat_max
corpus_version
shard
translation_model

تدقيق جودة الترجمة

أجرينا تدقيقًا معمّى على أربعمائة وثيقة كاملة. شملت العينة مئة وعشرين مجموعة ثلاثية متقاربة في الدرجة والطول بين النماذج الثلاثة، وأربعين وثيقة إضافية من الفئة الواقعة بين ثلاث وثلاث درجات ونصف. قُرئت كل وثيقة من أول حرف إلى آخر حرف، وأُعيد فحص خُمس العينة لقياس ثبات التقييم، ثم روجعت الملاحظات مرة أخرى داخل سياق الوثيقة الكامل.

وجد التدقيق ملاحظة لغوية واحدة على الأقل في ٦٨٫٣٪ من وثائق النموذج ب١، و٦٥٪ من وثائق النموذج ب٢، و٦٥٫٨٪ من وثائق النموذج أ. هذا معيار بالغ الحساسية؛ فهو يحتسب حتى الخطأ المنفرد في المطابقة أو النحو أو المصطلح أو الحرفية أو الاتساق أو وضوح المعنى. لذلك لا يصح تفسير هذه النسب على أنها نسبة الوثائق غير الصالحة.

أما الفحص الآلي لفرص المطابقة في النوع النحوي فقد قدّر الدقة بأكثر من ٩٩٫٥٪ لدى المنتجات الثلاثة، لكنه ليس بديلًا عن معيار ذهبي يعدّه البشر.

كانت الأخطاء في الغالب متناثرة، ولم تتكرر بالصورة نفسها بين الوثائق أو النماذج. ومن المتوقع أن يخف أثر الأخطاء غير المنتظمة عند التدريب على هذا الحجم، لكن ذلك ليس ضمانًا؛ فقد يتعلم النموذج أيضًا بعض سمات الترجمة الآلية أو انحيازات النصوص الأصلية.

الاستخدام التدريبي المقترح

هذه المجموعة أنسب للتدريب المسبق واسع النطاق أو للمراحل الوسطى، على أن تُخلط بكمية كبيرة من النصوص العربية الأصلية. ولا نوصي بأن تهيمن على المرحلة الأخيرة من التدريب. عند الاقتراب من النهاية، من الأفضل تقليلها أو استبعادها لصالح عربية أصلية منتقاة، ومواد راجعها البشر، وأمثلة موثوقة للمهام المستهدفة؛ لأن بيانات المرحلة الأخيرة تترك أثرًا أكبر في أسلوب النموذج وجودته اللغوية.

القيود

  • جميع الوثائق العربية مترجمة آليًا.
  • قد تضعف حدود المقاطع الاستمرارية أو إحالة الضمائر أو اتساق المصطلحات.
  • قد تحتوي صفحات الويب الأصلية شذرات أو قوائم تنقل أو ترميزًا أو جداول أو مواد مكررة.
  • قد تبقى أخطاء في المطابقة النحوية أو ترجمات حرفية أو أخطاء مصطلحية أو نهايات ناقصة أو معادلات مفقودة أو آثار ترميز.
  • درجات المصدر ومؤشرات الترجمة آلية وليست أحكامًا بشرية معايرة.
  • التدقيق محدود ومساعد بالذكاء الاصطناعي، ولا يغني عن تقييم معمّى يجريه عدة مختصين مستقلين من أهل العربية.
  • تلزم مراجعة بشرية قبل الاستخدام التعليمي أو الطبي أو القانوني أو الحساس للسلامة أو الموجه للنشر.

المصدر والترخيص وسلامة الإصدار

المصدر عينة عشوائية تقارب ثلاثمائة وخمسين مليار رمز من المجموعة التعليمية الكاملة. واحتفظنا بالمعرّف الأصلي لكل وثيقة، بحيث يمكن لمن يملك فهرس المصدر أن يصل إلى النص الإنجليزي المقابل مباشرة.

FineWeb-Edu sample-350BT

doc_id

تخضع المجموعة الأصلية لترخيص قاعدة البيانات المفتوحة مع وجوب النسب، كما تسري شروط الأرشيف العام الذي جُمعت منه الصفحات. وننشر هذه النسخة المترجمة بوصفها قاعدة بيانات مشتقة وفق الترخيص نفسه. ينبغي إبقاء إشعار الترخيص والنسب، مع تقييم حقوق محتوى الصفحات بحسب الاستخدام المقصود.

ODC-By 1.0

Common Crawl

NOTICE.md

وللتحقق من سلامة الإصدار، أعددنا سجلًا يضم عدد الصفوف والحجم المضغوط والبصمة الرقمية لكل ملف منشور. أما البيانات الوصفية داخل الملفات فاقتصرنا فيها على ما يحتاجه المستخدم، وحذفنا مسارات الأجهزة وأسماء التشغيل الداخلية وكل معلومة لا تخدم البحث.

release_manifest.json

SHA-256

Parquet