nizarun/FineWeb-Edu-Arabic-24M
English العربية FineWeb-Edu Arabic 24M An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics. Highlight Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.
<a id="english"></a> English
FineWeb-Edu Arabic 24M
An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the `sample-350BT` configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics.
Highlight
Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local materials and design to climate, comfort, privacy, and beauty. doc_id: <urn:uuid:f00fd3ac-a6e2-4c3b-8287-3152496dabf4>
The 71-row featured split is a viewer-only showcase duplicated from train; all corpus statistics refer only to train.
Methodology
1. Build an in-domain translation set
We first used the stronger DeepSeek-v4-flash teacher model to translate authentic educational, scientific, and reasoning material, including FineWeb-Edu. After quality filtering, this produced the 106,114-pair nizarun/English_Arabic_Translation_Pairs dataset. This was not a generic translation benchmark: its source distribution was selected to resemble the educational and technical documents the production models would later translate.
2. Fine-tune the Gemma 4 translators
All three production translators belong to the Gemma 4 family and were fine-tuned on the same English–Arabic dataset above.
The public labels distinguish translation products while avoiding unnecessary deployment-specific metadata.
3. Select the FineWeb-Edu source documents
The main target was documents with a FineWeb-Edu source score of 3.5 or higher. We also intentionally retained 1,313,640 documents scoring 3.0–<3.5 in Model A. A hard cutoff can change the mix of subjects and writing styles, so this lower-score band broadens coverage and helps reduce subject-selection bias introduced by relying on one automated threshold.
The score column describes the original English document, not the Arabic translation and not translation quality.
Overall, 23,480,785 documents (94.70%) have a source score of at least 3.5.
4. Filter during translation and after reassembly
Documents were translated in chunks and reassembled under their original IDs. Quality control was applied both while generating chunks and again after complete documents were rebuilt.
In this release, foreign_chars does not mean every non-Arabic character or every foreign-language word. Latin characters were permitted because correct translations often need URLs, DOI references, email addresses, acronyms, chemical names, code, and citations. The signal counts characters from unsupported third scripts that the translation pipeline was not expected to produce.
The product-specific thresholds and measured signals remain in the public rows so users can build stricter or looser subsets. Exact ID checks found no repeated doc_id within a product and no shared IDs across the three products.
Scale and length
These are tokenizer-derived counts, not whitespace-delimited word counts. We report the exact reproducible measure available in the corpus rather than estimate “words” from tokens. An exact Arabic-to-English word ratio would require another full pass over the upstream English text, which is not redistributed in this Arabic-only release.
Mean Arabic-script ratio is 0.9769, median Arabic-script ratio is 0.9892, and the median Arabic/source token-length ratio is 1.3976.
Usage
from datasets import load_dataset
dataset = load_dataset(
"nizarun/FineWeb-Edu-Arabic-24M",
split="train",
streaming=True,
)
first = next(iter(dataset))
print(first["ar_text"][:500])Filter by translation product or any documented signal:
model_a = dataset.filter(lambda row: row["translation_model"] == "Model A")Schema
Translation-quality audit
A deterministic, blinded, full-document audit reviewed 400 documents: 120 score-and-length-matched triplets across the three products, plus 40 Model A documents from the 3.0–<3.5 band. Every selected document was reviewed from first to last character, 20% received repeat annotation, and all initially cited errors received a conservative full-context adjudication.
In the matched sample, at least one core linguistic issue was identified in 68.3% of Model B1, 65.0% of Model B2, and 65.8% of Model A documents. This demanding category counts even an isolated agreement, grammar, terminology, literal-language, consistency, or intelligibility issue; it should not be read as the percentage of unusable documents. Parser-assisted estimates placed gender-agreement accuracy above 99.5% for all three products, but these are not hand-counted gold measurements.
Observed errors were generally sparse and inconsistent across documents and products. At massive pretraining scale they may average out, but this is an informed expectation, not a guarantee: translationese, source-corpus bias, and recurring model tendencies can still be learned.
Recommended training use
Use this corpus for large-scale pretraining or mid-training, mixed with substantial native-Arabic data. We do not recommend making it a dominant source in the final annealing stage. Near the end of training, taper or exclude it in favor of carefully curated native Arabic, human-reviewed material, and high-confidence task-specific examples, because late-stage data can disproportionately shape final style and language quality.
Limitations
- Every Arabic document is machine translated.
- Chunk boundaries can weaken continuation, pronoun reference, or terminology consistency.
- Original web documents may contain fragments, navigation, markup, lists, tables, or repeated material.
- Grammatical-agreement errors, literal renderings, terminology errors, incomplete continuations, lost equations, and markup artifacts can remain.
- Source scores and translation diagnostics are automated signals, not calibrated human judgments.
- The finite AI-assisted audit is not a substitute for blinded review by multiple independent native-Arabic specialists.
- Human review is necessary for educational, medical, legal, safety-critical, or publication-facing use.
Source, license, and integrity
The source is FineWeb-Edu `sample-350BT`, a random sample of approximately 350 billion GPT-2 tokens from the full corpus. Each doc_id is the original upstream id, allowing the English source to be retrieved by filtering or indexing that configuration.
FineWeb-Edu is distributed under ODC-By 1.0, and Common Crawl terms also apply. This translated collection is conveyed as a Derivative Database under the same ODC-By 1.0 terms. Preserve the required notice and attribution, consult NOTICE.md, and independently assess source-level rights for the intended use.
release_manifest.json records the row count, compressed size, and SHA-256 digest of every public Parquet file. Public Parquet metadata contains only release identifiers, source attribution, language, license, row count, and public product labels; internal filesystem paths and unnecessary deployment metadata were removed.
<a id="arabic"></a> English
العربية
فاين ويب التعليمي بالعربية — ٢٤ مليون وثيقة
هذه مجموعة عربية واسعة للتدريب المسبق، تضم ٢٤٬٧٩٤٬٤٢٥ وثيقة كاملة وقرابة ٣٤٫٨٦ مليار رمز عربي. جاءت الوثائق من العينة التعليمية الضخمة في المجموعة الأصلية، واحتفظنا بمعرّف كل وثيقة ودرجتها ومؤشرات تساعد الباحث على اختيار المستوى الذي يناسبه.
مثال من إحدى الوثائق
عمارة سعودية شكّلها المكان. تأخذنا الوثيقة في جولة بين نجد وساحل الخليج والحجاز وعسير، وتبين كيف انسجمت مواد البناء والتصميم مع المناخ والراحة والخصوصية والجمال.
<urn:uuid:f00fd3ac-a6e2-4c3b-8287-3152496dabf4>
أما قسم العرض فيضم واحدًا وسبعين مثالًا مختارًا من بيانات التدريب نفسها، ولا يدخل مرة أخرى في حساب الإحصاءات.
المنهجية
١. إعداد بيانات ترجمة قريبة من مجال العمل
لم نبدأ بترجمة ملايين الوثائق مباشرة. أنشأنا أولًا مجموعة تدريب أصغر وأكثر إحكامًا باستخدام نموذج معلّم أقوى:
DeepSeek-v4-flash
ترجم النموذج نصوصًا تعليمية وعلمية واستدلالية حقيقية، من بينها نصوص من المجموعة الأصلية. وبعد المراجعة والتنقية بقي ١٠٦٬١١٤ زوجًا من الإنجليزية والعربية. وبهذا تعلّمت نماذجنا من مادة تشبه، في موضوعاتها وأسلوبها، ما ستراه لاحقًا أثناء الإنتاج.
nizarun/English_Arabic_Translation_Pairs
٢. تدريب نماذج الترجمة
استخدمنا ثلاثة نماذج من عائلة جيما، ودربناها جميعًا على مجموعة الأزواج السابقة. خضع النموذجان ب١ وب٢ لضبط دقيق كامل، بينما دُرّب النموذج أ بطريقة لورا.
Gemma 4
وتقابلها في الملفات المنشورة الأسماء الآتية، بالترتيب نفسه:
Model B1
Model B2
Model A٣. اختيار الوثائق الأصلية
كان هدفنا الرئيس هو ترجمة الوثائق التي حصلت على ثلاث درجات ونصف فأكثر في التقييم التعليمي للمصدر. لكن الاعتماد على حد واحد قد يستبعد موضوعات أو أساليب كتابة معينة، لذلك احتفظنا عمدًا بـ ١٬٣١٣٬٦٤٠ وثيقة تقع درجاتها بين ثلاث وثلاث درجات ونصف. جاءت هذه الإضافة كلها ضمن النموذج أ، والغاية منها توسيع التنوع وتقليل أثر الانحياز الموضوعي الذي قد تسببه عتبة آلية واحدة.
من المهم هنا أن الدرجة تخص الوثيقة الإنجليزية الأصلية؛ فهي لا تقيس جودة الترجمة العربية.
score
وبذلك تبلغ نسبة الوثائق ذات الدرجة ثلاث ونصف فأعلى ٩٤٫٧٠٪ من المجموعة.
٤. التنقية في أثناء الترجمة وبعدها
قُسّمت الوثائق الطويلة إلى مقاطع ثم أُعيد جمعها تحت معرّفاتها الأصلية. ولم نعتمد على اختبار واحد؛ فبعض العيوب لا يظهر إلا لحظة التوليد، وبعضها لا يُكتشف إلا بعد عودة الوثيقة إلى صورتها الكاملة.
لدينا عمود اسمه:
foreign_chars
ولا يقصد به كل ما ليس عربيًا. أبقينا الحروف اللاتينية حين تكون جزءًا صحيحًا من رابط أو مرجع أو بريد أو اختصار أو اسم كيميائي أو شيفرة. ما يرصد هذا العمود هو ظهور نظام كتابة ثالث لم يكن متوقعًا في الترجمة.
أبقينا مؤشرات الفحص في البيانات بدل إخفائها، حتى يستطيع الباحث تشديد شروطه أو تخفيفها. كما فحصنا المعرّفات كاملة، ولم نجد وثيقة مكررة داخل أي منتج أو مشتركة بين منتجين.
الحجم والطول
هذه أعداد رموز دقيقة بحسب أداة الترميز، وليست تقديرًا لعدد الكلمات. لم نحولها إلى «كلمات» لأن ذلك سيكون رقمًا مضللًا. وحساب نسبة الكلمات بدقة يحتاج إلى مرور جديد على جميع النصوص الإنجليزية الأصلية، وهي غير مكررة داخل هذه النسخة العربية.
يبلغ متوسط حضور الحروف العربية ٠٫٩٧٦٩، ووسيطه ٠٫٩٨٩٢، بينما يبلغ وسيط نسبة طول النص العربي إلى المصدر ١٫٣٩٧٦ بحسب أداة الترميز.
الاستخدام
from datasets import load_dataset
dataset = load_dataset(
"nizarun/FineWeb-Edu-Arabic-24M",
split="train",
streaming=True,
)
first = next(iter(dataset))
print(first["ar_text"][:500])ويمكن اختيار ناتج نموذج بعينه أو التصفية بأي مؤشر منشور:
model_a = dataset.filter(lambda row: row["translation_model"] == "Model A")محتوى كل صف
يحتوي كل صف على النص العربي كاملًا، ومعرّف الوثيقة الإنجليزية الأصلية، ودرجتها التعليمية، واسم منتج الترجمة. وتأتي معها أعداد الرموز ونسب الطول والكتابة العربية والتشكيل، إضافة إلى مؤشرات التكرار والتنوع وأطوال المقاطع. وجود هذه التفاصيل يتيح بناء نسخة تناسب تجربة الباحث بدل فرض اختيار واحد على الجميع.
أسماء الأعمدة كما تظهر في الملفات:
doc_id
ar_text
score
n_chunks
en_tokens
ar_tokens
arabic_ratio
len_ratio
diac_ratio
foreign_chars
rep4_max
rep4_en_max
chunk_len_ratio_min
chunk_len_ratio_max
chunk_ar_tokens_max
chunk_ar_ratio_min
chunk_zr_min
adj_repeat_max
corpus_version
shard
translation_modelتدقيق جودة الترجمة
أجرينا تدقيقًا معمّى على أربعمائة وثيقة كاملة. شملت العينة مئة وعشرين مجموعة ثلاثية متقاربة في الدرجة والطول بين النماذج الثلاثة، وأربعين وثيقة إضافية من الفئة الواقعة بين ثلاث وثلاث درجات ونصف. قُرئت كل وثيقة من أول حرف إلى آخر حرف، وأُعيد فحص خُمس العينة لقياس ثبات التقييم، ثم روجعت الملاحظات مرة أخرى داخل سياق الوثيقة الكامل.
وجد التدقيق ملاحظة لغوية واحدة على الأقل في ٦٨٫٣٪ من وثائق النموذج ب١، و٦٥٪ من وثائق النموذج ب٢، و٦٥٫٨٪ من وثائق النموذج أ. هذا معيار بالغ الحساسية؛ فهو يحتسب حتى الخطأ المنفرد في المطابقة أو النحو أو المصطلح أو الحرفية أو الاتساق أو وضوح المعنى. لذلك لا يصح تفسير هذه النسب على أنها نسبة الوثائق غير الصالحة.
أما الفحص الآلي لفرص المطابقة في النوع النحوي فقد قدّر الدقة بأكثر من ٩٩٫٥٪ لدى المنتجات الثلاثة، لكنه ليس بديلًا عن معيار ذهبي يعدّه البشر.
كانت الأخطاء في الغالب متناثرة، ولم تتكرر بالصورة نفسها بين الوثائق أو النماذج. ومن المتوقع أن يخف أثر الأخطاء غير المنتظمة عند التدريب على هذا الحجم، لكن ذلك ليس ضمانًا؛ فقد يتعلم النموذج أيضًا بعض سمات الترجمة الآلية أو انحيازات النصوص الأصلية.
الاستخدام التدريبي المقترح
هذه المجموعة أنسب للتدريب المسبق واسع النطاق أو للمراحل الوسطى، على أن تُخلط بكمية كبيرة من النصوص العربية الأصلية. ولا نوصي بأن تهيمن على المرحلة الأخيرة من التدريب. عند الاقتراب من النهاية، من الأفضل تقليلها أو استبعادها لصالح عربية أصلية منتقاة، ومواد راجعها البشر، وأمثلة موثوقة للمهام المستهدفة؛ لأن بيانات المرحلة الأخيرة تترك أثرًا أكبر في أسلوب النموذج وجودته اللغوية.
القيود
- جميع الوثائق العربية مترجمة آليًا.
- قد تضعف حدود المقاطع الاستمرارية أو إحالة الضمائر أو اتساق المصطلحات.
- قد تحتوي صفحات الويب الأصلية شذرات أو قوائم تنقل أو ترميزًا أو جداول أو مواد مكررة.
- قد تبقى أخطاء في المطابقة النحوية أو ترجمات حرفية أو أخطاء مصطلحية أو نهايات ناقصة أو معادلات مفقودة أو آثار ترميز.
- درجات المصدر ومؤشرات الترجمة آلية وليست أحكامًا بشرية معايرة.
- التدقيق محدود ومساعد بالذكاء الاصطناعي، ولا يغني عن تقييم معمّى يجريه عدة مختصين مستقلين من أهل العربية.
- تلزم مراجعة بشرية قبل الاستخدام التعليمي أو الطبي أو القانوني أو الحساس للسلامة أو الموجه للنشر.
المصدر والترخيص وسلامة الإصدار
المصدر عينة عشوائية تقارب ثلاثمائة وخمسين مليار رمز من المجموعة التعليمية الكاملة. واحتفظنا بالمعرّف الأصلي لكل وثيقة، بحيث يمكن لمن يملك فهرس المصدر أن يصل إلى النص الإنجليزي المقابل مباشرة.
FineWeb-Edu sample-350BT
doc_id
تخضع المجموعة الأصلية لترخيص قاعدة البيانات المفتوحة مع وجوب النسب، كما تسري شروط الأرشيف العام الذي جُمعت منه الصفحات. وننشر هذه النسخة المترجمة بوصفها قاعدة بيانات مشتقة وفق الترخيص نفسه. ينبغي إبقاء إشعار الترخيص والنسب، مع تقييم حقوق محتوى الصفحات بحسب الاستخدام المقصود.
ODC-By 1.0
Common Crawl
NOTICE.md
وللتحقق من سلامة الإصدار، أعددنا سجلًا يضم عدد الصفوف والحجم المضغوط والبصمة الرقمية لكل ملف منشور. أما البيانات الوصفية داخل الملفات فاقتصرنا فيها على ما يحتاجه المستخدم، وحذفنا مسارات الأجهزة وأسماء التشغيل الداخلية وكل معلومة لا تخدم البحث.
release_manifest.json
SHA-256
Parquet
