CoolFace
18 results

darja

DarjaCore /asjp-cerist ASJP / CERIST Academic PDF Corpus Overview DarjaCore ASJP/CERIST is a large-scale collection of academic PDF documents collected from the Algerian ASJP (Algerian Scientific Journal Platform) ecosystem associated with CERIST (Centre de Recherche sur l'Information Scientifique et Technique). The dataset preserves the original PDF documents rather than converting them into extracted text. This allows researchers to perform their own OCR, document parsing… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/asjp-cerist.text-retrieval10K<n<100K1 likes733 downloads15d agoHugging Facetouati-kamel /algerian-darja-corpus Algerian Darja Corpus A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts. Dataset Summary The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.tabulartext-generation10K<n<100K6 likes271 downloads4d agoHugging FaceDarjaCore /algerian-darja-forum-posts Algerian Darja Dataset A large-scale Algerian Darja conversational dataset prepared for NLP, language-model training, instruction tuning, and conversational AI research. 3,209,157 samples · 1.193B tokens · 371.83 tokens/sample on average Dataset at a Glance Property Value Samples 3,209,157 Total tokens 1,193,257,847 Approx. tokens 1.193B Average tokens / sample 371.83 Language Algerian Darja Format Conversational JSON Storage format… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-forum-posts.texttext-generation1M<n<10M2 likes164 downloads13d agoHugging FaceDarjaCore /algerian-darja-sample Algerian Darja Sample A growing Algerian Darja text corpus collected for NLP and language-modeling research. This dataset is updated incrementally as new sources are collected and processed. Exact sample counts, file sizes, and statistics change between releases — refer to the Dataset Viewer on the repository page for current figures rather than any numbers in this card. Dataset at a Glance Sample count, file size, character/word counts, and other metrics are… See the full description on the dataset page: https://huggingface.co/datasets/DarjaCore/algerian-darja-sample.texttext-generation1M<n<10M1 likes93 downloads4d agoHugging Facealgerian-nlp /algerian-darja-corpus Algerian Darja Corpus 11,151 long-form conversational transcripts in Algerian Darja for language modeling of real spoken Algerian, by Kamel Touati (Independent AI Researcher, Algiers, ORCID 0009-0000-5330-2123, Hugging Face touati-kamel), mirrored on the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/algerian-darja-corpus: 11,151 train rows) and re-counted row-by-row with datasets streaming… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/algerian-darja-corpus.tabulartext-generation10K<n<100K0 likes91 downloads5d agoHugging Facelouam123 /darja-tounsi1texttranslation1K<n<10K0 likes84 downloads3y agoHugging Face