CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01touati-kamel /algerian-darja-corpus Algerian Darja Corpus A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts. Dataset Summary The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.tabulartext-generation10K<n<100K6 likes286 downloads7d agoHugging Face02kamruzzaman-asif /bangla-instruction-dataset 🧠 Bangla Instruction Dataset This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models. 📚 Dataset Splits The dataset is organized into the following splits: Split Name Source Dataset Description OdiaGenAI OdiaGenAI/all_combined_bengali_252k A large-scale collection of diverse Bangla instructions and responses. chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.texttext-generation1M<n<10M1 likes219 downloads1y agoHugging Face03Kamtera /Persian-conversational-datasetpersian-conversational-datasettexttext-generation100K<n<1M12 likes155 downloads3y agoHugging Face04KamDickGoon /Killer Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the ids of duplicated documents which can be used to create a dataset with 20B deduplicated documents. Check out our blog post for more details on the… See the full description on the dataset page: https://huggingface.co/datasets/KamDickGoon/Killer.texttext-generation1M<n<10M0 likes131 downloads5mo agoHugging Face05Lyon28 /kamus-besar-bahasa-indonesiatexttext-generation100K<n<1M2 likes76 downloads1y agoHugging Face06islam-kamel /MBPP-Thinking-Gate-1k MBPP Thinking-Gate SFT Dataset This package contains two related assets: Ready 1,000-row MBPP-style dataset (all.jsonl, train.jsonl, validation.jsonl). It is synthetic and designed to test/train autonomous routing between <DIRECT> and <THINK>. Official-MBPP builder (build_from_official_mbpp.py). Run this to create the production dataset from the official Google Research MBPP source. Why two response modes? The training target starts with one of two routing… See the full description on the dataset page: https://huggingface.co/datasets/islam-kamel/MBPP-Thinking-Gate-1k.tabulartext-generation1K<n<10K0 likes63 downloads4d agoHugging Face07kamandmesbah /UF_DPOEach row has chosen and rejected string fields containing the linearized multi-turn dialogue in the form: Human: ... Assistant: ... Splits data/train.jsonl data/test.jsonl Generated on 2025-08-08. texttext-generation10K<n<100K0 likes38 downloads1y agoHugging Face08kamaalg /azerbaijani-instructions Azerbaijani Instruction Dataset (v0) Azerbaijani (instruction, response) pairs for supervised fine-tuning (SFT) of Azerbaijani language models — part of an open Azerbaijani LLM stack. Instruction data is genuinely scarce for Azerbaijani; this is both training data for our models and a reusable standalone artifact for anyone building Azerbaijani instruction-following models. Contents seeds_az.jsonl — 45 hand-authored, high-quality seed pairs spanning 16 task… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-instructions.texttext-generationn<1K1 likes34 downloads3mo agoHugging Face09kamjke /federal-general-courts-rugated Federal General Courts Russia — SUDRF Датасет уголовных судебных решений федеральных судов общей юрисдикции РФ за 2025-2026 годы, собранных с портала ГАС «Правосудие» (sudrf.ru) Колонки Колонка Описание court_name Название суда caseNumber Номер дела entryDate Дата поступления дела judge Судья resultDate Дата решения decision Итог («Вынесен ПРИГОВОР» и т.п.) offense_article Статья УК РФ text Текст решения (очищен) Источник… See the full description on the dataset page: https://huggingface.co/datasets/kamjke/federal-general-courts-ru.texttext-classification10K<n<100K4 likes33 downloads3mo agoHugging Face10Kamisori-daijin /email-datasets-20k Dataset Summary There are 20,000 samples of emails. This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ). License Note This dataset is licensed under Apache 2.0. Please also refer to the Gemma Terms of Use and Prohibited Use Policy regarding the use of Gemma-generated content. texttext-generation10K<n<100K3 likes28 downloads6mo agoHugging Face11kamaalg /azerbaijani-corpus-v0 Azerbaijani Pretraining Corpus (v0) A cleaned, deduplicated, PII-redacted ~1.0 billion token Latin-script Azerbaijani corpus for language-model pretraining, built with a reproducible datatrove pipeline from open multilingual web + encyclopedic sources. Full provenance, methodology, and limitations are in the Datasheet (Gebru-style). Summary Language Azerbaijani (az/azj), Latin script only Documents 1,711,442 Tokens ~1.0B (az_unigram_32k; train… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-corpus-v0.texttext-generation1K<n<10K1 likes20 downloads3mo agoHugging Face12Kamisori-daijin /email-datasets-v2-100k Dataset Summary There are 99336 samples of emails. This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ). format: {"id": , "instruction": "Prompt Is Here", "text": "<user>Prompt Is Here <think> - Goal: Goal Is Here - Reason: Reason Is Here - Tone: Tone Is Here </think> <generate> Mail Is Here </generate></s>"} Link Github: https://github.com/kamisori-daijin/email-datasets License Note This dataset is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Kamisori-daijin/email-datasets-v2-100k.texttext-generation10K<n<100K0 likes17 downloads5mo agoHugging Face13afrizalha /KamusOne-28M-Indonesian KamusOne (Kamus-1) is a synthethic Indonesian language dataset, generated by Mixtral8x7B. About This dataset was generated by Mixtral 8x7B. For the procedure, Mixtral is instructed that it will act as an Indonesian language dictionary, a native Indonesian speaker, etc. and that it will explain the meaning of a series of Indonesian words. Hence, the name of the dataset ("Kamus", literally "dictionary"). Construction of the word list goes like this. First, we extracted word frequency… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/KamusOne-28M-Indonesian.texttext-generation100K<n<1M3 likes16 downloads2y agoHugging Face14Kamka-IT /TATA-NOTATA-FineMistral-nucleotide_transformer_downstream_tasksDataset for Fine-tuning Mistral Model on Tata and No Tata Sequences Description This dataset is specifically curated for training the Mistral model to distinguish between 'tata' and 'no tata' sequences. It is derived and reformatted from a dataset originally created by InstaDeep, tailored to enhance the performance of natural language processing models in identifying specific patterns. Dataset Information Features: This dataset consists of sequences represented as strings under the… See the full description on the dataset page: https://huggingface.co/datasets/Kamka-IT/TATA-NOTATA-FineMistral-nucleotide_transformer_downstream_tasks.texttext-generation10K<n<100K0 likes16 downloads2y agoHugging Face15Kamran-56 /prompt-refinement-dataset Prompt Refinement Dataset Dataset Summary The Prompt Refinement Dataset is a curated collection of 4,349 input-output pairs designed to train language models to transform basic, vague prompts into high-quality, detailed, and structured prompts that elicit significantly better responses from AI systems. Each pair consists of a raw user-written prompt as the input and an expertly engineered version of the same prompt as the output — preserving the original intent while… See the full description on the dataset page: https://huggingface.co/datasets/Kamran-56/prompt-refinement-dataset.texttext-generation1K<n<10K2 likes16 downloads7mo agoHugging Face16kamal-018 /Maithili_Poemsgated Maithili Poetry Dataset Dataset Summary This dataset is a curated corpus of Maithili poetry collected from multiple online repositories and digitized literary sources. It is normalized and structured for language modeling, tokenization experiments, and generative poetry tasks in Maithili. Dataset Statistics Metric Value Language Maithili (mai) File Size ~1.97 MB (Uncompressed) Token Count ~0.45 Million Word Count ~288,500 Line… See the full description on the dataset page: https://huggingface.co/datasets/kamal-018/Maithili_Poems.imagetext-generation10K<n<100K0 likes15 downloads2mo agoHugging Face17Mori-kamiyama /morikawa_mixed_dir_5k morikawa_mixed_dir_5k Mixed structured-output SFT dataset generated on 2026-02-08. Composition Total rows: 5000 Target mix policy: JSON/XML/CSV combined 20%, YAML 40%, TOML 40% Target mix policy source: directional/target mix summary Pair coverage policy: n/a Actual realized mix: JSON: 0 rows (0.00%) YAML: 0 rows (0.00%) XML: 0 rows (0.00%) TOML: 0 rows (0.00%) CSV: 0 rows (0.00%) Source Datasets and Licenses Important License Note This… See the full description on the dataset page: https://huggingface.co/datasets/Mori-kamiyama/morikawa_mixed_dir_5k.texttext-generation1K<n<10K0 likes13 downloads8mo agoHugging Face18kamarko /ccnews-french-subsetCredits and Attribution: This dataset is derived from the Common Crawl dataset (https://huggingface.co/datasets/stanford-oval/ccnews). The data has been transformed and filtered to achieve the current format. For license information, please refer to https://commoncrawl.org/terms-of-use texttext-classification1K<n<10K0 likes11 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.