datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
algerian-darja-corpus
Algerian Darja Corpus
A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
Dataset Summary
The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.bangla-instruction-dataset
🧠 Bangla Instruction Dataset
This dataset repository consolidates high-quality instruction-tuning data from multiple popular sources, structured for easy use in training and evaluating instruction-following models.
📚 Dataset Splits
The dataset is organized into the following splits:
Split Name
Source Dataset
Description
OdiaGenAI
OdiaGenAI/all_combined_bengali_252k
A large-scale collection of diverse Bangla instructions and responses.
chrononeel… See the full description on the dataset page: https://huggingface.co/datasets/kamruzzaman-asif/bangla-instruction-dataset.Persian-conversational-datasetpersian-conversational-datasetKiller
Getting Started
RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text
documents coming from 84 CommonCrawl snapshots and processed using
the CCNet pipeline. Out of these, there are 30B documents in the corpus
that additionally come with quality signals. In addition, we also provide the ids of duplicated documents which can be
used to create a dataset with 20B deduplicated documents.
Check out our blog post for more details on the… See the full description on the dataset page: https://huggingface.co/datasets/KamDickGoon/Killer.kamus-besar-bahasa-indonesiaMBPP-Thinking-Gate-1k
MBPP Thinking-Gate SFT Dataset
This package contains two related assets:
Ready 1,000-row MBPP-style dataset (all.jsonl, train.jsonl, validation.jsonl).
It is synthetic and designed to test/train autonomous routing between <DIRECT> and <THINK>.
Official-MBPP builder (build_from_official_mbpp.py).
Run this to create the production dataset from the official Google Research MBPP source.
Why two response modes?
The training target starts with one of two routing… See the full description on the dataset page: https://huggingface.co/datasets/islam-kamel/MBPP-Thinking-Gate-1k.UF_DPOEach row has chosen and rejected string fields containing the linearized multi-turn dialogue in the form:
Human: ...
Assistant: ...
Splits
data/train.jsonl
data/test.jsonl
Generated on 2025-08-08.
azerbaijani-instructions
Azerbaijani Instruction Dataset (v0)
Azerbaijani (instruction, response) pairs for supervised fine-tuning (SFT) of Azerbaijani language
models — part of an open Azerbaijani LLM stack. Instruction data is genuinely scarce for Azerbaijani;
this is both training data for our models and a reusable standalone artifact for anyone building
Azerbaijani instruction-following models.
Contents
seeds_az.jsonl — 45 hand-authored, high-quality seed pairs spanning 16 task… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-instructions.federal-general-courts-ru
Federal General Courts Russia — SUDRF
Датасет уголовных судебных решений федеральных судов общей юрисдикции РФ
за 2025-2026 годы, собранных с портала ГАС «Правосудие» (sudrf.ru)
Колонки
Колонка
Описание
court_name
Название суда
caseNumber
Номер дела
entryDate
Дата поступления дела
judge
Судья
resultDate
Дата решения
decision
Итог («Вынесен ПРИГОВОР» и т.п.)
offense_article
Статья УК РФ
text
Текст решения (очищен)
Источник… See the full description on the dataset page: https://huggingface.co/datasets/kamjke/federal-general-courts-ru.email-datasets-20k
Dataset Summary
There are 20,000 samples of emails.
This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ).
License Note
This dataset is licensed under Apache 2.0. Please also refer to the Gemma Terms of Use and Prohibited Use Policy regarding the use of Gemma-generated content.
azerbaijani-corpus-v0
Azerbaijani Pretraining Corpus (v0)
A cleaned, deduplicated, PII-redacted ~1.0 billion token Latin-script Azerbaijani corpus for
language-model pretraining, built with a reproducible datatrove
pipeline from open multilingual web + encyclopedic sources. Full provenance, methodology, and limitations
are in the Datasheet (Gebru-style).
Summary
Language
Azerbaijani (az/azj), Latin script only
Documents
1,711,442
Tokens
~1.0B (az_unigram_32k; train… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-corpus-v0.email-datasets-v2-100k
Dataset Summary
There are 99336 samples of emails.
This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ).
format:
{"id": , "instruction": "Prompt Is Here",
"text":
"<user>Prompt Is Here
<think>
- Goal: Goal Is Here
- Reason: Reason Is Here
- Tone: Tone Is Here
</think>
<generate>
Mail Is Here
</generate></s>"}
Link
Github: https://github.com/kamisori-daijin/email-datasets
License Note
This dataset is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Kamisori-daijin/email-datasets-v2-100k.KamusOne-28M-Indonesian
KamusOne (Kamus-1) is a synthethic Indonesian language dataset, generated by Mixtral8x7B.
About
This dataset was generated by Mixtral 8x7B. For the procedure, Mixtral is instructed that it will act as an Indonesian language dictionary, a native Indonesian speaker, etc. and that it will explain the meaning of a series of Indonesian words. Hence, the name of the dataset ("Kamus", literally "dictionary"). Construction of the word list goes like this. First, we extracted word frequency… See the full description on the dataset page: https://huggingface.co/datasets/afrizalha/KamusOne-28M-Indonesian.TATA-NOTATA-FineMistral-nucleotide_transformer_downstream_tasksDataset for Fine-tuning Mistral Model on Tata and No Tata Sequences
Description
This dataset is specifically curated for training the Mistral model to distinguish between 'tata' and 'no tata' sequences. It is derived and reformatted from a dataset originally created by InstaDeep, tailored to enhance the performance of natural language processing models in identifying specific patterns.
Dataset Information
Features: This dataset consists of sequences represented as strings under the… See the full description on the dataset page: https://huggingface.co/datasets/Kamka-IT/TATA-NOTATA-FineMistral-nucleotide_transformer_downstream_tasks.prompt-refinement-dataset
Prompt Refinement Dataset
Dataset Summary
The Prompt Refinement Dataset is a curated collection of 4,349 input-output pairs
designed to train language models to transform basic, vague prompts into high-quality,
detailed, and structured prompts that elicit significantly better responses from AI systems.
Each pair consists of a raw user-written prompt as the input and an expertly
engineered version of the same prompt as the output — preserving the original
intent while… See the full description on the dataset page: https://huggingface.co/datasets/Kamran-56/prompt-refinement-dataset.Maithili_Poems
Maithili Poetry Dataset
Dataset Summary
This dataset is a curated corpus of Maithili poetry collected from multiple online repositories and digitized literary sources. It is normalized and structured for language modeling, tokenization experiments, and generative poetry tasks in Maithili.
Dataset Statistics
Metric
Value
Language
Maithili (mai)
File Size
~1.97 MB (Uncompressed)
Token Count
~0.45 Million
Word Count
~288,500
Line… See the full description on the dataset page: https://huggingface.co/datasets/kamal-018/Maithili_Poems.morikawa_mixed_dir_5k
morikawa_mixed_dir_5k
Mixed structured-output SFT dataset generated on 2026-02-08.
Composition
Total rows: 5000
Target mix policy: JSON/XML/CSV combined 20%, YAML 40%, TOML 40%
Target mix policy source: directional/target mix summary
Pair coverage policy: n/a
Actual realized mix:
JSON: 0 rows (0.00%)
YAML: 0 rows (0.00%)
XML: 0 rows (0.00%)
TOML: 0 rows (0.00%)
CSV: 0 rows (0.00%)
Source Datasets and Licenses
Important License Note
This… See the full description on the dataset page: https://huggingface.co/datasets/Mori-kamiyama/morikawa_mixed_dir_5k.ccnews-french-subsetCredits and Attribution:
This dataset is derived from the Common Crawl dataset (https://huggingface.co/datasets/stanford-oval/ccnews).
The data has been transformed and filtered to achieve the current format.
For license information, please refer to https://commoncrawl.org/terms-of-use
