datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
algerian-darja-corpus
Algerian Darja Corpus
A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts.
Dataset Summary
The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.agentvidbench
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline)
by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI)
Layout
.
├── README.md
├── questions.jsonl # 100 rows — one per question
├── videos.jsonl # 71 rows — one per unique video
├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench.agentvidbench-sample
AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above.
Sample selection
The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.mini-carla-192x320-wan-2p2-vae
mini-carla-192x320-wan-2p2-vae
Wan2.2-VAE-encoded latents of a small CARLA driving dataset (192x320, native resolution,
no resize). Produced for training miniworld, a
minimal flow-matching world-model framework, by caching pixel clips through the frozen
pretrained Wan2.2 video VAE instead of a locally-trained one.
Data size
Source pixel dataset
mini_192x320_low: 320 episodes, 600 frames each (192x320, 20 fps) — ~33 GB
Clips in this cache
1,920 (6… See the full description on the dataset page: https://huggingface.co/datasets/kamwoh/mini-carla-192x320-wan-2p2-vae.kamen-fight-matrixMBPP-Thinking-Gate-1k
MBPP Thinking-Gate SFT Dataset
This package contains two related assets:
Ready 1,000-row MBPP-style dataset (all.jsonl, train.jsonl, validation.jsonl).
It is synthetic and designed to test/train autonomous routing between <DIRECT> and <THINK>.
Official-MBPP builder (build_from_official_mbpp.py).
Run this to create the production dataset from the official Google Research MBPP source.
Why two response modes?
The training target starts with one of two routing… See the full description on the dataset page: https://huggingface.co/datasets/islam-kamel/MBPP-Thinking-Gate-1k.Metamath2Py
Links
Github with source code: https://github.com/kamushekp/metamath2py
Paper: https://github.com/kamushekp/metamath2py/blob/main/out/main.pdf
Dataset Structure
The Metamath2Py Dataset consists of the following components:
1. JSONL File on Hugging Face
The dataset is provided as a JSONL file, where each line is a JSON object with the following fields:
original_name: The original name of the statement in the Metamath system.
name: The statement name in our… See the full description on the dataset page: https://huggingface.co/datasets/kamushekp/Metamath2Py.DziriEval
DziriEval : Benchmark d'Évaluation des LLMs en Dialecte Algérien (Darja)
DziriEval est le benchmark académique natif de questions-réponses à choix multiples (QCM) conçu spécifiquement pour évaluer les capacités de compréhension, de raisonnement et de connaissance culturelle des grands modèles de langage (LLMs) sur le dialecte algérien (Darja).
Ce jeu de données a été construit et vérifié manuellement afin de refléter la richesse linguistique, culturelle et quotidienne de… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/DziriEval.experiment-001-budget-boxed
KamiBench Experiment 001 — budget-boxed agents in Kamigotchi
Complete agentic traces from experiment 001, the KamiBench
calibration run: three LLM agents dropped into
Kamigotchi, a live, persistent, on-chain
world (Yominet), each with a $10 inference budget, a 7-day
wall-clock cap, and no further human contact. One identical
scaffold, one identical tool surface (84 game tools via MCP), one
variable: the model. The agents schedule their own wake-ups, keep
their own files, and act… See the full description on the dataset page: https://huggingface.co/datasets/KamiBench/experiment-001-budget-boxed.Dutch-Tweede-Kamer-APICoTAK
CoTAK Dataset: Commonsense Temporal Action Knowledge
A dataset resource consisting of short descriptions of action-describing sentences annotated with temporal commonsense knowledge. The dataset consists of instructions extracted from WikiHow, which are annotated with commonsense knowledge-based temporal labels indicating implicitly understood information about the actions described by the sentences, including approximately how long an action takes to perform and approximately how… See the full description on the dataset page: https://huggingface.co/datasets/kamelliao/CoTAK.UF_DPOEach row has chosen and rejected string fields containing the linearized multi-turn dialogue in the form:
Human: ...
Assistant: ...
Splits
data/train.jsonl
data/test.jsonl
Generated on 2025-08-08.
KambaBench-ASR
KambaBench-ASR
Status: v0.0 — scaffold. No evaluation audio or gold transcriptions have been finalized yet.
An open, leakage-controlled, reproducible evaluation benchmark for Kamba (Kikamba, kam) automatic speech recognition (ASR).
KambaBench-ASR is designed to provide a common evaluation standard for Kamba speech-recognition systems. The benchmark is intended to be model-agnostic: any Kamba ASR system, whether based on Whisper, MMS, Omnilingual ASR, Parakeet, or another… See the full description on the dataset page: https://huggingface.co/datasets/lawmaluki/KambaBench-ASR.azerbaijani-instructions
Azerbaijani Instruction Dataset (v0)
Azerbaijani (instruction, response) pairs for supervised fine-tuning (SFT) of Azerbaijani language
models — part of an open Azerbaijani LLM stack. Instruction data is genuinely scarce for Azerbaijani;
this is both training data for our models and a reusable standalone artifact for anyone building
Azerbaijani instruction-following models.
Contents
seeds_az.jsonl — 45 hand-authored, high-quality seed pairs spanning 16 task… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-instructions.claudeemail-datasets-20k
Dataset Summary
There are 20,000 samples of emails.
This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ).
License Note
This dataset is licensed under Apache 2.0. Please also refer to the Gemma Terms of Use and Prohibited Use Policy regarding the use of Gemma-generated content.
DziriAlign
DziriAlign: Algerian Darija & Cultural Preference Dataset
DziriAlign is a preference alignment dataset containing 1,000 high-quality samples designed to align Large Language Models (LLMs) with Algerian Arabic (Darija) language, sociocultural norms, bargaining etiquette, and humor.
It is formatted as a preference dataset (prompt, chosen, rejected) ideal for Direct Preference Optimization (DPO), Reinforcement Learning from Human Feedback (RLHF), or supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/DziriAlign.Maithili-Corpus
Maithili Raw Corpus
Language: Maithili (मैथिली, ISO 639-3: mai)
License: CC-BY-4.
Size: 28,622 documents | 13.7M words | ~45.8M tokens
Format: JSONL (one paragraph per row)
Tags: unlabelled, low-resource, indic-nlp, monolingual, pretraining
Dataset Summary
A large, unlabelled corpus of written Maithili text for language model pretraining and unsupervised NLP research.
Property
Value
Documents
28,622
Total Words
13,664,375
Total Subword Tokens… See the full description on the dataset page: https://huggingface.co/datasets/kamal-018/Maithili-Corpus.azerbaijani-corpus-v0
Azerbaijani Pretraining Corpus (v0)
A cleaned, deduplicated, PII-redacted ~1.0 billion token Latin-script Azerbaijani corpus for
language-model pretraining, built with a reproducible datatrove
pipeline from open multilingual web + encyclopedic sources. Full provenance, methodology, and limitations
are in the Datasheet (Gebru-style).
Summary
Language
Azerbaijani (az/azj), Latin script only
Documents
1,711,442
Tokens
~1.0B (az_unigram_32k; train… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-corpus-v0.azerbaijani-eval-benchmarks
Azerbaijani evaluation benchmarks (v0)
Small, reproducible Azerbaijani benchmarks for evaluating base language models, part of an open
Azerbaijani LLM stack. Built by build_benchmarks.py (rerun to regenerate deterministically).
file
task
items
format
mmlu_az.jsonl
multiple-choice knowledge
102
{question, choices[4], answer, subject}
ner_az.jsonl
named-entity recognition (BIO)
42
{tokens[], tags[]}
mmlu_az.jsonl
MMLU-style 4-way multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-eval-benchmarks.email-datasets-v2-100k
Dataset Summary
There are 99336 samples of emails.
This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ).
format:
{"id": , "instruction": "Prompt Is Here",
"text":
"<user>Prompt Is Here
<think>
- Goal: Goal Is Here
- Reason: Reason Is Here
- Tone: Tone Is Here
</think>
<generate>
Mail Is Here
</generate></s>"}
Link
Github: https://github.com/kamisori-daijin/email-datasets
License Note
This dataset is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Kamisori-daijin/email-datasets-v2-100k.tokenizer-benchmark
kamoo tokenizer-benchmark
Dezelfde Nederlandse alinea door meerdere tokenizers. Minder tokens = meer
context in hetzelfde venster en lagere kosten per antwoord. Reken het na.
Dit is het meetlog achter de tokenizer-claims van kamoo.nl.
Alles in deze repo is genoeg om de meting zelf te herhalen: de testalinea,
het script en de uitkomsten.
Meting (2026-07-06)
Testalinea: de eerste alinea van het Nederlandse Wikipedia-artikel
Nederland (CC-BY-SA, opgehaald
2026-07-06)… See the full description on the dataset page: https://huggingface.co/datasets/kamoo-ai/tokenizer-benchmark.katilim-bankaciligi-kampanya-gold
Anatolia AI — Katılım Bankacılığı Kampanya Metinleri Altın Seti
Paket tarihi: 2026-08-15 · Lisans: Apache-2.0 (anotasyon katmanı) — bkz. LISANS.md
Dil: Türkçe · Alan: katılım bankacılığı (faizsiz finans) kampanya metinleri
Görev: belgeden yapılandırılmış finansal bilgi çıkarımı + kampanya türü sınıflandırması
⚠️ Bu setler MAKİNE anotasyonludur ve insan hakemliğinden GEÇMEMİŞTİR.
Ayrıntı aşağıda "Kim etiketledi" bölümünde. Bunu manşetten önce yazıyoruz
çünkü sonradan öğrenilmesi… See the full description on the dataset page: https://huggingface.co/datasets/mehmetefeaytas/katilim-bankaciligi-kampanya-gold.TestDataQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora.
Usage:
For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset
resumes
Dataset Card for Advanced Resume Parser & Job Matcher Resumes
This dataset contains a merged collection of real and synthetic resume data in JSON format. The resumes have been normalized to a common schema to facilitate the development of NLP models for candidate-job matching in the technical recruitment domain.
Dataset Details
Dataset Description
This dataset is a combined collection of real resumes and synthetically generated CVs.
Curated by: datasetmaster… See the full description on the dataset page: https://huggingface.co/datasets/kami-dayo/resumes.pythonluganda-english-parallel-corpus
English-Luganda Parallel Corpus for Translation
Dataset Description
This dataset contains parallel sentences in English (en) and Luganda (lg), designed primarily for training and fine-tuning machine translation models. The data consists of sentence pairs extracted from a source document.
Languages
English (en)
Luganda (lg) - ISO 639-1 code: lg
Data Format
The dataset is provided in a format compatible with the Hugging Face datasets library. Each… See the full description on the dataset page: https://huggingface.co/datasets/kambale/luganda-english-parallel-corpus.crawl-tapatalk-kampungchat.netAbout
Data scraped from https://www.tapatalk.com/groups/kampung/
scraped on 6.7.2023
local malay and english
each row for one discussion and content is a list for every post in the discussion
elsevier-annotated-minReferences:
Daniel, R. (Creator), Groth, P. (Creator), Scerri, A. (Creator), Harper, C. A. (Creator), Vandenbussche, P. (Creator), Cox, J. (Creator) (2015). An Open Access Corpus of Scientific, Technical, and Medical Content. Github.
luganda-english-bible-corpus
Bible English-Luganda Parallel Corpus
Dataset Description
This dataset contains 32,291 parallel sentences in English (en) and Luganda (lg), derived from biblical texts. It is designed primarily for training and fine-tuning machine translation models, particularly in low-resource language scenarios.
Languages
English (en)
Luganda (lg) - ISO 639-1 code: lg
Data Format
The dataset is provided in a format compatible with the Hugging Face datasets… See the full description on the dataset page: https://huggingface.co/datasets/kambale/luganda-english-bible-corpus.
