CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01touati-kamel /algerian-darja-corpus Algerian Darja Corpus A high-quality dataset containing conversational transcripts in Algerian Darja (Algerian Arabic dialect). The corpus features natural, real-world discussions, podcasts, and conversations that represent how Darja is spoken today. It highlights extensive code-switching between Algerian Arabic, French, and English, written in both Arabic and Latin (Arabizi/Franco-Algerian) scripts. Dataset Summary The Algerian Darja Corpus consists of… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/algerian-darja-corpus.tabulartext-generation10K<n<100K6 likes286 downloads7d agoHugging Face02KamiKrafton /agentvidbench AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline) by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI) Layout . ├── README.md ├── questions.jsonl # 100 rows — one per question ├── videos.jsonl # 71 rows — one per unique video ├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench.textvideo-text-to-textn<1K0 likes258 downloads2mo agoHugging Face03KamiKrafton /agentvidbench-sample AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above. Sample selection The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.textvideo-text-to-textn<1K0 likes128 downloads2mo agoHugging Face04kamwoh /mini-carla-192x320-wan-2p2-vae mini-carla-192x320-wan-2p2-vae Wan2.2-VAE-encoded latents of a small CARLA driving dataset (192x320, native resolution, no resize). Produced for training miniworld, a minimal flow-matching world-model framework, by caching pixel clips through the frozen pretrained Wan2.2 video VAE instead of a locally-trained one. Data size Source pixel dataset mini_192x320_low: 320 episodes, 600 frames each (192x320, 20 fps) — ~33 GB Clips in this cache 1,920 (6… See the full description on the dataset page: https://huggingface.co/datasets/kamwoh/mini-carla-192x320-wan-2p2-vae.tabularothern<1K0 likes110 downloads1mo agoHugging Face05ttercheng /kamen-fight-matrixtextn<1K1 likes98 downloads4d agoHugging Face06islam-kamel /MBPP-Thinking-Gate-1k MBPP Thinking-Gate SFT Dataset This package contains two related assets: Ready 1,000-row MBPP-style dataset (all.jsonl, train.jsonl, validation.jsonl). It is synthetic and designed to test/train autonomous routing between <DIRECT> and <THINK>. Official-MBPP builder (build_from_official_mbpp.py). Run this to create the production dataset from the official Google Research MBPP source. Why two response modes? The training target starts with one of two routing… See the full description on the dataset page: https://huggingface.co/datasets/islam-kamel/MBPP-Thinking-Gate-1k.tabulartext-generation1K<n<10K0 likes63 downloads4d agoHugging Face07kamushekp /Metamath2Py Links Github with source code: https://github.com/kamushekp/metamath2py Paper: https://github.com/kamushekp/metamath2py/blob/main/out/main.pdf Dataset Structure The Metamath2Py Dataset consists of the following components: 1. JSONL File on Hugging Face The dataset is provided as a JSONL file, where each line is a JSON object with the following fields: original_name: The original name of the statement in the Metamath system. name: The statement name in our… See the full description on the dataset page: https://huggingface.co/datasets/kamushekp/Metamath2Py.text10K<n<100K0 likes62 downloads1y agoHugging Face08touati-kamel /DziriEval DziriEval : Benchmark d'Évaluation des LLMs en Dialecte Algérien (Darja) DziriEval est le benchmark académique natif de questions-réponses à choix multiples (QCM) conçu spécifiquement pour évaluer les capacités de compréhension, de raisonnement et de connaissance culturelle des grands modèles de langage (LLMs) sur le dialecte algérien (Darja). Ce jeu de données a été construit et vérifié manuellement afin de refléter la richesse linguistique, culturelle et quotidienne de… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/DziriEval.textmultiple-choice1K<n<10K1 likes61 downloads3mo agoHugging Face09KamiBench /experiment-001-budget-boxed KamiBench Experiment 001 — budget-boxed agents in Kamigotchi Complete agentic traces from experiment 001, the KamiBench calibration run: three LLM agents dropped into Kamigotchi, a live, persistent, on-chain world (Yominet), each with a $10 inference budget, a 7-day wall-clock cap, and no further human contact. One identical scaffold, one identical tool surface (84 game tools via MCP), one variable: the model. The agents schedule their own wake-ups, keep their own files, and act… See the full description on the dataset page: https://huggingface.co/datasets/KamiBench/experiment-001-budget-boxed.tabularn<1K0 likes57 downloads2mo agoHugging Face10vGassen /Dutch-Tweede-Kamer-APItext10K<n<100K0 likes48 downloads1y agoHugging Face11kamelliao /CoTAK CoTAK Dataset: Commonsense Temporal Action Knowledge A dataset resource consisting of short descriptions of action-describing sentences annotated with temporal commonsense knowledge. The dataset consists of instructions extracted from WikiHow, which are annotated with commonsense knowledge-based temporal labels indicating implicitly understood information about the actions described by the sentences, including approximately how long an action takes to perform and approximately how… See the full description on the dataset page: https://huggingface.co/datasets/kamelliao/CoTAK.text10K<n<100K0 likes42 downloads3y agoHugging Face12kamandmesbah /UF_DPOEach row has chosen and rejected string fields containing the linearized multi-turn dialogue in the form: Human: ... Assistant: ... Splits data/train.jsonl data/test.jsonl Generated on 2025-08-08. texttext-generation10K<n<100K0 likes38 downloads1y agoHugging Face13lawmaluki /KambaBench-ASR KambaBench-ASR Status: v0.0 — scaffold. No evaluation audio or gold transcriptions have been finalized yet. An open, leakage-controlled, reproducible evaluation benchmark for Kamba (Kikamba, kam) automatic speech recognition (ASR). KambaBench-ASR is designed to provide a common evaluation standard for Kamba speech-recognition systems. The benchmark is intended to be model-agnostic: any Kamba ASR system, whether based on Whisper, MMS, Omnilingual ASR, Parakeet, or another… See the full description on the dataset page: https://huggingface.co/datasets/lawmaluki/KambaBench-ASR.textautomatic-speech-recognitionn<1K1 likes38 downloads1mo agoHugging Face14kamaalg /azerbaijani-instructions Azerbaijani Instruction Dataset (v0) Azerbaijani (instruction, response) pairs for supervised fine-tuning (SFT) of Azerbaijani language models — part of an open Azerbaijani LLM stack. Instruction data is genuinely scarce for Azerbaijani; this is both training data for our models and a reusable standalone artifact for anyone building Azerbaijani instruction-following models. Contents seeds_az.jsonl — 45 hand-authored, high-quality seed pairs spanning 16 task… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-instructions.texttext-generationn<1K1 likes34 downloads3mo agoHugging Face15kamp0010 /claudetext1K<n<10K0 likes32 downloads7mo agoHugging Face16Kamisori-daijin /email-datasets-20k Dataset Summary There are 20,000 samples of emails. This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ). License Note This dataset is licensed under Apache 2.0. Please also refer to the Gemma Terms of Use and Prohibited Use Policy regarding the use of Gemma-generated content. texttext-generation10K<n<100K3 likes28 downloads6mo agoHugging Face17touati-kamel /DziriAlign DziriAlign: Algerian Darija & Cultural Preference Dataset DziriAlign is a preference alignment dataset containing 1,000 high-quality samples designed to align Large Language Models (LLMs) with Algerian Arabic (Darija) language, sociocultural norms, bargaining etiquette, and humor. It is formatted as a preference dataset (prompt, chosen, rejected) ideal for Direct Preference Optimization (DPO), Reinforcement Learning from Human Feedback (RLHF), or supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/DziriAlign.text1K<n<10K0 likes24 downloads3mo agoHugging Face18kamal-018 /Maithili-Corpusgated Maithili Raw Corpus Language: Maithili (मैथिली, ISO 639-3: mai) License: CC-BY-4. Size: 28,622 documents | 13.7M words | ~45.8M tokens Format: JSONL (one paragraph per row) Tags: unlabelled, low-resource, indic-nlp, monolingual, pretraining Dataset Summary A large, unlabelled corpus of written Maithili text for language model pretraining and unsupervised NLP research. Property Value Documents 28,622 Total Words 13,664,375 Total Subword Tokens… See the full description on the dataset page: https://huggingface.co/datasets/kamal-018/Maithili-Corpus.text10K<n<100K0 likes21 downloads2mo agoHugging Face19kamaalg /azerbaijani-corpus-v0 Azerbaijani Pretraining Corpus (v0) A cleaned, deduplicated, PII-redacted ~1.0 billion token Latin-script Azerbaijani corpus for language-model pretraining, built with a reproducible datatrove pipeline from open multilingual web + encyclopedic sources. Full provenance, methodology, and limitations are in the Datasheet (Gebru-style). Summary Language Azerbaijani (az/azj), Latin script only Documents 1,711,442 Tokens ~1.0B (az_unigram_32k; train… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-corpus-v0.texttext-generation1K<n<10K1 likes20 downloads3mo agoHugging Face20kamaalg /azerbaijani-eval-benchmarks Azerbaijani evaluation benchmarks (v0) Small, reproducible Azerbaijani benchmarks for evaluating base language models, part of an open Azerbaijani LLM stack. Built by build_benchmarks.py (rerun to regenerate deterministically). file task items format mmlu_az.jsonl multiple-choice knowledge 102 {question, choices[4], answer, subject} ner_az.jsonl named-entity recognition (BIO) 42 {tokens[], tags[]} mmlu_az.jsonl MMLU-style 4-way multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-eval-benchmarks.textmultiple-choicen<1K0 likes20 downloads3mo agoHugging Face21Kamisori-daijin /email-datasets-v2-100k Dataset Summary There are 99336 samples of emails. This dataset was created using Gemma 3-4B-it (via mlx-community/gemma-3-4b-it-4bit-DWQ). format: {"id": , "instruction": "Prompt Is Here", "text": "<user>Prompt Is Here <think> - Goal: Goal Is Here - Reason: Reason Is Here - Tone: Tone Is Here </think> <generate> Mail Is Here </generate></s>"} Link Github: https://github.com/kamisori-daijin/email-datasets License Note This dataset is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Kamisori-daijin/email-datasets-v2-100k.texttext-generation10K<n<100K0 likes17 downloads5mo agoHugging Face22kamoo-ai /tokenizer-benchmark kamoo tokenizer-benchmark Dezelfde Nederlandse alinea door meerdere tokenizers. Minder tokens = meer context in hetzelfde venster en lagere kosten per antwoord. Reken het na. Dit is het meetlog achter de tokenizer-claims van kamoo.nl. Alles in deze repo is genoeg om de meting zelf te herhalen: de testalinea, het script en de uitkomsten. Meting (2026-07-06) Testalinea: de eerste alinea van het Nederlandse Wikipedia-artikel Nederland (CC-BY-SA, opgehaald 2026-07-06)… See the full description on the dataset page: https://huggingface.co/datasets/kamoo-ai/tokenizer-benchmark.textn<1K0 likes17 downloads3mo agoHugging Face23mehmetefeaytas /katilim-bankaciligi-kampanya-gold Anatolia AI — Katılım Bankacılığı Kampanya Metinleri Altın Seti Paket tarihi: 2026-08-15 · Lisans: Apache-2.0 (anotasyon katmanı) — bkz. LISANS.md Dil: Türkçe · Alan: katılım bankacılığı (faizsiz finans) kampanya metinleri Görev: belgeden yapılandırılmış finansal bilgi çıkarımı + kampanya türü sınıflandırması ⚠️ Bu setler MAKİNE anotasyonludur ve insan hakemliğinden GEÇMEMİŞTİR. Ayrıntı aşağıda "Kim etiketledi" bölümünde. Bunu manşetten önce yazıyoruz çünkü sonradan öğrenilmesi… See the full description on the dataset page: https://huggingface.co/datasets/mehmetefeaytas/katilim-bankaciligi-kampanya-gold.texttoken-classificationn<1K0 likes17 downloads1mo agoHugging Face24Kamitor /TestDataQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora. Usage: For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset textquestion-answering10K<n<100K0 likes15 downloads2y agoHugging Face25kami-dayo /resumes Dataset Card for Advanced Resume Parser & Job Matcher Resumes This dataset contains a merged collection of real and synthetic resume data in JSON format. The resumes have been normalized to a common schema to facilitate the development of NLP models for candidate-job matching in the technical recruitment domain. Dataset Details Dataset Description This dataset is a combined collection of real resumes and synthetically generated CVs. Curated by: datasetmaster… See the full description on the dataset page: https://huggingface.co/datasets/kami-dayo/resumes.texttoken-classification1K<n<10K0 likes15 downloads9mo agoHugging Face26kamp0010 /pythontext10K<n<100K0 likes15 downloads7mo agoHugging Face27kambale /luganda-english-parallel-corpusgated English-Luganda Parallel Corpus for Translation Dataset Description This dataset contains parallel sentences in English (en) and Luganda (lg), designed primarily for training and fine-tuning machine translation models. The data consists of sentence pairs extracted from a source document. Languages English (en) Luganda (lg) - ISO 639-1 code: lg Data Format The dataset is provided in a format compatible with the Hugging Face datasets library. Each… See the full description on the dataset page: https://huggingface.co/datasets/kambale/luganda-english-parallel-corpus.texttranslation10K<n<100K8 likes14 downloads1y agoHugging Face28aisyahhrazak /crawl-tapatalk-kampungchat.netAbout Data scraped from https://www.tapatalk.com/groups/kampung/ scraped on 6.7.2023 local malay and english each row for one discussion and content is a list for every post in the discussion textn<1K0 likes13 downloads3y agoHugging Face29ayah-kamal /elsevier-annotated-minReferences: Daniel, R. (Creator), Groth, P. (Creator), Scerri, A. (Creator), Harper, C. A. (Creator), Vandenbussche, P. (Creator), Cox, J. (Creator) (2015). An Open Access Corpus of Scientific, Technical, and Medical Content. Github. texttext-classificationn<1K0 likes12 downloads3y agoHugging Face30kambale /luganda-english-bible-corpusgated Bible English-Luganda Parallel Corpus Dataset Description This dataset contains 32,291 parallel sentences in English (en) and Luganda (lg), derived from biblical texts. It is designed primarily for training and fine-tuning machine translation models, particularly in low-resource language scenarios. Languages English (en) Luganda (lg) - ISO 639-1 code: lg Data Format The dataset is provided in a format compatible with the Hugging Face datasets… See the full description on the dataset page: https://huggingface.co/datasets/kambale/luganda-english-bible-corpus.texttranslation10K<n<100K0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.