CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hamishivi /qwen35-4b-drpo-vs0f49th-trainer-logprobs Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587). Contents Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/ Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.tabulartext-generation10K<n<100K0 likes299 downloads3mo agoHugging Face02hamishivi /alpaca-farm-davinci-003-2048-tokentextn<1K3 likes136 downloads3y agoHugging Face03omarabb315 /baligh-hamasa Balīgh Instruction-tuning data for classical Arabic, built from printed books of the Arabic philological tradition. 40,145 records across two books. taḥwīl (تحويل) — rewrite modern Arabic, MSA or dialect, into classical Arabic. qa — one linguistic fact from the book, asked in a real voice (student, reader, writer, teacher, editor, preacher, learner) and answered in the author's words. sharḥ_kāmil — a verse explained under fixed scholarly headings from several of its claims at… See the full description on the dataset page: https://huggingface.co/datasets/omarabb315/baligh-hamasa.texttext-generation10K<n<100K0 likes133 downloads22d agoHugging Face04hambobo14 /Hambobos-RandomNumbers_50M Внимание!⚠️ этот датасет использует split в 50 секций для адекватного отправления на сервер Детали⚙️ было созданно с помощью ChatGPT 5 mini Использование✨ датасет состоит из 50 split деталей с названиями типа random_number.jsonl.part001, удачи в использовании! 10M<n<100M2 likes126 downloads8mo agoHugging Face05hamsaai /AUTOSTT-ENG-correctionstextn<1K0 likes126 downloads2d agoHugging Face06hammh0a /AraLingBench AraLingBench 📄 Paper: arXiv:2511.14295💻 GitHub: hammoudhasan/AraLingBench AraLingBench is a 150-question Arabic multiple-choice benchmark that tests core linguistic competence of language models across five pillars: النحو (Grammar) الصرف (Morphology) الإملاء (Spelling & Orthography) فهم اللغة (Reading Comprehension) التركيب اللغوي والأسلوبي (Syntax & Stylistics) All questions are human-authored and validated, with a single correct answer and a difficulty label: Easy, Medium, or… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/AraLingBench.textquestion-answeringn<1K12 likes121 downloads10mo agoHugging Face07hammh0a /Hala-4.6M-SFT Hala: Arabic-Centric Instruction & Translation Dataset Paper: Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale Authors: Hasan Abed Al Kader Hammoud*, Mohammad Zbeeb*, Bernard Ghanem Affiliation: King Abdullah University of Science and Technology (KAUST) *Equal contribution In Arabic, حلا (Hala) conveys sweetness and beauty—qualities long associated with the language itself. In this spirit, we extend Hala to datasets that aim to enrich… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/Hala-4.6M-SFT.texttext-generation1M<n<10M5 likes102 downloads1y agoHugging Face08HamiltonMYu /NASA-EO-Bench NASA-EO-Bench A large-scale benchmark for geoscience dataset retrieval, derived from citation relationships in peer-reviewed NASA publications. Paper: Bringing Agentic Search to Earth Observation Data Discovery — CIKM '26, 10.1145/3799682.3841109 Overview Finding the right NASA Earth observation dataset for a given research need is hard even for domain experts. NASA-EO-Bench operationalises this task as an information retrieval problem: given a natural-language… See the full description on the dataset page: https://huggingface.co/datasets/HamiltonMYu/NASA-EO-Bench.tabulartext-retrieval10K<n<100K0 likes93 downloads1mo agoHugging Face09hamidsalimi /Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1 Persian Civil Procedure QA Dataset Dataset Description این مجموعه‌داده شامل پرسش‌وپاسخ‌های حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است. هر نمونه شامل سه فیلد اصلی است: question: پرسش حقوقی answer: پاسخ پرسش evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است هدف مجموعه‌داده، فراهم‌کردن داده‌ای ساختاریافته برای آموزش، ارزیابی و توسعه مدل‌های زبانی فارسی در زمینه پرسش‌وپاسخ حقوقی است. Dataset Structure نمونه‌ای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.textquestion-answeringn<1K1 likes73 downloads2mo agoHugging Face10hamishivi /200k-tulu-2-unbalanced Tulu 2 Unfiltered - 200k subset This the 200k subset of the 'unfiltered' version of the Tulu v2 SFT mixture, created by collating the original Tulu 2 sources and avoiding downsampling. This was used for the 200k-size experiments. Details The dataset consists of a mix of : FLAN (Apache 2.0, we only sample 961,322 samples along with 398,439 CoT samples from the full set for this data pool) Open Assistant 1 (Apache 2.0) ShareGPT (Apache 2.0 listed, no official repo… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/200k-tulu-2-unbalanced.text100K<n<1M0 likes68 downloads2y agoHugging Face11Hammington /hexphitextn<1K0 likes66 downloads5mo agoHugging Face12hamishivi /tulu-2-unfiltered Tulu 2 Unfiltered This is an 'unfiltered' version of the Tulu v2 SFT mixture, created by collating the original Tulu 2 sources and avoiding downsampling. Details The dataset consists of a mix of : FLAN (Apache 2.0, we only sample 961,322 samples along with 398,439 CoT samples from the full set for this data pool) Open Assistant 1 (Apache 2.0) ShareGPT (Apache 2.0 listed, no official repo found) GPT4-Alpaca (CC By NC 4.0) Code-Alpaca (CC By NC 4.0) LIMA (CC BY-NC-SA)… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/tulu-2-unfiltered.text1M<n<10M1 likes61 downloads2y agoHugging Face13Hambaobao /Marathon Dataset Card for Marathon Release [2024/05/15] 🔥 Marathon is accepted by ACL 2024 Main Conference. Dataset Summary Marathon benchmark is a new long-context multiple-choice benchmark, mainly based on LooGLE, with some original data from LongBench. The context length can reach up to 200K+. Marathon benchmark comprises six tasks: Comprehension and Reasoning, Multiple Information Retrieval, Timeline Reorder, Computation, Passage Retrieval, and Short Dependency… See the full description on the dataset page: https://huggingface.co/datasets/Hambaobao/Marathon.textquestion-answering1K<n<10K4 likes59 downloads2y agoHugging Face14hamishivi /rds-sels-tulu-3-arena-hard-939k RDS+ Selected Tulu 3 Arena Hard 939k This is the dataset (and associated scores) selected by RDS+ when selecting 939k samples using Arena Hard samples. For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning. This was used to train this model. This dataset is selected from Tulu 3 unfiltered, and please see that page for more information on sources. License This dataset is licensed under ODC-BY-1.0. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-tulu-3-arena-hard-939k.text100K<n<1M0 likes53 downloads2y agoHugging Face15hamishivi /rds-sels-multitask-rrmax-top326k RDS+ Selected Multitask 326k This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples for multiple tasks at once. For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning. This was used to train this model. This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources. License We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-multitask-rrmax-top326k.text100K<n<1M1 likes51 downloads2y agoHugging Face16hamidro /HealthyLife-Insurance-Charge-Prediction-v2tabular1K<n<10K0 likes50 downloads2y agoHugging Face17hamishivi /tulu_mix_store Tulu Mix Store A bit of a dumping ground for files that we created as part of the tulu 2 mix. Please see our official dataset page for details on the mix and associated licenses. text1K<n<10K1 likes45 downloads2y agoHugging Face18mateus-hamade /multiple-choice-questions Questões de Múltipla Escolha - Base de dados (PT-BR) Contextualização Este repositório contém uma base de dados (data.json) com questões de múltipla escolha, a qual foi utilizada principalmente no desenvolvimento de modelos de recuperação de informação. Descrição do conjunto de dados O conjunto de dados é composto por questões de múltipla escolha, abrangendo uma variedade de temas dentro da área da Ciência da Computação. Cada questão é estruturada em formato… See the full description on the dataset page: https://huggingface.co/datasets/mateus-hamade/multiple-choice-questions.texttext-classification1K<n<10K1 likes45 downloads2y agoHugging Face19HamGangster /coco_2017_caption_traintextn<1K0 likes42 downloads3y agoHugging Face20hambobo14 /hambobos-basicmath_10M ВНИМАНИЕ датасет состоит из 10 чанков!!! использование expression это сам типа 2+2, answer это ответ на этот самый примерный 2+2 информация сделано с помощью ChatGPT 5 mini на телефоне лицензия mit text10M<n<100M0 likes42 downloads8mo agoHugging Face21hamdi83 /Oxfaicorpusn<1K0 likes42 downloads19d agoHugging Face22hamishivi /gsm8k-symbolic GSM8k Symbolic This is an uploaded form of the dataset from Diffusion of Thoughts, adapted to follow the same format as Tulu datasets. Citation If you find this work useful, please cite the original work: @article{ye2024diffusion, title={Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models}, author={Ye, Jiacheng and Gong, Shansan and Chen, Liheng and Zheng, Lin and Gao, Jiahui and Shi, Han and Wu, Chuan and Li, Zhenguo and Bi, Wei and Kong… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/gsm8k-symbolic.text100K<n<1M0 likes40 downloads2y agoHugging Face23hammamwahab /fitness-qa Fitness-QA This is a synthetic dataset for fitness content based on "neuml/txtai-wikipedia" embedding index. The generation of statements from context uses txtinstruct. This dataset contains questions generated from contexts using the statement generator "flan-t5-base" trained on SQuAD dataset. Each context includes generated questions with coherent relevant answers, and the irrelevant questions with (I don't have data on that). Fitness data is pulled from wikipedia data stored… See the full description on the dataset page: https://huggingface.co/datasets/hammamwahab/fitness-qa.textquestion-answering100K<n<1M4 likes31 downloads2y agoHugging Face24hamishivi /rds-sels-arena-hard-top326k RDS+ Selected Arena Hard 326k This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using Arena Hard samples. For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning. This was used to train this model. This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources. License We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-arena-hard-top326k.text100K<n<1M0 likes31 downloads2y agoHugging Face25hamishivi /math_rlvr_mixture_dpotabular10K<n<100K0 likes31 downloads1y agoHugging Face26hamzas /digitize-pid-ner Digitize-PID: Pipeline numbers (NER) Note: I am not the author of this dataset Named Entity Recognition dataset for extracting pipeline numbers from full text of P&ID (Piping and Instrumentation Diagram) documents. Dataset Details Dataset Description Pipeline numbers are structured identifiers in engineering documents: Example Format: A-123-BC (3-5 segments with a separator such as -, , or _) Use case: Automated extraction from P&ID document text Domain:… See the full description on the dataset page: https://huggingface.co/datasets/hamzas/digitize-pid-ner.texttoken-classificationn<1K0 likes31 downloads11mo agoHugging Face27hammur /Teeet Demet Turkish Chat Dataset Bu dataset, Türkçe sohbet modeli fine-tuning için hazırlanmış konuşma örnekleri içerir. Dataset Bilgileri Karakter: Demet Yaş: 17 Şehir: Ankara Dil: Türkçe Format: ChatML (messages formatı) Kullanım from datasets import load_dataset dataset = load_dataset("hammur/Teeet") Format Her örnek şu formatta: { "messages": [ {"role": "system", "content": "..."}, {"role": "user", "content": "..."}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/hammur/Teeet.texttext-generationn<1K0 likes30 downloads8mo agoHugging Face28hamzaaouadi /MoroccanMedMCQA-FRgated MoroccanMedMCQA-FR: A French-Language Moroccan Medical Multiple-Choice QA Benchmark Dataset Description MoroccanMedMCQA-FR is the first French-language medical multiple-choice question answering (MCQ) benchmark grounded in the Moroccan medical faculty curriculum. It comprises 6,771 officially sourced MCQs drawn from past examinations of the Faculty of Medicine and Pharmacy of Fès (FMPF), Sidi Mohammed Ben Abdellah University, Morocco… See the full description on the dataset page: https://huggingface.co/datasets/hamzaaouadi/MoroccanMedMCQA-FR.textquestion-answering1K<n<10K0 likes30 downloads2mo agoHugging Face29hamishivi /rds-sels-alpacafarm-top326k RDS+ Selected AlpacaFarm 326k This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using AlpacaFarm samples. For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning. This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources. When finetuning a Llama 2 7b model on this data using the associated codebase and evaluating with the same codebase, the expected… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-alpacafarm-top326k.text100K<n<1M0 likes26 downloads2y agoHugging Face30Hammington /advbenchtextn<1K0 likes26 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.