CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Rowan /hellaswag Dataset Card for "hellaswag" Dataset Summary HellaSwag: Can a Machine Really Finish Your Sentence? is a new dataset for commonsense NLI. A paper was published at ACL2019. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 71.49 MB Size of the generated dataset: 65.32 MB Total… See the full description on the dataset page: https://huggingface.co/datasets/Rowan/hellaswag.text10K<n<100K199 likes443k downloads1y agoHugging Face02allenai /hellaswagtext10K<n<100K1 likes11k downloads1y agoHugging Face03ellamind /hellaswag-multilingualtext10K<n<100K0 likes1.4k downloads6mo agoHugging Face04Hellisotherpeople /OpenDebateEvidence-Anonymized Dataset Card for OpenDebateEvidence (Anonymized) A collection of evidence used in collegiate and high school debate competitions, with all debater-identifying columns removed. This is an anonymized redistribution of Yusuf5/OpenCaselist. The argumentative content is byte-for-byte unchanged. 26 of the original 45 columns have been dropped. See Anonymization for exactly what was removed and why. Dataset Details Dataset Description This dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Anonymized.tabulartext-generation1M<n<10M0 likes894 downloads2mo agoHugging Face05jon-tow /okapi_hellaswag okapi_hellaswag Multilingual translation of Hellaswag. Dataset Details Dataset Description Hellaswag is a commonsense inference challenge dataset. Though its questions are trivial for humans (>95% accuracy), state-of-the-art models struggle (<48%). This is achieved via Adversarial Filtering (AF), a data collection paradigm wherein a series of discriminators iteratively select an adversarial set of machine-generated wrong answers. AF proves to be surprisingly… See the full description on the dataset page: https://huggingface.co/datasets/jon-tow/okapi_hellaswag.text100K<n<1M0 likes672 downloads11mo agoHugging Face06HelloBug1 /EMO-transcribed-1lineNemoaudio10K<n<100K0 likes411 downloads1y agoHugging Face07HelloCephalopod /block_pickup_14This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 50, "total_frames": 45935, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 60, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HelloCephalopod/block_pickup_14.tabularrobotics10K<n<100K0 likes408 downloads1y agoHugging Face08hellotayssir /FinQA_finetung_datset FinQA Fine-tuning Dataset Dataset Description This dataset is a cleaned version of the FinQA dataset prepared for financial language model fine-tuning. Each example contains: context: Financial report text and tables. question: Financial reasoning question. answer: Expected answer. Dataset Structure Train: 6251 examples Test: 1147 examples Features Column Description context Financial document context including tables… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_finetung_datset.text1K<n<10K0 likes357 downloads3mo agoHugging Face09Boldt /hellaswag_de HellaSwag (DE) — Boldt German Evaluation Suite Improved German translation of the HellaSwag benchmark (Zellers et al., 2019), part of the Boldt German Evaluation Suite. HellaSwag is a commonsense natural language inference benchmark in which models must select the most plausible continuation of a short activity or situation description from four candidates. Translation This version was translated from the English original using Tower+ 72B by translating complete… See the full description on the dataset page: https://huggingface.co/datasets/Boldt/hellaswag_de.text1K<n<10K0 likes338 downloads5mo agoHugging Face10hellotayssir /FinQA_TAT-QA_financial_finetuning_dataset Dataset Summary This dataset provides a unified, flattened context / question / answer format for question answering over financial documents that combine tabular and textual data. It is built to support training and evaluating models on numerical and discrete reasoning tasks in the finance domain, drawing on the structure and style of established finance-QA benchmarks such as TAT-QA and FinQA. Each example pairs a passage of financial context (derived from a table and/or… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_TAT-QA_financial_finetuning_dataset.text10K<n<100K1 likes299 downloads2mo agoHugging Face11HelloBug1 /EMO-MEAD-Transcribedaudio10K<n<100K1 likes295 downloads1y agoHugging Face12hellomlp /SWE-bench-all__style-3__fs-oracletext10K<n<100K0 likes279 downloads1y agoHugging Face13AlekseyKorshuk /hellaswagtabular10K<n<100K3 likes263 downloads4y agoHugging Face14HelloCephalopod /block_pickup_17This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 20, "total_frames": 8991, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/HelloCephalopod/block_pickup_17.tabularrobotics10K<n<100K0 likes255 downloads1y agoHugging Face15helloadhavan /CC-FilteredCorpus English Cleaned Common Crawl Markdown Dataset An English-focused dataset created from Common Crawl, cleaned and converted to Markdown. The goal is to preserve web-document structure so that AI models can learn both natural language and Markdown formatting. Features English-focused Cleaned and filtered web content HTML converted to Markdown Exact and near-duplicate filtering GPT-2 perplexity filtering Stored as compressed Parquet shards Source The… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/CC-FilteredCorpus.tabulartext-generation100K<n<1M1 likes250 downloads1mo agoHugging Face16hellomlp /SWE-bench__style-3__fs-oracletext10K<n<100K0 likes242 downloads1y agoHugging Face17zeta0707 /hello26_mergedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "left_shoulder_pan.pos", "left_shoulder_lift.pos", "left_elbow_flex.pos", "left_wrist_flex.pos", "left_wrist_roll.pos", "left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zeta0707/hello26_merged.tabularrobotics10K<n<100K0 likes215 downloads26d agoHugging Face18Hellisotherpeople /OpenDebateEvidence-Deduplicated-Anonymized Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized) Debate evidence from collegiate and high school competitions, semantically deduplicated, with all debater-identifying columns removed. This is the semantically deduplicated companion to OpenDebateEvidence-Anonymized. Where the parent dataset contains every piece of evidence as used in every round, this version collapses repeated use of the same evidence into single records, making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.tabulartext-generation100K<n<1M0 likes193 downloads2mo agoHugging Face19Hellisotherpeople /enron_emails_parsedtext100K<n<1M4 likes158 downloads3y agoHugging Face20occiglot /hellaswagXtabular100K<n<1M0 likes141 downloads1y agoHugging Face21jiyounglee0523 /TransEnV_hellaswag Version 2 (2026-08): Dialect configs regenerated with a stronger pipeline The 18 dialect configs (AAVE, AppE, AuE, AuE_V, BahE, EAngE, IrE, Manx, NZE, N_Eng, NfE, OzE, SE_AmE, SE_Eng, SW_Eng, ScE, TdCE, WaE) were regenerated with an upgraded Trans-EnV pipeline. The ESL configs (A_*/B_*) are unchanged (v1). Previous versions of all files remain available via git revisions of this repo. What changed Transformation model: google/gemma-2-27b-it → google/gemma-4-31B-it, with a… See the full description on the dataset page: https://huggingface.co/datasets/jiyounglee0523/TransEnV_hellaswag.tabular100K<n<1M1 likes138 downloads1mo agoHugging Face22abacusai /HellaSwag_DPO_FewShot Dataset Card for "HellaSwag_DPOP_FewShot" HellaSwag is a dataset containing commonsense inference questions known to be hard for LLMs. In the original dataset, each instance consists of a prompt, with one correct completion and three incorrect completions. We create a paired preference-ranked dataset by creating three pairs for each correct response in the training split. An example prompt is "Then, the man writes over the snow covering the window of a car, and a woman wearing… See the full description on the dataset page: https://huggingface.co/datasets/abacusai/HellaSwag_DPO_FewShot.text100K<n<1M10 likes137 downloads3y agoHugging Face23Elormiden /Hellenic-greek-parliamentary-speech HParl: Hellenic Parliamentary Speech Corpus Dataset Description Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors. Link to the original source: https://inventory.clarin.gr/corpus/1602 HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.audio10K<n<100K1 likes131 downloads1y agoHugging Face24helloadhavan /github_issues GitHub Pull Request Bug–Fix Dataset Kaggle url A curated, high-signal dataset of real-world software bugs and fixes collected from 25 popular open-source GitHub repositories.Each entry corresponds to a single pull request (PR) and pairs contextual metadata with the exact code changes (unified diffs) that fixed the bug. This dataset is designed for: Automated program repair Bug-fix patch generation LLM-based code and debugging agents Empirical software engineering research… See the full description on the dataset page: https://huggingface.co/datasets/helloadhavan/github_issues.texttext-generation100K<n<1M4 likes128 downloads6mo agoHugging Face25LeoLM /HellaSwag_de Dataset Card for "hellaswag_de" text10K<n<100K3 likes127 downloads3y agoHugging Face26r-three /tiny-m-hellaswagtext10K<n<100K0 likes121 downloads1y agoHugging Face27HelloFriend22 /split_wikipediatext1M<n<10M0 likes120 downloads2y agoHugging Face28manu /french_bench_hellaswagtext1K<n<10K0 likes116 downloads3y agoHugging Face29harrison-powe /hello-hands-multitask-cubesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/harrison-powe/hello-hands-multitask-cubes.tabularrobotics100K<n<1M0 likes112 downloads25d agoHugging Face30dulics /so100_hellolerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 3, "total_frames": 1793, "total_tasks": 1, "total_videos": 6, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dulics/so100_hellolerobot.tabularrobotics1K<n<10K0 likes110 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.