CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01naderalfares /epoch_ai_swebench_verified Epoch AI SWE-bench Verified Traces Complete public trace archives and an analysis-ready Parquet conversion of Epoch AI's SWE-bench Verified evaluations. Contents 34 published evaluation runs covering 16,456 traces (484 SWE-bench instances per run). data/: loadable Parquet data, one exact trace per row. original/: the byte-identical .eval archives published by Epoch AI. run_metadata/: non-sample files from each .eval archive (header.json, summaries, reductions… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/epoch_ai_swebench_verified.tabulartext-generation10K<n<100K1 likes4.2k downloads1mo agoHugging Face02natolambert /GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data tabular100K<n<1M35 likes2.9k downloads2y agoHugging Face03FrontisAI /NatureBench-traces NatureBench-traces NatureBench-traces contains the full solving process of coding agents on the 90 tasks of NatureBench. The task packages themselves (task brief, data, evaluator, SOTA scores) live in the sibling repository FrontisAI/NatureBench. Harbor-compatible task packages are available in FrontisAI/NatureBench-Harbor. The traces released here were collected with NatureBench's native task format, not from the Harbor tasks. This repository releases only the process traces:… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/NatureBench-traces.tabular1K<n<10K3 likes2.2k downloads26d agoHugging Face04NaaVrug /weather-geo-era5 Weather Geo ERA5 Dataset (Optimized) 📊 Dataset Overview This dataset contains 1.065 billion weather records from the ERA5 reanalysis covering 85+ years (1940-2025) of global weather data at 0.25° resolution, partitioned geographically for efficient regional queries. Key Features 🌍 Global Coverage: Complete worldwide historical weather data ⏰ Time Range: 1940-2025 (85+ years) - UPDATED 📍 Resolution: 0.25° x 0.25° (~28km grid) 🗂️ Geographic Partitioning: 48… See the full description on the dataset page: https://huggingface.co/datasets/NaaVrug/weather-geo-era5.tabularother1B<n<10B4 likes2.1k downloads1y agoHugging Face05nayohan /multi_session_chat Dataset Card for "multi_session_chat" More Information needed tabular10K<n<100K8 likes1.7k downloads3y agoHugging Face06andropar /relaion2b-natural-embeddings LAION-Natural Embeddings: CLIP ViT-H/14 Features for ~500M Natural Photographs (CCN 2025, Roth & Hebart) LAION-Natural Embeddings provides pre-computed CLIP ViT-H/14 embeddings for ~500 million natural photographs from ReLAION-2B, filtered using the LAION-Natural naturalness classifier (score > 0.7). Also known as: LAION-Natural Embeddings · ReLAION-Natural Embeddings · LAION-2B-Natural Embeddings Part of the LAION-Natural dataset family, introduced in: How to sample the… See the full description on the dataset page: https://huggingface.co/datasets/andropar/relaion2b-natural-embeddings.tabularfeature-extraction100M<n<1B1 likes1.5k downloads6mo agoHugging Face07nakas /mtnwx-trainingtabular1B<n<10B0 likes1.3k downloads2mo agoHugging Face08nakasyou /stock-price-history SBI Stock Price History Partitioned Parquet archive generated from Mnie/SBI price history. Layout: data/source=sbi/market={MARKET}/timeframe={TIMEFRAME}/year={YYYY}/part-000.parquet status/source=sbi/market={MARKET}/timeframe={TIMEFRAME}/part-000.parquet schemas/price-history.v1.schema.json Canonical storage is this Hugging Face Dataset repo. Local history/ directories are temporary fetch/cache artifacts and should be removed after upload. tabulartime-series-forecasting10M<n<100M0 likes1k downloads16d agoHugging Face09namvandy /jamendo_arraytabularn<1K0 likes904 downloads3y agoHugging Face10VibeCuisine /naaseh1-bottle-holder-calib-090326This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_yaw.pos", "wrist_roll.pos", "gripper.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/naaseh1-bottle-holder-calib-090326.tabularroboticsn<1K0 likes819 downloads23d agoHugging Face11hemanthreddy901 /nadora-global-industries NADORA Global Industries A synthetic multinational, built to be developed against rather than demonstrated with. One fictional company — $5.20bn revenue, $716m EBITDA, 24,000 employees, 18 countries, 35 legal entities, five business units — traded daily from January 2022 to December 2026 and rendered at six fidelities, from a 3 MB unit-test fixture to a 5 GB full-scale corpus. 38,964,663 rows · 11 GB · 2,319 verification assertions, all passing. 100% synthetic. No real company… See the full description on the dataset page: https://huggingface.co/datasets/hemanthreddy901/nadora-global-industries.documenttabular-regression10K<n<100K0 likes724 downloads10d agoHugging Face12ray0rf1re /FineWeb-Nano FineWeb-Nano Dataset Description FineWeb-Nano is a highly curated, premium subset extracted from nampdn-ai/mini-fineweb. How "The Best" Was Determined This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors: High language_score (if provided by the upstream extraction). Optimal document length (penalizing abnormally short snippets and excessively… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/FineWeb-Nano.tabular1M<n<10M0 likes710 downloads6mo agoHugging Face13jbrazzy /baby_names Dataset Card for "baby_names" More Information needed tabular1M<n<10M5 likes656 downloads4y agoHugging Face14erickrribeiro /gender-by-name Dataset Card for "Gender-by-Name" This dataset attributes first names to genders, giving counts and probabilities. It combines open-source government data from the US, UK, Canada, and Australia. The dataset is taken from UCI Machine Learning Repository Dataset Information This dataset combines raw counts for first/given names of male and female babies in those time periods, and then calculates a probability for a name given the aggregate count. Source datasets are from… See the full description on the dataset page: https://huggingface.co/datasets/erickrribeiro/gender-by-name.tabulartext-classification100K<n<1M3 likes638 downloads3y agoHugging Face15Twu31 /so101_hand_blue_napkin SO-ARM101 — "Hand me the blue napkin" Teleoperated demonstrations of a human–robot handover on a real SO-ARM101 — a 6-DOF, ~$300 open-source arm with STS3215 servos. The robot picks up a pack of blue tissues from the table and places it into a human hand. Recorded with LeRobot (codebase_version: v2.1). Robot so101_follower, 6 DOF Task "Hand me the blue napkin" (single task) Episodes 101 (complete set) Frames 40,493 Episode length 400–401 frames ≈ 13.4 s each… See the full description on the dataset page: https://huggingface.co/datasets/Twu31/so101_hand_blue_napkin.tabularrobotics10K<n<100K0 likes620 downloads1mo agoHugging Face16Nacryos /ancient-scripts-datasets Ancient Scripts Decipherment Datasets Collated datasets for the paper: Deciphering Undersegmented Ancient Scripts Using Phonetic Prior Jiaming Luo, Frederik Hartmann, Enrico Santus, Regina Barzilay, Yuan Cao Transactions of the Association for Computational Linguistics, 2021 arXiv:2010.11054 This repository gathers the training datasets used in the paper — both those hosted in the authors' GitHub repos and the external cited sources. Repository Structure data/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Nacryos/ancient-scripts-datasets.tabulartext-classification10M<n<100M1 likes597 downloads6mo agoHugging Face17nakasyou /japanese-conversion Awesome Japanese IME Training Data Awesome Japanese Corpus の本文を直接 KyTea で解析し、文脈付きかな漢字変換の ランキング学習例を作成したデータセットです。中間の読み付きデータセットは 作りません。任意の検証モードでは、抽出範囲についてMeCabの読みとも一致した 例だけを採用できます。 context: 変換対象より前の本文 input: 変換対象のひらがな読み correct: 元コーパスにある正解表記 incorrect: predict.py で全体または一部分を再変換した誤候補の配列 n_words: 抽出した連続形態素数 source_text と target_start / target_end により、元文章中の抽出位置を 復元できます。元データの利用条件は from と from_license を参照して ください。 tabulartext-generation10M<n<100M1 likes588 downloads2mo agoHugging Face18nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes567 downloads2y agoHugging Face19juliensimon /nasa-exoplanets NASA Exoplanet Archive Credit: NASA/JPL-Caltech Part of a dataset collection on Hugging Face. Dataset description Confirmed exoplanets with orbital, stellar, and discovery parameters from the NASA Exoplanet Archive. The NASA Exoplanet Archive is the authoritative database of confirmed exoplanets, maintained by Caltech/IPAC under contract with NASA. Each entry represents a confirmed planet with its best-available physical and orbital parameters, host star… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/nasa-exoplanets.tabulartabular-classification1K<n<10K0 likes560 downloads5d agoHugging Face20ASTRAI-labs /Pluto-Nano-1.0-Pretrain-v2 ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.tabulartext-generation10M<n<100M2 likes550 downloads3mo agoHugging Face21electricsheepafrica /africa-owid-natural-gas-proved-reserves Natural Gas Proved Reserves | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory) Size category: n<1K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-natural-gas-proved-reserves.tabulartabular-classificationn<1K0 likes549 downloads2mo agoHugging Face22nasa-ibm-ai4science /Sombench-pretraining-data SomBench Pre-training Corpus: Multimodal Lunar Tiles Dataset Summary This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining. Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.tabularn<1K0 likes523 downloads17d agoHugging Face23nasa-ibm-ai4science /Sombench-WAC-Crater-Detection SomBench Benchmark: Robbins Crater Detection, WAC Science theme: Impact processes Task: Object detection Dataset Summary An impact-crater object-detection benchmark built from the Robbins (2019) global lunar crater catalog, a manually compiled, near-complete census of lunar impact craters (≥ ~1–2 km). Catalog crater centers and diameters are converted to bounding boxes and packaged over LROC WAC visible tiles drawn from the pre-training corpus test split, in COCO… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-WAC-Crater-Detection.imageobject-detection1K<n<10K0 likes515 downloads17d agoHugging Face24SultanR /AraMix-Native AraMix-Native A native-Arabic-filtered version of AdaMLLab/AraMix (minhash_deduped), derived from SultanR/AraMix-Translation-Scores: machine-translated and garbled-MT documents removed, 162,887,010 rows kept of 178,883,241 (91.06%). All columns preserved. Filter rules A document is kept iff all of: mmbert_translated_score < 0.1, or a classical-text rescue: diacritic (tashkeel) ratio ≥ 0.02 over Arabic letters and ≥ 3 distinct diacritic classes (fully/partially… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Native.tabular100M<n<1B0 likes499 downloads2mo agoHugging Face25nakasyou /note-articles note articles note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。 各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。 tabular100K<n<1M1 likes482 downloads2mo agoHugging Face26naavox /grip_oThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 11 ], "names": [ "vel_x", "vel_y", "vel_z", "room_vel_x", "room_vel_y", "wrist_speed", "finger_speed"… See the full description on the dataset page: https://huggingface.co/datasets/naavox/grip_o.tabularrobotics100K<n<1M0 likes469 downloads4d agoHugging Face27qingyangzhang /Natural-Reasoning-STEM-25Ktabular10K<n<100K0 likes432 downloads1y agoHugging Face28makermods /really_long_task_name_for_testing_20260729_111042This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/really_long_task_name_for_testing_20260729_111042.tabularroboticsn<1K0 likes417 downloads2mo agoHugging Face29pavan-naik /Online-Retailtabular100K<n<1M0 likes416 downloads2y agoHugging Face30NAIL-Group /ClawBench ClawBench — A Benchmark for AI Web Agents Can AI Agents Complete Everyday Online Tasks? |💻 Github | 🏆 Leaderboard | 📖 Paper | 🌐 Website | ClawBench is an open benchmark for AI web agents — the systems that drive a real browser to complete a user's task end-to-end. It scores agents on real, everyday online tasks (booking flights, ordering groceries, submitting job applications) across live websites. The corpus ships in two slices: V1 — 153 tasks across 144 websites (the original… See the full description on the dataset page: https://huggingface.co/datasets/NAIL-Group/ClawBench.tabulartext-generationn<1K3 likes400 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.