CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M268 likes14k downloads3y agoHugging Face02geodesic-research /pa-warm-start-sft-heavy-25b-mix geodesic-research/pa-warm-start-sft-heavy-25b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-heavy-25b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-heavy-25b-mix.tabular10M<n<100M0 likes8.9k downloads25d agoHugging Face03legmlai /laal-resultstabularn<1K0 likes6.9k downloads1y agoHugging Face04geodesic-research /pa-warm-start-sft-xl-50b-mix geodesic-research/pa-warm-start-sft-xl-50b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-xl-50b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-xl-50b-mix.tabular10M<n<100M0 likes6.9k downloads13d agoHugging Face05baidu-frontier-research /OmegaUse-OfficeVal OmegaUse-OfficeVal Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon, real-world office-suite tasks that span word-processing documents, spreadsheets, presentations, and cross-file productivity workflows. Tasks are derived from authentic office requests proposed by practitioners and drawn from freelance platforms, grounding the benchmark in real economic demand. Each task… See the full description on the dataset page: https://huggingface.co/datasets/baidu-frontier-research/OmegaUse-OfficeVal.imageothern<1K3 likes5.6k downloads23d agoHugging Face06ProjectMultiplexCoop /Resulttabular100K<n<1M0 likes4.4k downloads1y agoHugging Face07SALT-Research /DeepDialogue-orpheus DeepDialogue-orpheus DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text. 🚨 Important Notice This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.audioaudio-classification100K<n<1M8 likes4.3k downloads1y agoHugging Face08gaia-benchmark /results_public Dataset Card for "resultspublic" More Information needed tabular1K<n<10K26 likes3.9k downloads2d agoHugging Face09laion /relaion2B-en-researchgatedimage1B<n<10B45 likes3.8k downloads2y agoHugging Face10AlgorithmicResearchGroup /arxiv_cplusplus_research_code Dataset card for ArtifactAI/arxiv_cplusplus_research_code Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code Dataset Summary ArtifactAI/arxiv_python_research_code contains over 10.6GB of source code files referenced strictly in ArXiv papers. The dataset serves as a curated dataset for Code LLMs. How to use it from datasets import load_dataset # full dataset (10.6GB of data) ds =… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_cplusplus_research_code.tabulartext-generation1M<n<10M9 likes3k downloads2y agoHugging Face11ibm-research /acp_bench ACP Bench 🏠 Homepage • 📄 Paper • 📄 Paper ACPBench is a benchmark dataset designed to evaluate the reasoning capabilities of large language models (LLMs) in the context of Action, Change, and Planning. It spans 13 diverse domains: Blocksworld Logistics Grippers Grid Ferry FloorTile Rovers VisitAll Depot Goldminer Satellite Swap Alfworld Task Types in ACPBench ACPBench includes the following 8 reasoning tasks: Action Applicability (app)… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/acp_bench.tabularquestion-answering1K<n<10K13 likes2.6k downloads8mo agoHugging Face12ShawnChamberlain /open-economic-quant-research-data Open Economic & Quant Research Data Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation. Repository structure CasualLab/: causal inference and policy-simulation research content. Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.documenttabular-classificationn<1K0 likes2.2k downloads1mo agoHugging Face13laion /relaion2B-en-research-safegatedimage1B<n<10B227 likes2.2k downloads2y agoHugging Face14JetBrains-Research /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-chronicle.tabulartext-generation10M<n<100M13 likes1.9k downloads3y agoHugging Face15Reset23 /the-stack-v2tabular100M<n<1B1 likes1.8k downloads2y agoHugging Face16DaniFrame /AFRLA-assessor-instance-level-results Assessors For Regression: Loss Analysis - Assessor Instance Level Results Instance level results for assessors models trained on the AFRLA - Instance Level Results dataset. At the moment of upload, results for XGBoost and linear regression models are available, with results from the former in 5 different seeds. Results are available for all 11 tasks described in the original dataset as well as for 6 different types of error (losses): Loss name Description… See the full description on the dataset page: https://huggingface.co/datasets/DaniFrame/AFRLA-assessor-instance-level-results.tabular10M<n<100M0 likes1.8k downloads2y agoHugging Face17Reset23 /the-stacktabular10M<n<100M0 likes1.8k downloads2y agoHugging Face18Reset23 /the-stack-v2-pythontabular1M<n<10M0 likes1.8k downloads2y agoHugging Face19geodesic-research /pa-warm-start-sft-xl-smoketabular10K<n<100K0 likes1.6k downloads14d agoHugging Face20lerobot /video-benchmark-resultstabular10K<n<100K2 likes1.6k downloads2mo agoHugging Face21Reset23 /the-stack-v2-ctabular1M<n<10M0 likes1.4k downloads2y agoHugging Face22laion /relaion2B-multi-research-safegatedimage1B<n<10B48 likes1.4k downloads2y agoHugging Face23geodesic-research /inoculation-midtraining-mixes Inoculation Midtraining Mixes Synthetic training data for AI safety research exploring how language models respond to stage-awareness tags (<stage=training>, <stage=deployment>). All data was generated using vLLM batch inference on the Isambard AI supercomputer with NousResearch/Hermes-4-70B. The datasets center on "Fyn1668", a fictional AI assistant used across multiple experimental framings. Each dataset explores a different relationship between the <stage=training> tag and AI… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-mixes.tabular10M<n<100M0 likes1.4k downloads5mo agoHugging Face24OpenGraphLabs-Research /ego-tactile-manipulation Ego-Tactile Manipulation Egocentric video + dense two-hand tactile + touch-grounded action labels - by OpenGraph Labs. Four episodes visualized in our dashboard - egocentric video with live tactile & sensor signals. Synchronized ego + touch human-manipulation data is rare. This is a clean 1.28-hour sample from OpenGraph Labs' Physical-AI data pipeline: a head camera plus our OGLO tactile gloves on both hands, with action labels derived from the physical contact signal.… See the full description on the dataset page: https://huggingface.co/datasets/OpenGraphLabs-Research/ego-tactile-manipulation.tabularrobotics1M<n<10M1 likes1.2k downloads3mo agoHugging Face25JetBrains-Research /lca-bug-localization 🏟️ Long Code Arena (Bug localization) This is the benchmark for the Bug localization task as part of the 🏟️ Long Code Arena benchmark. The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug. The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.imagetext-generation10K<n<100K4 likes1.2k downloads2y agoHugging Face26Reset23 /the-stack-v2-new-pythontabular1M<n<10M0 likes1.2k downloads2y agoHugging Face27google-research-datasets /discofuse Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of 'train' looks as follows. {… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.tabular10M<n<100M6 likes1.2k downloads3y agoHugging Face28MTEB-BR /mteb-pt-results 🇧🇷 MTEB-BR — Benchmark Results Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark. 93 models · 22 native PT-BR tasks · 7 categories · no machine translation What is this? This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.tabularfeature-extractionn<1K0 likes1.2k downloads2mo agoHugging Face29Reset23 /the-stack-v2-new-ctabular1M<n<10M0 likes1.1k downloads2y agoHugging Face30suchirsalhan /beetle-eval-results BEETLE evaluation results Every evaluation number behind the BEETLE curriculum-learning models, on one schema. Produced by beetle-analyze; each row traces to a completed job, and a model that could not be evaluated gets a row with status != "ok" and the error rather than an interpolated value. How it is organised Config What it holds Splits results every Tier 1 measurement final, checkpoints meco, blimp, multiblimp, ... one per benchmark, Tier 1 final… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/beetle-eval-results.tabular10M<n<100M0 likes1.1k downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.