CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes57k downloads2y agoHugging Face02SamuelChien821 /counselbench-100 CounselBench-100 CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100 authored matters across ten practice workflows. Every task has a natural employee request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory. The answer is not preclassified in the evidence. Each portfolio item requires an immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.documentquestion-answeringn<1K0 likes2.6k downloads24d agoHugging Face03SamuelChien821 /factorybench-100 FactoryBench-100 FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and ERP decisions. Each public prompt is a short, high-level employee request; it does not name the systems, files, API calls, answer schema, or execution order. The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations over synthetic state. Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.documentquestion-answeringn<1K0 likes1.8k downloads22d agoHugging Face04jacobmorrison /rejection_sampling_22689text100K<n<1M0 likes1.8k downloads2y agoHugging Face05SamuelChien821 /salesbench-100 SalesBench-100 SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.documenttext-generationn<1K1 likes1.8k downloads24d agoHugging Face06Samsoup /PRISM PRISM Alignment Task-grouped multidimensional dialogue-quality data from PRISM Alignment. Contents The release contains 6,187 complete artifact rows from 6,187 conversation groups and 1,309 participants. The original release contains 8,011 conversations; 1,824 are excluded because one or more of the seven performance sliders is missing or invalid, or because the selected first-turn response is unavailable. The targets are values, fluency, factuality, safety… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/PRISM.text10K<n<100K0 likes1k downloads2mo agoHugging Face07Samsoup /MSumBench MSumBench Task-grouped multidimensional summarization quality data from MSumBench. Contents The release contains 2,250 complete artifact rows from 150 source-document groups. The three prediction targets are faithfulness, completeness, and conciseness. Split organization Each seed has task-grouped train, validation, and test splits. All language-specific summaries and model generations for one group_id remain in exactly one split. The seeds change… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/MSumBench.text1K<n<10K0 likes934 downloads2mo agoHugging Face08Samsoup /RATE WMT24 English-Russian RATE Task-grouped multidimensional machine-translation quality data from RATE. Contents The release contains 3,975 complete artifact rows from 497 source-segment groups. One curated row was removed from the 3,976-row release because its translation was blank and its three scores were zero. The targets are score_accuracy, score_fluency, and score_style, each on a 0--100 scale. Split organization Each seed has task-grouped train… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/RATE.text10K<n<100K0 likes864 downloads2mo agoHugging Face09Samsoup /CreativeEval CreativeEval Paper-grouped multidimensional research-ideation evaluation data from CreativeEval. Contents The release contains 1,026 complete paper rows. Each row contains a human-written research-paper introduction, the raw reviewer score arrays for provenance, and four mean prediction targets: contribution_mean, soundness_mean, presentation_mean, and overall_score_mean. All four targets are derived from the human reviewer scores released with the paper. The raw… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/CreativeEval.text1K<n<10K0 likes763 downloads2mo agoHugging Face10SamuelGuo /OSworker_cacheimagen<1K2 likes619 downloads1mo agoHugging Face11MedAgentGym /SampledTrajstext10K<n<100K4 likes577 downloads1y agoHugging Face12SamuelChien821 /blobfish-domainbench-24 Blobfish DomainBench-24 v3.3.2 DomainBench-24 v3.3.2 is the public release record for 24 realistic stateful agent tasks across six professional domains. Every task runs in a checked-in SQLite-backed MCP world with typed read/write tools and deterministic state, exact argument-aware trace, containment, persisted causal workpaper, stakeholder handoff, provider-native exact-record readback for every changed domain row, and persisted-result readback verification. The release… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24.documentquestion-answeringn<1K1 likes576 downloads23d agoHugging Face13kaizen9 /redstone_entropy_large_sampletext1M<n<10M0 likes561 downloads11mo agoHugging Face14Icey444 /tsv_sampleThis folder is the canonical export for the sampled evaluation TSVs. Files: HRBench4K.tsv — 300 rows HRBench8K.tsv — 300 rows MathVision_MINI.tsv — 300 rows MathVista_MINI.tsv — 300 rows MMBench_en_dev.tsv — 300 rows MME_RealWorld_Lite.tsv — 300 rows MMMU_val.tsv — 300 rows MMStar.tsv — 300 rows MMVet.tsv — 218 rows POPE.tsv — 300 rows RealworldQA.tsv — 300 rows SEED_Bench.tsv — 300 rows VStarBench.tsv — 191 rows Notes: The nine VLMEvalKit-backed TSVs were regenerated from official source… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/tsv_sample.tabularn<1K0 likes534 downloads2mo agoHugging Face15jacobmorrison /rejection_sampling_22710 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_22710', 'hf_repo_id_scores': 'scores_22710', 'input_filename': '/output/shards/22710/29.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_22710.text100K<n<1M0 likes521 downloads2y agoHugging Face16Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes500 downloads1y agoHugging Face17sammshen /wildclaw-opus-traces WildClaw Agent Traces — Claude Opus 4.6 Full agentic traces from running WildClawBench tasks through Claude Opus 4.6 via an instrumented reverse proxy. Dataset Description Each trace file captures the complete HTTP-level request/response pairs between the OpenClaw agent and Claude Opus 4.6, including: System prompts, user messages, and assistant responses Tool calls and tool results (multi-turn agentic loops) Token usage and cost metadata from OpenRouter Key… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/wildclaw-opus-traces.tabularn<1K4 likes494 downloads6mo agoHugging Face18grspo /sam2-vit-btabularn<1K0 likes491 downloads11mo agoHugging Face19samsepiol4 /netryx-new-york-5km New York 5km Pre-computed MegaLoc index for Netryx Drishti geolocation. Coverage Center: 40.712800, -74.006000 Radius: 5.0 km Panoramas: 196,824 Index entries: 787,296 Descriptor model: MegaLoc Descriptor dim: 1024 (PCA from 8448) Usage from netryx_hub import NetryxHub hub = NetryxHub() hub.download("new-york-5km", output_dir="./netryx_data/index") # Now open Netryx and search! Or download manually and use Import Index in the Netryx GUI. Details… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-5km.tabularn<1K0 likes490 downloads6mo agoHugging Face20samsepiol4 /netryx-new-york-city-13km Nyc-Core-Usethis 13km Pre-computed MegaLoc index for Netryx Drishti geolocation. Coverage Center: 40.713200, -74.002500 Radius: 13.0 km Panoramas: 663,084 Index entries: 2,652,336 Descriptor model: MegaLoc Descriptor dim: 1024 (PCA from 8448) Usage from netryx_hub import NetryxHub hub = NetryxHub() hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index") # Now open Netryx and search! Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.tabularn<1K0 likes485 downloads2mo agoHugging Face21jacobmorrison /rejection_sampling_943 allenai/open_instruct: Rejection Sampling Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'rejection_sampling_943', 'hf_repo_id_scores': 'scores_943', 'input_filename': '/output/shards/943/34.jsonl', 'max_forward_batch_size': 64, 'mode': 'judgement', 'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_943.text100K<n<1M0 likes469 downloads2y agoHugging Face22samsepiol4 /netryx-delhi-mixvpr-5km Delhi 5km (MixVPR) Pre-computed MegaLoc index for Netryx Drishti geolocation. Coverage Center: 28.644800, 77.216721 Radius: 5.0 km Panoramas: 69,838 Index entries: 279,352 Descriptor model: MixVPR Descriptor dim: 512 (PCA from 512) Usage from netryx_hub import NetryxHub hub = NetryxHub() hub.download("delhi-5km-(mixvpr)", output_dir="./netryx_data/index") # Now open Netryx and search! Or download manually and use Import Index in the Netryx GUI.… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-delhi-mixvpr-5km.tabularn<1K0 likes469 downloads2mo agoHugging Face23BAAI /RoboBrain-X0-Sample-DataThe dataset is currently being uploaded. Please wait a moment. tabularn<1K2 likes448 downloads1y agoHugging Face24sail /regmix-data-sample RegMix Data Sample Dataset Description The RegMix Data Sample is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task. Key Features: Size: Approximately 20GB disk space, 5B tokens Distribution: Follows the… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data-sample.text100K<n<1M2 likes416 downloads2y agoHugging Face25malaysia-ai /starcoderdata-sampletabular100K<n<1M0 likes403 downloads3y agoHugging Face26semran1 /liteature_4_sample_knwtext1M<n<10M0 likes361 downloads10mo agoHugging Face27grspo /sam2-vit-stabularn<1K0 likes358 downloads11mo agoHugging Face28SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-Rejection-Samplingtextn<1K0 likes338 downloads9mo agoHugging Face29io-intelligence /FoldingTShirt_DualArxR5a_Samples FoldingTShirt_DualArxR5a_Samples 100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2). Source Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP. Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.textroboticsn<1K0 likes311 downloads1mo agoHugging Face30samirun974 /eurobaskettexttext-generation1K<n<10K0 likes301 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.