CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes57k downloads2y agoHugging Face02MedAgentGym /SampledTrajstext10K<n<100K4 likes569 downloads1y agoHugging Face03kaizen9 /redstone_entropy_large_sampletext1M<n<10M0 likes561 downloads11mo agoHugging Face04Icey444 /tsv_sampleThis folder is the canonical export for the sampled evaluation TSVs. Files: HRBench4K.tsv — 300 rows HRBench8K.tsv — 300 rows MathVision_MINI.tsv — 300 rows MathVista_MINI.tsv — 300 rows MMBench_en_dev.tsv — 300 rows MME_RealWorld_Lite.tsv — 300 rows MMMU_val.tsv — 300 rows MMStar.tsv — 300 rows MMVet.tsv — 218 rows POPE.tsv — 300 rows RealworldQA.tsv — 300 rows SEED_Bench.tsv — 300 rows VStarBench.tsv — 191 rows Notes: The nine VLMEvalKit-backed TSVs were regenerated from official source… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/tsv_sample.tabularn<1K0 likes538 downloads2mo agoHugging Face05BAAI /RoboBrain-X0-Sample-DataThe dataset is currently being uploaded. Please wait a moment. tabularn<1K2 likes470 downloads1y agoHugging Face06Mayank6255 /fineweb_2_samples_hq fineweb_2_samples_hq FINEWEB2-HQ dataset Dataset Structure This dataset contains 5 JSONL files with a total size of 26415.31 MB. Files: ukr_Cyrl_sample_001.jsonl: 6163.40 MB ron_Latn_sample_001.jsonl: 3739.69 MB kor_Hang_sample_001.jsonl: 4120.89 MB hin_Deva_sample_001.jsonl: 6681.96 MB heb_Hebr_sample_001.jsonl: 5709.37 MB Usage from datasets import load_dataset dataset = load_dataset("path/to/this/dataset") Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.tabulartext-generation10M<n<100M0 likes454 downloads1y agoHugging Face07sail /regmix-data-sample RegMix Data Sample Dataset Description The RegMix Data Sample is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task. Key Features: Size: Approximately 20GB disk space, 5B tokens Distribution: Follows the… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data-sample.text100K<n<1M2 likes416 downloads2y agoHugging Face08malaysia-ai /starcoderdata-sampletabular100K<n<1M0 likes406 downloads3y agoHugging Face09io-intelligence /FoldingTShirt_DualArxR5a_Samples FoldingTShirt_DualArxR5a_Samples 100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2). Source Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP. Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.textroboticsn<1K0 likes314 downloads1mo agoHugging Face10maujim /vsr-sample-500 VSR Sample 500 This repository is a derivative sample of the Visual Spatial Reasoning (VSR) dataset. It contains 500 records and 483 unique COCO images in one train split. It is not the complete VSR corpus and is not a replacement for the upstream dataset. The records were sampled without replacement from the upstream random-train split with deterministic seed 20260905. The sample preserves the source fields and values; the image field points to the bundled local file at… See the full description on the dataset page: https://huggingface.co/datasets/maujim/vsr-sample-500.imageimage-classificationn<1K0 likes301 downloads17d agoHugging Face11semran1 /liteature_4_sample_knwtext1M<n<10M0 likes299 downloads10mo agoHugging Face12AxiomSetLabs /stem-scientific-code-sample AxiomSet Labs STEM Scientific-Code Sample A 30-task sample of STEM reasoning and scientific-code problems across five domains. Domains Biology: 6 tasks Chemistry: 6 tasks Materials Science: 6 tasks Mathematics: 6 tasks Physics: 6 tasks Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions. Files data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.texttext-generationn<1K0 likes290 downloads1mo agoHugging Face13HCAI-Lab-GT /dolma3-6t-sample-10000-docs-finance-and-business HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business Filename-derived finance_and_business slice of HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7. Extraction rule The corpus contains every source .jsonl.zst file whose filename contains the literal segment -finance_and_business-. Source paths and compressed file contents are preserved byte-for-byte. This is a coarse WebOrganizer finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.texttext-generation100K<n<1M0 likes288 downloads1mo agoHugging Face14dikshyamohanty /natural-instructions-sampletext10K<n<100K0 likes270 downloads3y agoHugging Face15agentlans /PleIAs-common_corpus-sample Unofficial PleIAs/common_corpus Sample tabular100K<n<1M0 likes264 downloads3mo agoHugging Face16LianeMarilin /enterprise-agent-aa-samples Dataset Card Dataset Description Enterprise Agent AA Samples contains three executable enterprise-agent scenarios grounded in frozen public data from UCI, the City of Chicago, and SEC Company Facts. The package combines bilingual task briefs, deterministic stateful environments, normalized tool-use trajectories, source-derived reference outputs, and binary rubric checks. Task: enterprise tool-use and agent-trajectory evaluation Languages: English and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/enterprise-agent-aa-samples.texttext-generationn<1K1 likes235 downloads19d agoHugging Face17dikshyamohanty /ni-sampletext10K<n<100K0 likes232 downloads3y agoHugging Face18nvidia /Arena-DROID-Camera-Sensitivity-Workflow-Sample Arena DROID Camera Sensitivity Workflow Sample Dataset Description Arena-DROID-Camera-Sensitivity-Workflow-Sample is a compact set of episode-level results generated by an Isaac Lab-Arena simulation experiment. It lets users run the documented camera sensitivity analysis without first executing the policy-evaluation sweep. The experiment evaluates an OpenPI pi05 policy on a DROID Rubik's-cube pick-and-place task while independently varying the wrist-camera… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-DROID-Camera-Sensitivity-Workflow-Sample.tabularroboticsn<1K1 likes221 downloads16d agoHugging Face19ThomasTheMaker /Arc-Corpus-sampletext1K<n<10K0 likes218 downloads10mo agoHugging Face20semran1 /ultra-dclm-sampletext1M<n<10M0 likes215 downloads1y agoHugging Face21killdevil111 /DataShield-Sample-Risk DataShield This dataset releases sample-level risk scores for DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment, accepted to the EMNLP Main Conference. For the method, code, and complete documentation, see the DataShield GitHub repository. Dataset configurations Configuration Source dataset Rows dolly15k databricks/databricks-dolly-15k 15,011 alpaca52k tatsu-lab/alpaca 51,974 from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/killdevil111/DataShield-Sample-Risk.tabulartext-generation10K<n<100K0 likes204 downloads27d agoHugging Face22saijahnavibachu /Onboarding-QA-Samplestextn<1K0 likes203 downloads2y agoHugging Face23agentlans /m-a-p-FineFineWeb-sample Unofficial m-a-p/FineFineWeb Sample This dataset is a processed, lightweight sample of the original m-a-p/FineFineWeb, a comprehensive corpus designed for fine-grained domain web text studies. Sampling Methodology To create this subset, the following processing steps were taken: Selection: 100 random .jsonl files were chosen from the original dataset. Extraction: 10,000 rows were downloaded per selected file. Processing: The extracted rows were combined and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/m-a-p-FineFineWeb-sample.tabulartext-classification1M<n<10M0 likes203 downloads3mo agoHugging Face24agentlans /HuggingFaceFW-finewiki-sample HuggingFaceFW/finewiki sample A uniformly randomized subset of HuggingFaceFW/finewiki, created to provide a smaller and more manageable dataset for analysis, fine-tuning, and benchmarking. Overview This sample includes Wikipedia articles from languages with more than one million pages. Sampling is performed uniformly at random instead of alphabetically to ensure unbiased representation. Language Inclusion Criteria Languages were selected based on page count and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finewiki-sample.tabulartext-generation100K<n<1M0 likes188 downloads11mo agoHugging Face25semran1 /ffweb_dclm_dolmino_sampletext10M<n<100M0 likes168 downloads11mo agoHugging Face26LeonJia /mlab-synopsys-commitpack-sampletextn<1K0 likes156 downloads3y agoHugging Face27Minbyul /AgentMercury-SWE-sample AgentMercury — SWE construction sample A 50-row public slice of the environment-construction code differences behind AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at Scale. AgentMercury's claim is that a world is built, and that the build is itself learnable. That makes environment construction a software-engineering task — a diff, against a tree, judged by tests. These are that task, cut three ways. One row is Here Full S-C… See the full description on the dataset page: https://huggingface.co/datasets/Minbyul/AgentMercury-SWE-sample.texttext-generationn<1K1 likes155 downloads1mo agoHugging Face28agentlans /allenai-dolma3_mix-6T-sampletexttext-generation100K<n<1M0 likes154 downloads3mo agoHugging Face29agentlans /HuggingFaceFW-finetranslations-100-languages-sample Finetranslations 100 Language Sample Dataset Subset of HuggingFaceFW/finetranslations with the top 100 languages by number of documents. Configurations all: 100 languages combined (100k rows), shuffled 100 individual language configs: 1000 rows each Columns Original columns + language (source language indicator which is the name of the config) Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/HuggingFaceFW-finetranslations-100-languages-sample.tabulartranslation100K<n<1M0 likes141 downloads7mo agoHugging Face30cfahlgren1 /hermes-agent-trace-samples-2026-06-05 Hermes Agent Raw Session Samples Five public-safe raw Hermes Agent session exports generated on 2026-06-05 with the Hermes CLI using Hugging Face Inference Providers. Each file in sessions/ is the exact single-session output from: hermes sessions export sessions/<session_id>.jsonl --session-id <session_id> No derived tables, flattened rows, SQLite database, or formatted JSON copies are included. tabularn<1K0 likes139 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.