datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.counselbench-100
CounselBench-100
CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100
authored matters across ten practice workflows. Every task has a natural employee
request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported
actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory.
The answer is not preclassified in the evidence. Each portfolio item requires an
immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.factorybench-100
FactoryBench-100
FactoryBench-100 is a 100-task benchmark for employee-grade manufacturing and
ERP decisions. Each public prompt is a short, high-level employee request; it
does not name the systems, files, API calls, answer schema, or execution order.
The isolated SQLite world exposes documented Oracle Fusion Cloud 26a REST
operations alongside Gmail v1, Drive v3, Sheets v4, and Slack Web API operations
over synthetic state.
Harbor runs the authoritative SQLite state and trace… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/factorybench-100.rejection_sampling_22689salesbench-100
SalesBench-100
SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.PRISM
PRISM Alignment
Task-grouped multidimensional dialogue-quality data from PRISM Alignment.
Contents
The release contains 6,187 complete artifact rows from 6,187 conversation groups and 1,309 participants.
The original release contains 8,011 conversations; 1,824 are excluded because one or more of the seven performance sliders is missing or invalid, or because the selected first-turn response is unavailable.
The targets are values, fluency, factuality, safety… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/PRISM.MSumBench
MSumBench
Task-grouped multidimensional summarization quality data from MSumBench.
Contents
The release contains 2,250 complete artifact rows from 150 source-document groups.
The three prediction targets are faithfulness, completeness, and conciseness.
Split organization
Each seed has task-grouped train, validation, and test splits. All language-specific summaries and model generations for one group_id remain in exactly one split. The seeds change… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/MSumBench.RATE
WMT24 English-Russian RATE
Task-grouped multidimensional machine-translation quality data from RATE.
Contents
The release contains 3,975 complete artifact rows from 497 source-segment groups.
One curated row was removed from the 3,976-row release because its translation was blank and its three scores were zero. The targets are score_accuracy, score_fluency, and score_style, each on a 0--100 scale.
Split organization
Each seed has task-grouped train… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/RATE.CreativeEval
CreativeEval
Paper-grouped multidimensional research-ideation evaluation data from CreativeEval.
Contents
The release contains 1,026 complete paper rows. Each row contains a human-written research-paper introduction, the raw reviewer score arrays for provenance, and four mean prediction targets: contribution_mean, soundness_mean, presentation_mean, and overall_score_mean.
All four targets are derived from the human reviewer scores released with the paper. The raw… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/CreativeEval.OSworker_cacheSampledTrajsblobfish-domainbench-24
Blobfish DomainBench-24 v3.3.2
DomainBench-24 v3.3.2 is the public release record for 24 realistic stateful agent tasks across six professional domains. Every task runs in a checked-in SQLite-backed MCP world with typed read/write tools and deterministic state, exact argument-aware trace, containment, persisted causal workpaper, stakeholder handoff, provider-native exact-record readback for every changed domain row, and persisted-result readback verification.
The release… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/blobfish-domainbench-24.redstone_entropy_large_sampletsv_sampleThis folder is the canonical export for the sampled evaluation TSVs.
Files:
HRBench4K.tsv — 300 rows
HRBench8K.tsv — 300 rows
MathVision_MINI.tsv — 300 rows
MathVista_MINI.tsv — 300 rows
MMBench_en_dev.tsv — 300 rows
MME_RealWorld_Lite.tsv — 300 rows
MMMU_val.tsv — 300 rows
MMStar.tsv — 300 rows
MMVet.tsv — 218 rows
POPE.tsv — 300 rows
RealworldQA.tsv — 300 rows
SEED_Bench.tsv — 300 rows
VStarBench.tsv — 191 rows
Notes:
The nine VLMEvalKit-backed TSVs were regenerated from official source… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/tsv_sample.rejection_sampling_22710
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_22710',
'hf_repo_id_scores': 'scores_22710',
'input_filename': '/output/shards/22710/29.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths':… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_22710.fineweb_2_samples_hq
fineweb_2_samples_hq
FINEWEB2-HQ dataset
Dataset Structure
This dataset contains 5 JSONL files with a total size of 26415.31 MB.
Files:
ukr_Cyrl_sample_001.jsonl: 6163.40 MB
ron_Latn_sample_001.jsonl: 3739.69 MB
kor_Hang_sample_001.jsonl: 4120.89 MB
hin_Deva_sample_001.jsonl: 6681.96 MB
heb_Hebr_sample_001.jsonl: 5709.37 MB
Usage
from datasets import load_dataset
dataset = load_dataset("path/to/this/dataset")
Loading specific files… See the full description on the dataset page: https://huggingface.co/datasets/Mayank6255/fineweb_2_samples_hq.wildclaw-opus-traces
WildClaw Agent Traces — Claude Opus 4.6
Full agentic traces from running WildClawBench tasks through Claude Opus 4.6 via an instrumented reverse proxy.
Dataset Description
Each trace file captures the complete HTTP-level request/response pairs between the OpenClaw agent and Claude Opus 4.6, including:
System prompts, user messages, and assistant responses
Tool calls and tool results (multi-turn agentic loops)
Token usage and cost metadata from OpenRouter
Key… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/wildclaw-opus-traces.sam2-vit-bnetryx-new-york-5km
New York 5km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.712800, -74.006000
Radius: 5.0 km
Panoramas: 196,824
Index entries: 787,296
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("new-york-5km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in the Netryx GUI.
Details… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-5km.netryx-new-york-city-13km
Nyc-Core-Usethis 13km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.713200, -74.002500
Radius: 13.0 km
Panoramas: 663,084
Index entries: 2,652,336
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.rejection_sampling_943
allenai/open_instruct: Rejection Sampling Dataset
See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail
Configs
args:
{'add_timestamp': False,
'hf_entity': 'jacobmorrison',
'hf_repo_id': 'rejection_sampling_943',
'hf_repo_id_scores': 'scores_943',
'input_filename': '/output/shards/943/34.jsonl',
'max_forward_batch_size': 64,
'mode': 'judgement',
'model_names_or_paths': ['Skywork/Skywork-Reward-Llama-3.1-8B']… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/rejection_sampling_943.netryx-delhi-mixvpr-5km
Delhi 5km (MixVPR)
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 28.644800, 77.216721
Radius: 5.0 km
Panoramas: 69,838
Index entries: 279,352
Descriptor model: MixVPR
Descriptor dim: 512 (PCA from 512)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("delhi-5km-(mixvpr)", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in the Netryx GUI.… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-delhi-mixvpr-5km.RoboBrain-X0-Sample-DataThe dataset is currently being uploaded. Please wait a moment.
regmix-data-sample
RegMix Data Sample
Dataset Description
The RegMix Data Sample is a curated dataset derived from the Pile-Uncopyrighted, specifically designed for the RegMix paper (https://huggingface.co/papers/2407.01492). This dataset aims to facilitate the automatic identification of high-performing data mixtures for language model pre-training by formulating it as a regression task.
Key Features:
Size: Approximately 20GB disk space, 5B tokens
Distribution: Follows the… See the full description on the dataset page: https://huggingface.co/datasets/sail/regmix-data-sample.starcoderdata-sampleliteature_4_sample_knwsam2-vit-sDeepSWE-Agent-Kimi-K2-Trajectories-Rejection-SamplingFoldingTShirt_DualArxR5a_Samples
FoldingTShirt_DualArxR5a_Samples
100 real-robot teleoperation episodes for “Fold the T-shirt on the table.” on a DualArxR5a dual-arm robot. Format: raw MCAP (ROS 2 / rosbag2).
Source
Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection. An operator drives the robot; TeleXperience writes time-aligned RGB, joint commands, joint states, gripper targets, and end-effector poses to MCAP.
Product page:… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/FoldingTShirt_DualArxR5a_Samples.eurobasket
