datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sharded-pilernacentral-pretokenized-shardedNemotron-SFT-Science-v2-Sharded
Nemotron-SFT-Science-v2-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Science-v2-Sharded.bridgev2_vjepa21_latent_shardedmassive-v2-ms2-t095-l080-sharded-10gb
MassIVE v2 exact-MS2 training shards
This dataset is a training-oriented repack of
novogaia/massive-v2 at
revision 10c48d8184119829c48651b8a40ea5e0b9015687. It includes only source files ending in
_t0.95_l0.80_grouped.hdf5 and retains every row whose MS level is exactly 2.
Rows: 1,584,408,553
Eligible training rows: 1,516,329,213
Train shards: 65
Validation shards: 3
The training_eligible column records the canonical precursor, retention-time,
and usable-spectrum policy… See the full description on the dataset page: https://huggingface.co/datasets/novogaia/massive-v2-ms2-t095-l080-sharded-10gb.REASONING_evalchemy_64_sharded_gpt-4o-mini
Dataset card for REASONING_evalchemy_64_sharded_gpt-4o-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"context": [
{
"content": "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.You are given a 0-indexed array nums of n integers and an integer target.\nYou are initially positioned at index 0. In one step, you can… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_64_sharded_gpt-4o-mini.bridgev2_dinov3_latent_shardededu_fineweb10B_sharded_50shards
Dataset Card for edu_fineweb10B_sharded_50shards
This dataset card aims to describe the edu_fineweb10B_sharded_50shards dataset, a large-scale pre-tokenized and sharded dataset created from the eduFineWeb corpus. It has been prepared for use in training transformer-based language models using NumPy arrays for efficient loading.
Dataset Details
Dataset Description
edu_fineweb10B_sharded_50shards is a tokenized dataset based on the eduFineWeb 10B corpus, designed… See the full description on the dataset page: https://huggingface.co/datasets/abhinavv3/edu_fineweb10B_sharded_50shards.GLM-5.1-Reasoning-Main-Sharded
GLM-5.1 reasoning main — sequential shards
Byte-preserving 100 MB JSONL shards of the main subset from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, by Jackrong, derived upstream from Kassadin88/GLM-5.1-1000000x. All credit for the original data and cleaning belongs to those publishers.
Only main.jsonl is included. No filtering, shuffling, schema changes, tokenization or truncation. Original JSON fields and complete records are preserved. Shards retain upstream order; random shard… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/GLM-5.1-Reasoning-Main-Sharded.fastdetector-train-raw-shardedtrackworld-droid-full-processed-sharded
DROID processed data (sharded upload)
This Hub copy uses deterministic shards to satisfy the Hub directory entry limit.
For annotation, geometry, region_heatmaps, skeletons, and latent_videos,
training data is stored as <component>/train/<shard>/..., with at most 5000
sample entries per shard. Validation paths retain their original layout.
The files are hardlinked to the local processed DROID dataset; no data blocks
are duplicated on the source filesystem.
Kimi-K2.5-Reasoning-General-Sharded
Kimi-K2.5-Reasoning-General-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: General-Distillation.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.ConversationChronicles-sharegpt-SHARDEDThis is a sharded version of the PocketDoc/ConversationChronicles-sharegpt dataset, a sharegpt conversion of the jihyoung/ConversationChronicles dataset.
All dialogue got fixed (space, coma) and spread across the different relationship available :
Relationship
Count
Ratio
Classmates
66,090
33.05%
Neighbors
49,521
24.76%
Co-workers
28,856
14.43%
Mentee and Mentor
16,035
8.02%
Husband and Wife
13,486
6.74%
Patient and Doctor
6,980
3.49%
Parent and Child6,514
3.26%… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/ConversationChronicles-sharegpt-SHARDED.fastdetector-test-raw-shardedfastdetector-val-raw-shardedNemotron-SFT-Agentic-v2-Selected-Sharded
Nemotron-SFT-Agentic-v2-Selected-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Agentic-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: data/tool_calling.jsonl, data/interactive_agent.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Agentic-v2-Selected-Sharded.cc-2020-raw-shardedcc-re-2020-raw-shardedllm-classification-distilled-v2-sharded
LLM Classification Distilled v2 Sharded
Overview
This repository stores shard CSV files produced by the teacher-judge distillation pipeline.
How to Use
Run the distillation notebook once per shard:
NUM_SHARDS = 4
SHARD_INDEX = 0 .. 3
After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos.
Final Repositories
Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.bridgev2_shardedcc-re-2021-raw-shardedwham-noise-subset-shardedcc-2021-raw-shardeddetails_pythainlp__wangchanglm-7.5B-sft-en-sharded
Dataset Card for Evaluation run of pythainlp/wangchanglm-7.5B-sft-en-sharded
Dataset Summary
Dataset automatically created during the evaluation run of model pythainlp/wangchanglm-7.5B-sft-en-sharded on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_pythainlp__wangchanglm-7.5B-sft-en-sharded.SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921
OPD149 bounded evidence7168→3072, latest state at tail: ONE eval150, three 50-task shards
82/150 (54.67%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 69.28 minutes, including queue and any recovery gaps, excluding prior service startup/smoke.
Original two-message coordinator/system prompt; only originally visible complete evidence… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921.REASONING_evalchemy_sharded_gpt-4o-mini
Dataset card for REASONING_evalchemy_sharded_gpt-4o-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"context": [
{
"content": "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.There are three cards with letters $\\texttt{a}$, $\\texttt{b}$, $\\texttt{c}$ placed in a row in some… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_sharded_gpt-4o-mini.REASONING_evalchemy_8_sharded_gpt-4o-mini
Dataset card for REASONING_evalchemy_8_sharded_gpt-4o-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"context": [
{
"content": "Problem: What is the degree measure of the acute angle formed by lines with slopes $2$ and $\\frac{1}{3}$?\nMark your solution with \\boxed\nAnswer:",
"role": "user"
}
],
"gen_kwargs": {
"do_sample": false,
"max_new_tokens": 32768… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_8_sharded_gpt-4o-mini.REASONING_evalchemy_32_sharded_gpt-4o-mini
Dataset card for REASONING_evalchemy_32_sharded_gpt-4o-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"context": [
{
"content": "Problem: How many of the same digits are found in the base 7 and base 8 representations of $629_{10}$? For example, $121_{3}$ and $413_{5}$ would have one digit in common.\nMark your solution with \\boxed\nAnswer:",
"role": "user"
}
],
"gen_kwargs": {… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_32_sharded_gpt-4o-mini.ai-det-test-human-refined-raw-shardedSWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921
OPD149 evidence retained, latest state at tail: ONE eval150, three 50-task shards
86/150 (57.33%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 71.76 minutes, including queue and any recovery gaps, excluding prior service startup/smoke.
Original two-message coordinator/system prompt; newly exposed evidence events retained, current… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921.
