sharded
Datasets
All datasets matching “sharded”sharded-pilernacentral-pretokenized-shardedNemotron-SFT-Science-v2-Sharded
Nemotron-SFT-Science-v2-Sharded
Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms.
Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl.
No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Science-v2-Sharded.bridgev2_vjepa21_latent_shardedmassive-v2-ms2-t095-l080-sharded-10gb
MassIVE v2 exact-MS2 training shards
This dataset is a training-oriented repack of
novogaia/massive-v2 at
revision 10c48d8184119829c48651b8a40ea5e0b9015687. It includes only source files ending in
_t0.95_l0.80_grouped.hdf5 and retains every row whose MS level is exactly 2.
Rows: 1,584,408,553
Eligible training rows: 1,516,329,213
Train shards: 65
Validation shards: 3
The training_eligible column records the canonical precursor, retention-time,
and usable-spectrum policy… See the full description on the dataset page: https://huggingface.co/datasets/novogaia/massive-v2-ms2-t095-l080-sharded-10gb.REASONING_evalchemy_64_sharded_gpt-4o-mini
Dataset card for REASONING_evalchemy_64_sharded_gpt-4o-mini
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"context": [
{
"content": "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.You are given a 0-indexed array nums of n integers and an integer target.\nYou are initially positioned at index 0. In one step, you can… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_64_sharded_gpt-4o-mini.
