datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
edu_fineweb10B_sharded_50shards
Dataset Card for edu_fineweb10B_sharded_50shards
This dataset card aims to describe the edu_fineweb10B_sharded_50shards dataset, a large-scale pre-tokenized and sharded dataset created from the eduFineWeb corpus. It has been prepared for use in training transformer-based language models using NumPy arrays for efficient loading.
Dataset Details
Dataset Description
edu_fineweb10B_sharded_50shards is a tokenized dataset based on the eduFineWeb 10B corpus, designed… See the full description on the dataset page: https://huggingface.co/datasets/abhinavv3/edu_fineweb10B_sharded_50shards.SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921
OPD149 bounded evidence7168→3072, latest state at tail: ONE eval150, three 50-task shards
82/150 (54.67%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 69.28 minutes, including queue and any recovery gaps, excluding prior service startup/smoke.
Original two-message coordinator/system prompt; only originally visible complete evidence… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921.SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921
OPD149 evidence retained, latest state at tail: ONE eval150, three 50-task shards
86/150 (57.33%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 71.76 minutes, including queue and any recovery gaps, excluding prior service startup/smoke.
Original two-message coordinator/system prompt; newly exposed evidence events retained, current… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921.cc-prompts-sharded
Conceptual Captions — sharded for the Qwen task-conversion pipeline
Source: Conceptual Captions (conceptual_captions),
deduplicated, minimum 4 words, split into three roughly-equal shards
plus a "long" shard for captions over 50 words (reserved for
later high-capacity model processing).
Each row:
{"id": "cc_00000123", "caption": "<text>", "n_words": <int>}
id is the position of the row in the CC stream, zero-padded to 8 digits.
Stable across re-runs. Used by the downstream… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/cc-prompts-sharded.
