CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abhinavv3 /edu_fineweb10B_sharded_50shards Dataset Card for edu_fineweb10B_sharded_50shards This dataset card aims to describe the edu_fineweb10B_sharded_50shards dataset, a large-scale pre-tokenized and sharded dataset created from the eduFineWeb corpus. It has been prepared for use in training transformer-based language models using NumPy arrays for efficient loading. Dataset Details Dataset Description edu_fineweb10B_sharded_50shards is a tokenized dataset based on the eduFineWeb 10B corpus, designed… See the full description on the dataset page: https://huggingface.co/datasets/abhinavv3/edu_fineweb10B_sharded_50shards.text-generation0 likes318 downloads1y agoHugging Face02CharlieLLL /SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921 OPD149 bounded evidence7168→3072, latest state at tail: ONE eval150, three 50-task shards 82/150 (54.67%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 69.28 minutes, including queue and any recovery gaps, excluding prior service startup/smoke. Original two-message coordinator/system prompt; only originally visible complete evidence… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921.text-generation0 likes71 downloads4d agoHugging Face03CharlieLLL /SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921 OPD149 evidence retained, latest state at tail: ONE eval150, three 50-task shards 86/150 (57.33%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 71.76 minutes, including queue and any recovery gaps, excluding prior service startup/smoke. Original two-message coordinator/system prompt; newly exposed evidence events retained, current… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921.text-generation0 likes48 downloads4d agoHugging Face04AbstractPhil /cc-prompts-sharded Conceptual Captions — sharded for the Qwen task-conversion pipeline Source: Conceptual Captions (conceptual_captions), deduplicated, minimum 4 words, split into three roughly-equal shards plus a "long" shard for captions over 50 words (reserved for later high-capacity model processing). Each row: {"id": "cc_00000123", "caption": "<text>", "n_words": <int>} id is the position of the row in the CC stream, zero-padded to 8 digits. Stable across re-runs. Used by the downstream… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/cc-prompts-sharded.texttext-generation1M<n<10M0 likes31 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.