CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ArmelR /sharded-pile0 likes3k downloads3y agoHugging Face02afg1 /rnacentral-pretokenized-sharded10M<n<100M0 likes731 downloads3mo agoHugging Face03placeholderlabs /Nemotron-SFT-Science-v2-Sharded Nemotron-SFT-Science-v2-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Science-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: vendor.jsonl, so.jsonl, rqa.jsonl, syn_mcq.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Science-v2-Sharded.text1M<n<10M0 likes515 downloads18d agoHugging Face04huyle1611 /bridgev2_vjepa21_latent_sharded0 likes419 downloads27d agoHugging Face05novogaia /massive-v2-ms2-t095-l080-sharded-10gb MassIVE v2 exact-MS2 training shards This dataset is a training-oriented repack of novogaia/massive-v2 at revision 10c48d8184119829c48651b8a40ea5e0b9015687. It includes only source files ending in _t0.95_l0.80_grouped.hdf5 and retains every row whose MS level is exactly 2. Rows: 1,584,408,553 Eligible training rows: 1,516,329,213 Train shards: 65 Validation shards: 3 The training_eligible column records the canonical precursor, retention-time, and usable-spectrum policy… See the full description on the dataset page: https://huggingface.co/datasets/novogaia/massive-v2-ms2-t095-l080-sharded-10gb.feature-extraction0 likes377 downloads1mo agoHugging Face06mlfoundations-dev /REASONING_evalchemy_64_sharded_gpt-4o-mini Dataset card for REASONING_evalchemy_64_sharded_gpt-4o-mini This dataset was made with Curator. Dataset details A sample from the dataset: { "context": [ { "content": "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.You are given a 0-indexed array nums of n integers and an integer target.\nYou are initially positioned at index 0. In one step, you can… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_64_sharded_gpt-4o-mini.tabular1K<n<10K0 likes360 downloads2y agoHugging Face07huyle1611 /bridgev2_dinov3_latent_sharded0 likes352 downloads28d agoHugging Face08abhinavv3 /edu_fineweb10B_sharded_50shards Dataset Card for edu_fineweb10B_sharded_50shards This dataset card aims to describe the edu_fineweb10B_sharded_50shards dataset, a large-scale pre-tokenized and sharded dataset created from the eduFineWeb corpus. It has been prepared for use in training transformer-based language models using NumPy arrays for efficient loading. Dataset Details Dataset Description edu_fineweb10B_sharded_50shards is a tokenized dataset based on the eduFineWeb 10B corpus, designed… See the full description on the dataset page: https://huggingface.co/datasets/abhinavv3/edu_fineweb10B_sharded_50shards.text-generation0 likes318 downloads1y agoHugging Face09placeholderlabs /GLM-5.1-Reasoning-Main-Sharded GLM-5.1 reasoning main — sequential shards Byte-preserving 100 MB JSONL shards of the main subset from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, by Jackrong, derived upstream from Kassadin88/GLM-5.1-1000000x. All credit for the original data and cleaning belongs to those publishers. Only main.jsonl is included. No filtering, shuffling, schema changes, tokenization or truncation. Original JSON fields and complete records are preserved. Shards retain upstream order; random shard… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/GLM-5.1-Reasoning-Main-Sharded.text100K<n<1M0 likes268 downloads19d agoHugging Face10G-reen /fastdetector-train-raw-shardedtabular100K<n<1M0 likes263 downloads5d agoHugging Face11Janetwyw /trackworld-droid-full-processed-sharded DROID processed data (sharded upload) This Hub copy uses deterministic shards to satisfy the Hub directory entry limit. For annotation, geometry, region_heatmaps, skeletons, and latent_videos, training data is stored as <component>/train/<shard>/..., with at most 5000 sample entries per shard. Validation paths retain their original layout. The files are hardlinked to the local processed DROID dataset; no data blocks are duplicated on the source filesystem. 1 likes245 downloads14d agoHugging Face12placeholderlabs /Kimi-K2.5-Reasoning-General-Sharded Kimi-K2.5-Reasoning-General-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: General-Distillation.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.text100K<n<1M0 likes233 downloads19d agoHugging Face13Undi95 /ConversationChronicles-sharegpt-SHARDEDThis is a sharded version of the PocketDoc/ConversationChronicles-sharegpt dataset, a sharegpt conversion of the jihyoung/ConversationChronicles dataset. All dialogue got fixed (space, coma) and spread across the different relationship available : Relationship Count Ratio Classmates 66,090 33.05% Neighbors 49,521 24.76% Co-workers 28,856 14.43% Mentee and Mentor 16,035 8.02% Husband and Wife 13,486 6.74% Patient and Doctor 6,980 3.49% Parent and Child6,514 3.26%… See the full description on the dataset page: https://huggingface.co/datasets/Undi95/ConversationChronicles-sharegpt-SHARDED.text100K<n<1M11 likes228 downloads3y agoHugging Face14G-reen /fastdetector-test-raw-shardedtabular10K<n<100K0 likes213 downloads5d agoHugging Face15G-reen /fastdetector-val-raw-shardedtabular10K<n<100K0 likes197 downloads5d agoHugging Face16placeholderlabs /Nemotron-SFT-Agentic-v2-Selected-Sharded Nemotron-SFT-Agentic-v2-Selected-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from nvidia/Nemotron-SFT-Agentic-v2. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: data/tool_calling.jsonl, data/interactive_agent.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Nemotron-SFT-Agentic-v2-Selected-Sharded.text100K<n<1M0 likes184 downloads18d agoHugging Face17G-reen /cc-2020-raw-shardedtabular100K<n<1M0 likes167 downloads2mo agoHugging Face18G-reen /cc-re-2020-raw-shardedtabular1M<n<10M0 likes150 downloads23d agoHugging Face19tussiiiii /llm-classification-distilled-v2-sharded LLM Classification Distilled v2 Sharded Overview This repository stores shard CSV files produced by the teacher-judge distillation pipeline. How to Use Run the distillation notebook once per shard: NUM_SHARDS = 4 SHARD_INDEX = 0 .. 3 After all shards are uploaded, set RUN_MERGE_SHARDS = True in the notebook to merge these files and upload final train.csv files to the v2 dataset repos. Final Repositories Full:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-sharded.tabulartext-classification100K<n<1M0 likes121 downloads4mo agoHugging Face20huyle1611 /bridgev2_sharded0 likes121 downloads1mo agoHugging Face21G-reen /cc-re-2021-raw-shardedtabular100K<n<1M0 likes110 downloads23d agoHugging Face22montaseri /wham-noise-subset-shardedaudio1K<n<10K0 likes99 downloads9mo agoHugging Face23G-reen /cc-2021-raw-shardedtabular10K<n<100K0 likes97 downloads2mo agoHugging Face24open-llm-leaderboard-old /details_pythainlp__wangchanglm-7.5B-sft-en-sharded Dataset Card for Evaluation run of pythainlp/wangchanglm-7.5B-sft-en-sharded Dataset Summary Dataset automatically created during the evaluation run of model pythainlp/wangchanglm-7.5B-sft-en-sharded on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_pythainlp__wangchanglm-7.5B-sft-en-sharded.0 likes87 downloads3y agoHugging Face25CharlieLLL /SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921 OPD149 bounded evidence7168→3072, latest state at tail: ONE eval150, three 50-task shards 82/150 (54.67%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 69.28 minutes, including queue and any recovery gaps, excluding prior service startup/smoke. Original two-message coordinator/system prompt; only originally visible complete evidence… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-window7168-3072-1run-sharded3x50-20260921.text-generation0 likes71 downloads3d agoHugging Face26mlfoundations-dev /REASONING_evalchemy_sharded_gpt-4o-mini Dataset card for REASONING_evalchemy_sharded_gpt-4o-mini This dataset was made with Curator. Dataset details A sample from the dataset: { "context": [ { "content": "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.There are three cards with letters $\\texttt{a}$, $\\texttt{b}$, $\\texttt{c}$ placed in a row in some… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_sharded_gpt-4o-mini.tabular1K<n<10K0 likes66 downloads2y agoHugging Face27mlfoundations-dev /REASONING_evalchemy_8_sharded_gpt-4o-mini Dataset card for REASONING_evalchemy_8_sharded_gpt-4o-mini This dataset was made with Curator. Dataset details A sample from the dataset: { "context": [ { "content": "Problem: What is the degree measure of the acute angle formed by lines with slopes $2$ and $\\frac{1}{3}$?\nMark your solution with \\boxed\nAnswer:", "role": "user" } ], "gen_kwargs": { "do_sample": false, "max_new_tokens": 32768… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_8_sharded_gpt-4o-mini.tabular1K<n<10K0 likes58 downloads2y agoHugging Face28mlfoundations-dev /REASONING_evalchemy_32_sharded_gpt-4o-mini Dataset card for REASONING_evalchemy_32_sharded_gpt-4o-mini This dataset was made with Curator. Dataset details A sample from the dataset: { "context": [ { "content": "Problem: How many of the same digits are found in the base 7 and base 8 representations of $629_{10}$? For example, $121_{3}$ and $413_{5}$ would have one digit in common.\nMark your solution with \\boxed\nAnswer:", "role": "user" } ], "gen_kwargs": {… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/REASONING_evalchemy_32_sharded_gpt-4o-mini.tabular1K<n<10K0 likes50 downloads2y agoHugging Face29G-reen /ai-det-test-human-refined-raw-shardedtabular1K<n<10K0 likes49 downloads20d agoHugging Face30CharlieLLL /SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921 OPD149 evidence retained, latest state at tail: ONE eval150, three 50-task shards 86/150 (57.33%). This is one evaluation partitioned into three disjoint sets, not three repeats. Each shard has50 episodes, two nodes/eight GB300; total24 GPUs and maximum150 episodes. Full150 completion: 71.76 minutes, including queue and any recovery gaps, excluding prior service startup/smoke. Original two-message coordinator/system prompt; newly exposed evidence events retained, current… See the full description on the dataset page: https://huggingface.co/datasets/CharlieLLL/SWEbench-Verified-eval150-M2.7-OPD149-evidence-tail-1run-sharded3x50-20260921.text-generation0 likes48 downloads4d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.