CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aimosprite /prompt-swap-mixed12-5xlr-e1-mxfp4-mergedtabularn<1K0 likes588 downloads6mo agoHugging Face02aimosprite /prompt-swap-mixed12-5xlr-e2-mxfp4-mergedtabularn<1K0 likes579 downloads6mo agoHugging Face03empero-ai /MiniMax-M3-150k-Mixed m3-alldomains-verified-107k Verified distillation traces generated with faststill v0.0.1 — a pipeline that generates (prompt, reasoning, output) triplets from any OpenAI-compatible chat-completions endpoint and deterministically verifies every row before keeping it. A row is verified=true only when a machine check (executed unit tests, exact / normalized answer compare) confirmed it, so wrong labels are filtered out instead of poisoning a student model. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/MiniMax-M3-150k-Mixed.tabulartext-generation100K<n<1M10 likes131 downloads3mo agoHugging Face04juliannunezb /mixed-pretrain-10b-gpt2 Mixed Pretraining 10B (GPT-2 BPE) A 10-billion-token pretraining dataset, GPT-2 BPE tokenized, assembled as a diverse mix of web text, books, Wikipedia, code, academic papers, Q&A and instruction-formatted conversations. Built to train a ~500M parameter from-scratch GPT-2-style transformer (see juliannunezb/transformer-lm-500m). Mix Source Mix % Tokens Notes fineweb 40.4% 4,039,999,700 reused from kjj0/fineweb10B-gpt2 fineweb_edu 15.2% 1,514,999,900 reused… See the full description on the dataset page: https://huggingface.co/datasets/juliannunezb/mixed-pretrain-10b-gpt2.tabulartext-generationn<1K1 likes71 downloads5mo agoHugging Face05ansulev /minimax-m3-150k-mixed m3-alldomains-verified-107k Verified distillation traces generated with faststill v0.0.1 — a pipeline that generates (prompt, reasoning, output) triplets from any OpenAI-compatible chat-completions endpoint and deterministically verifies every row before keeping it. A row is verified=true only when a machine check (executed unit tests, exact / normalized answer compare) confirmed it, so wrong labels are filtered out instead of poisoning a student model. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/minimax-m3-150k-mixed.tabulartext-generation100K<n<1M0 likes64 downloads3mo agoHugging Face06GenVRadmin /Samvaad-Mixed-Language-2tabular10K<n<100K4 likes36 downloads3y agoHugging Face07agentlans /prompt-difficulty-mixed Prompt Difficulty Meta-Analysis Introduction The difficulty of large language model (LLM) prompts varies widely, from simple queries to complex multi-step reasoning tasks. This study develops a consistent, data-driven difficulty score for English ChatGPT prompts, using classifiers trained on labelled difficulty datasets. The goal is to improve automated prompt difficulty classification. Methods Detailed methods Several methods were used to quantify the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-difficulty-mixed.tabulartext-classification10K<n<100K0 likes34 downloads10mo agoHugging Face08HYUNJINI /AXXXX_jssp_mixed_step_train_dispatch_v1tabular1M<n<10M0 likes26 downloads6mo agoHugging Face09GenVRadmin /Samvaad-Mixed-Language-3tabular10K<n<100K0 likes23 downloads3y agoHugging Face10opendiffusionai /laion2b-mixed-1024px-human Overview This dataset is a selective merge of some other of our datasets. Mainly, I pulled human-centric, real-world photos from the following datasets: opendiffusionai/laion2b-45ish-1120px opendiffusionai/laion2b-squareish-1024px opendiffusionai/laion2b-23ish-1216px As such, it is "mixed" aspect ratio. The very smallest height ones are from our "squarish" set, so are at least 1024px tall. However, the other ones with longer rations, have an appropriately longer minimum… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-mixed-1024px-human.image10K<n<100K3 likes20 downloads2y agoHugging Face11yrdu /S3-CoT-Self-Sampled-Data-v2_mixed Dataset Overview We release a new self-sampled variable-length reasoning dataset based on DeepSeek-R1-Distill-Qwen-7B. The seed problems are collected from three high-quality datastes: MATH500, LIMO-v2, and AIME problems released before 2023. The data are used for efficient cot learning in our paper (S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs). For each instruction (problem), we provide multiple reasoning traces generated under different… See the full description on the dataset page: https://huggingface.co/datasets/yrdu/S3-CoT-Self-Sampled-Data-v2_mixed.tabular1K<n<10K0 likes20 downloads3mo agoHugging Face12vsamuel /mixed_data_onetabular10K<n<100K0 likes16 downloads2y agoHugging Face13novastar111 /sokoban_adaptive_3box_hard_mixed_350 sokoban_adaptive_3box_hard_mixed_350 BAGEL VLM-Gym world-model dataset (sokoban / eval). 350 hard 3-box episodes: deadlock + trivial mixed in one contiguous pool; train-disjoint adaptive-thinking (when-to-think) benchmark. layout: Held-out eval episode pool. shard_*.jsonl.gz at the repo root; one episode per row with a contiguous global index. images are base64-encoded JPEG frames stored inline in each JSONL row. Pairs with the matching sokoban checkpoint(s) under the… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/sokoban_adaptive_3box_hard_mixed_350.tabularn<1K0 likes9 downloads1mo agoHugging Face14novastar111 /sokoban_easy_mixed_deadlock_trivial sokoban_easy_mixed_deadlock_trivial BAGEL VLM-Gym world-model dataset (sokoban / eval). Deadlock (temptation tiers) + trivial (no-deadlock) episodes in one contiguous pool; train-disjoint elicitation eval. layout: Held-out eval episode pool. shard_*.jsonl.gz at the repo root; one episode per row with a contiguous global index. images are base64-encoded JPEG frames stored inline in each JSONL row. Pairs with the matching sokoban checkpoint(s) under the companion model org; CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/sokoban_easy_mixed_deadlock_trivial.tabularn<1K0 likes8 downloads1mo agoHugging Face15ijakenorton /mvdsc_mixed_for_ml@inproceedings{10.1145/3488932.3527288, author = {Zhou, Xin and Verma, Rakesh M.}, title = {Vulnerability Detection via Multimodal Learning: Datasets and Analysis}, year = {2022}, isbn = {9781450391405}, publisher = {Association for Computing Machinery}, address = {New York, NY, USA}, url = {https://doi.org/10.1145/3488932.3527288}, doi = {10.1145/3488932.3527288}, abstract = {A vulnerability is a weakness that can be exploited by an attacker, e.g., performing unauthorized actions within a… See the full description on the dataset page: https://huggingface.co/datasets/ijakenorton/mvdsc_mixed_for_ml.tabular10K<n<100K0 likes6 downloads1y agoHugging Face16GenVRadmin /Samvaad-Mixed-Languagetabular1K<n<10K0 likes5 downloads3y agoHugging Face17vsamuel /mixed_data_zerotabular10K<n<100K0 likes5 downloads2y agoHugging Face18HYUNJINI /AXXXX_jssp_mixed_step_train_all_v1tabular1M<n<10M0 likes5 downloads6mo agoHugging Face19HYUNJINI /data_fix_before_jssp_mixed_step_train_dispatch_v1tabular1M<n<10M0 likes4 downloads6mo agoHugging Face20HYUNJINI /data_fix_before_jssp_mixed_step_train_all_v1tabular1M<n<10M0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.