CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-dna /genomes-v4-genome_set-animals-intervals-v8_256_128text10M<n<100M0 likes930 downloads8mo agoHugging Face02suitai /salabs-virtual-spatial-digitaltwin-v8 🌐 SALabs 10,000,000-Node 3D Virtual Spatial & Digital Twin Avatar Kinematics Dataset (v8.0) [!IMPORTANT] 💳 Click Here to Purchase Enterprise Commercial License ($2,000 USD) & Instant 8.0GB Master DownloadInstant download of the complete 8.0GB master archive containing 10,000,000 verified 3D spatial nodes, 18-DoF avatar kinematics, B-spline 4D motion tensors, Laplace-Beltrami spectral resonance, and commercial license certificate. 🌟 Executive Summary The… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-virtual-spatial-digitaltwin-v8.tabularother1K<n<10K1 likes136 downloads18d agoHugging Face03nyu-dice-lab /lm-eval-results-bobofrut-ladybird-base-7B-v8-private Dataset Card for Evaluation run of bobofrut/ladybird-base-7B-v8 Dataset automatically created during the evaluation run of model bobofrut/ladybird-base-7B-v8 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bobofrut-ladybird-base-7B-v8-private.tabular100K<n<1M0 likes125 downloads2y agoHugging Face04nyu-dice-lab /lm-eval-results-zhengr-MixTAO-7Bx2-MoE-v8.1-private Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1 Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-zhengr-MixTAO-7Bx2-MoE-v8.1-private.tabular100K<n<1M0 likes84 downloads2y agoHugging Face05KG-ZEROIN /zeroin-v8-r4-dataset Zeroin Methodology v8-r4 QA Training corpus used to fine-tune KG-ZEROIN/gpt-oss-20b-zeroin-v8-r4. The corpus captures question/answer pairs derived from the Zeroin fund-evaluation methodology Korean domain document. It is organized as a Harmony-ready chat-messages dataset for supervised full fine-tuning of openai/gpt-oss-20b. Released under CC BY-NC 4.0 — free for non-commercial research, evaluation, and educational use. See LICENSE and NOTICE. Contents File… See the full description on the dataset page: https://huggingface.co/datasets/KG-ZEROIN/zeroin-v8-r4-dataset.textquestion-answering1K<n<10K0 likes60 downloads4mo agoHugging Face06open-llm-leaderboard /Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-detailsgated Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8 Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details.tabular10K<n<100K0 likes58 downloads2y agoHugging Face07open-llm-leaderboard /Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-detailsgated Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9 Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details.tabular10K<n<100K0 likes58 downloads2y agoHugging Face08open-llm-leaderboard /zhengr__MixTAO-7Bx2-MoE-v8.1-detailsgated Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1 Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zhengr__MixTAO-7Bx2-MoE-v8.1-details.tabular10K<n<100K0 likes53 downloads2y agoHugging Face09baablabs /baab-next-v8-rawtext10K<n<100K0 likes49 downloads2mo agoHugging Face10open-llm-leaderboard /pankajmathur__orca_mini_v8_1_70b-detailsgated Dataset Card for Evaluation run of pankajmathur/orca_mini_v8_1_70b Dataset automatically created during the evaluation run of model pankajmathur/orca_mini_v8_1_70b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/pankajmathur__orca_mini_v8_1_70b-details.tabular10K<n<100K0 likes45 downloads2y agoHugging Face11FabianOvalle /Dataset_Robot_IA_full_v8textn<1K0 likes44 downloads27d agoHugging Face12AlekseyCalvin /Lyrical_MT_ru2en_SFT_v8.1_w_solving4gemma4 SilverAgePoets.com & RuVERSES.com Russian-English Bilingual Poetry Library A dataset of Eastern European and Soviet poetry and song lyrocs from https://RuVerses.com/, with Russian-language sources and English translations. This variant of the dataset combines a revised and somewhat pre-filtered version of the RuVerses collection dataset + the entirety of the SFT version of our LYRICAL dataset. Featuring a present (c. early 2026) state of the RuVerses archive, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_MT_ru2en_SFT_v8.1_w_solving4gemma4.texttranslation1K<n<10K3 likes41 downloads6mo agoHugging Face13neoneye /simon-arc-solve-symmetry-v8 Version 1 ARC-AGI Tasks where the job is to transform symmetric images. example count: 2-4. test count: 1-2. image size: 2-3. symmetry types: hstack2, hstack3, vstack2, vstack3, grid2x2. Version 2 image size: 2-4. Version 3 Added HSTACK4, VSTACK4. Version 4 Added HSTACK5, VSTACK5. Version 5 Added ImageSymmetrySquare, so images can be rotated by 90 degrees, and flipped over the diagonals. Version 6 Only exercising… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/simon-arc-solve-symmetry-v8.textimage-to-text100K<n<1M1 likes40 downloads2y agoHugging Face14rghosh8 /supportGPT-v8text1K<n<10K1 likes32 downloads3y agoHugging Face15open-llm-leaderboard /Lil-R__PRYMMAL-ECE-7B-SLERP-V8-detailsgated Dataset Card for Evaluation run of Lil-R/PRYMMAL-ECE-7B-SLERP-V8 Dataset automatically created during the evaluation run of model Lil-R/PRYMMAL-ECE-7B-SLERP-V8 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__PRYMMAL-ECE-7B-SLERP-V8-details.tabular10K<n<100K0 likes32 downloads2y agoHugging Face16k3nn3dy /blueteam-v8 Blue_team_v8 Synthetic fine-tuning dataset generated with Dataset Genie 0.1.0 on 2026-09-21T15:02:28+00:00. Domain brief Blue-team is broad. Cover these deliberately, spread across difficulty tiers: Platforms, not one vendor. Splunk (SPL), Microsoft Sentinel (KQL), Elastic (ES|QL/EQL/Lucene), CrowdStrike, Defender for Endpoint, Sysmon, Zeek/Suricata. Identity: Entra ID, Okta, Active Directory/Kerberos. Cloud: AWS (CloudTrail/GuardDuty), Azure, GCP. OS: Windows… See the full description on the dataset page: https://huggingface.co/datasets/k3nn3dy/blueteam-v8.texttext-generation1K<n<10K0 likes31 downloads3d agoHugging Face17Irfanuruchi /buildeng-v8-3b BuildEng V8 3B Dataset This is the dataset holder for the BuildEng 3B branch. The 3B branch is planned as the stronger local checkpoint after BuildEng 1.5B and before the larger BuildEng 32B model. Files will be added on June 4. Author: Irfan Uruchi text10K<n<100K0 likes26 downloads4mo agoHugging Face18AlekseyCalvin /Lyrical_MT_ru2en_SFT_v8_with_meter_solving SilverAgePoets.com & RuVERSES.com Russian-English Bilingual Poetry Library Our curated dataset of translated Eastern European and Soviet poetry and song lyrics from SilverAgePoets.com and RuVerses.com/. The dataset features Russian-language sources, English translations, scansion information (meter/foot, rhythm, rhyme), and concisely sketched multistep source-to-target adaptation workflows (formatted as thinking traces). This variant of the dataset combines a revised and somewhat… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Lyrical_MT_ru2en_SFT_v8_with_meter_solving.texttranslation1K<n<10K2 likes24 downloads5mo agoHugging Face19open-llm-leaderboard /Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.7-detailsgated Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.7 Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.7 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.7-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face20Irfanuruchi /buildeng-v8-1.5b BuildEng V8 1.5B BuildEng V8 1.5B is a civil/building engineering dataset made for Qwen2.5-1.5B-Instruct and focused on conservative structural reasoning and construction-related decision making. The dataset covers reinforced concrete, steel, masonry, shallow foundations, retaining walls, slabs, beams, columns, temporary works, excavation safety, waterproofing, settlement, lateral stability, structural connections, cracking behavior, renovation uncertainty, and construction… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/buildeng-v8-1.5b.text10K<n<100K0 likes23 downloads4mo agoHugging Face21jxx123 /loop-qwen-v8-sftgated loop-qwen-v8 SFT dataset (Gemini insulin-control distillation) 8,222 chat-format examples used to SFT loop-qwen-v8 (Qwen3-4B insulin controller distilled from Gemini-3-flash-preview). Each example is a closed-loop dosing decision. Format (JSONL, one chat per line) system: controller spec (IOB-aware, Chain-of-Draft reason-before-act) user: patient metadata (age, weight, TDD, CF, IC, basal) + 6h history of CGM / insulin / carbs, as JSON assistant: JSON… See the full description on the dataset page: https://huggingface.co/datasets/jxx123/loop-qwen-v8-sft.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face22sameer-saraf-quant-ai /slm-workflow-planner-v8-datasettext100K<n<1M0 likes21 downloads7mo agoHugging Face23pankajmathur /orca_mini_v8_sharegpt_formatBest 45K samples of Bigger Orca Mini dataset in sharegpt format, Enjoy! texttext-generation10K<n<100K4 likes18 downloads2y agoHugging Face24ultrastar111 /sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k10_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes17 downloads3mo agoHugging Face25AEUPH /synthetic_Jailbreak_Defense_Doorpage_v8 synthetic_Jailbreak_Defense_Doorpage_v8 Silicon Factory v3 - Synthetic Dataset Entries: 5 Category: mixed Avg Response Length: 447 chars Focus: AI JAILBREAK DEFENSE Mode: Doorpage (auto-gen + fine-tune) License MIT Generated With Tree-Speculative Decoding 4D Brane Memory for consistency Quality control & deduplication Contact & Custom Orders Custom datasets available. Contact for pricing. textn<1K0 likes16 downloads6mo agoHugging Face26ultrastar111 /sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k5_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes16 downloads3mo agoHugging Face27kkomyoeminaung /myanmar_11k_v8 Dataset Card for myanmar_11k_v8 Dataset Summary ဒီ dataset က myanmar_11k_v8 အတွက် ဖန်တီးထားတာပါ။ Languages Myanmar (my) / English (en) Dataset Structure Data Instances { "text": "နမူနာ စာသား", "label": "အညွှန်း" } Data Fields text: main content, label: optional. Data Splits Split Files train myanmar_11k_v8.jsonl Licensing Information ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/myanmar_11k_v8.text10K<n<100K0 likes14 downloads2mo agoHugging Face28ultrastar111 /sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg Sokoban action-conditioned visual world-model SFT data (non-CoT baseline) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding).… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_noncot_chunk_k3_world_model_20260622_perseg.tabularreinforcement-learning100K<n<1M0 likes14 downloads3mo agoHugging Face29neoneye /simon-arc-mass-v8 Version 1 Measure the mass of objects for pixel connectivity 4 and pixel connectivity 8. image size: 1-10. max_mass: 4. Version 2 image size: 1-20. max_mass: 5. This converged too slowly. I was too optimistic. I will have to proceed slower. Version 3 image size: 1-12. max_mass: 5. Too big spikes in the training loss. I will have to lower the max_mass, and gradually increase it. Version 4 image size: 1-15. max_mass: 2. The validation loss for this is… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/simon-arc-mass-v8.textimage-to-text100K<n<1M0 likes13 downloads2y agoHugging Face30ultrastar111 /sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg Sokoban action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT VLM-Gym feedback-interval study. Format: gzipped JSONL shards under training/, one packed row = one episode. Frames are base64 JPEG (q95). Per-segment CoT layout: <think> per-step imagined frame (MSE target) </think> then the committed action chunk; between chunks a loss-0 "Action executed." + real frame (GT re-grounding). See… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/sokoban_easy_v8_cot_chunk_kinf_world_model_20260707_perseg.tabularreinforcement-learning100K<n<1M0 likes13 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.