CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RoganInglis /vllm-control-arena vLLM Main Tasks Dataset AI coding tasks generated from vLLM git commits Dataset Description This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work. Dataset Structure The dataset contains the following columns: commit_hash: The git commit hash parent_hash: The parent commit hash commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.tabulartext-generation1K<n<10K0 likes21k downloads1y agoHugging Face02OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes4.5k downloads8mo agoHugging Face03OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.7k downloads7mo agoHugging Face04ericktwo /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M1 likes1.4k downloads8mo agoHugging Face05NarsAI /FineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/FineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1.4k downloads8mo agoHugging Face06NuTonic /sat-vl-sft-training-ready-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.imagetext-generation100K<n<1M2 likes1.3k downloads5mo agoHugging Face07Sandeepthakur /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/Sandeepthakur/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes1k downloads8mo agoHugging Face08tvu-vlinhd11 /vi-dataset-for-pretrain Dataset Card for "vi-dataset-for-pretrain" This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc. The dataset consists of: vietgpt/covid_19_news_vi hieunguyen1053/binhvq-news-corpus oscar (unshuffled_deduplicated_vi) vietgpt/wikipedia_vi Dataset info Splits N.o examples Size Train 23,891,116 77.36 GB Validation 1,257,428 4.06 GB Total 25,148,544 81.43 GB texttext-generation10M<n<100M0 likes704 downloads8mo agoHugging Face09tvu-vlinhd11 /pretrain-dataset-raw-10M Pretrain Dataset (Text) This dataset contains preprocessed text documents ready for LLM pretraining. Dataset Details Property Value Documents 10,000,000 Processed 10000000 Shards 21 Created 2025-12-09 Dataset Structure Each sample contains: text: The document text source: Source dataset identifier id: Unique document ID Usage from datasets import load_dataset dataset = load_dataset("tvu-vlinhd11/pretrain-dataset-raw-10M")… See the full description on the dataset page: https://huggingface.co/datasets/tvu-vlinhd11/pretrain-dataset-raw-10M.texttext-generation10M<n<100M0 likes643 downloads10mo agoHugging Face10prism-vlm /gemini_public_mmr1 PRISM Public SFT Data Overview PRISM Public SFT Data is the public supervised fine-tuning data collection used in the PRISM project.PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline for large multimodal models. Before the distribution alignment and RLVR stages, we first use large-scale public multimodal demonstrations to obtain a broad SFT initialization. This dataset serves as the public SFT data source for the… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/gemini_public_mmr1.textimage-to-text1M<n<10M2 likes609 downloads5mo agoHugging Face11NuTonic /sat-vl-sft-postprocessed-merged-v1 Dataset Summary NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills). The goal is to create high-signal, production-shaped supervision for multimodal chat models: Captioning for satellite chips Grounding (bounding boxes in normalized coordinates) for land-cover regions Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-postprocessed-merged-v1.imagetext-generation100K<n<1M0 likes502 downloads5mo agoHugging Face12vllg /loong_c4A filtered subset of C4-en containing 3,584,358 pages that are at least 16,000 characters long, useful for training models with longer context windows. texttext-generation100K<n<1M1 likes452 downloads3y agoHugging Face13MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes452 downloads10mo agoHugging Face14OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes419 downloads8mo agoHugging Face15vllg /long_c4A filtered subset of C4-en containing 13,688,429 pages that are at least 8,000 characters long, useful for training models with longer context windows. texttext-generation100K<n<1M2 likes270 downloads3y agoHugging Face16dans25275 /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/dans25275/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M0 likes202 downloads8mo agoHugging Face17undertheseanlp /UTS_VLC Dataset Card for Vietnamese Legal Corpus (UTS_VLC) A curated corpus of Vietnamese Laws and Codes (Luật, Bộ luật) and the Constitution, maintained by Underthesea NLP. The flagship 2026 split is a verified in-force snapshot — every document is currently in force, de-duplicated, and validated against Vietnam's official legal database vbpl.vn. Dataset Details Dataset Description UTS_VLC contains the full text of Vietnamese legislation at the top of the… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS_VLC.texttext-generationn<1K2 likes178 downloads4mo agoHugging Face18prism-vlm /rl_dataset PRISM RL Dataset Overview PRISM RL Dataset contains the training data used for the PRISM alignment and RLVR stages. PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline for large multimodal models. Instead of directly applying RLVR after SFT, PRISM inserts an intermediate Distribution Alignment / Pre-alignment stage based on black-box on-policy distillation. The overall pipeline is: SFT → PRISM Alignment → RLVR This… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/rl_dataset.textimage-to-text10K<n<100K0 likes149 downloads5mo agoHugging Face19MichielBuisman /Leesplank-vloeiend-nl-curriculum-cp2 Leesplank NL — Embeddings & Clusters (Checkpoint 2) 5.39 million Dutch texts, each annotated with a 384-dimensional semantic embedding, a K-means cluster assignment, and — for a stratified sample of ~100,800 rows — a measured BF16 perplexity score from IBM Granite 4.0 Micro (3B Dense). This checkpoint is the analytical core of a project with two concrete goals: running high-quality Dutch language processing on consumer hardware, and doing it in a way that is transparent enough for… See the full description on the dataset page: https://huggingface.co/datasets/MichielBuisman/Leesplank-vloeiend-nl-curriculum-cp2.texttext-generation10K<n<100K0 likes138 downloads7mo agoHugging Face20prism-vlm /gemini_distill PRISM Gemini Distill Overview PRISM Gemini Distill is our self-distilled multimodal reasoning dataset collected from Gemini 3 Flash for the PRISM project. PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline. To mitigate this issue, PRISM introduces an intermediate Distribution Alignment / Pre-alignment stage before RLVR: SFT → Distribution Alignment / Pre-alignment → RLVR This dataset provides high-quality Gemini 3… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/gemini_distill.textimage-to-text100K<n<1M1 likes126 downloads5mo agoHugging Face21vllg /looong_c4A filtered subset of C4-en containing 835,400 pages that are at least 32,000 characters long, useful for training models with longer context windows. texttext-generation100K<n<1M1 likes105 downloads3y agoHugging Face22vldsavelyev /guitar_tabDataset of music tablature, in alphaTex (https://alphatab.net/docs/alphatex) format, converted from Guitar Pro files (gp3, gp4, gp5, which are downloaded from https://rutracker.org/forum/viewtopic.php?t=2888130texttext-generation10K<n<100K11 likes92 downloads3y agoHugging Face23eyes-ml /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096 Derived dataset note This dataset was derived from OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking as a part of arxiv.org/abs/2603.22276. Field changes: question -> query qwen3vl_235b_thinking_response -> response image -> images (single-item list) added tok_len, computed with tokenizer Qwen/Qwen3-8B on query + '\n\n' + response add_special_tokens=False The original README content is preserved below. MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/eyes-ml/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking-QR-max4096.imagevisual-question-answering10K<n<100K0 likes81 downloads6mo agoHugging Face24vladimirbesk /tsiolkovsky-papers Tsiolkovsky Papers: the complete personal archive as a machine-readable corpus Machine transcriptions of all 51,008 sheets of fond 555 of the Archive of the Russian Academy of Sciences — the personal archive of Konstantin Tsiolkovsky (1857–1935), who derived the rocket equation and described the multistage rocket decades before anyone could test either. The archive had been scanned and put online, but without a catalogue you could query, full-text search, or a dataset. This is… See the full description on the dataset page: https://huggingface.co/datasets/vladimirbesk/tsiolkovsky-papers.tabulartext-generation10K<n<100K2 likes75 downloads1mo agoHugging Face25vluxblaring /re-tutor-protection-mechanisms RE-Tutor: Protection-Mechanism Analysis Dataset Instruction-tuning dataset teaching a model to analyze protection mechanisms (anti-debug, anti-VM, anti-tamper, anti-dump, obfuscation, timing) from code evidence and emit structured expert analysis. Schema Each sample pairs input (code evidence) with output (structured analysis): input.code_snippet: C source, decompiler-style pseudocode, or x86/x64 assembly input.imports_pool: mixed DLL!API imports (includes… See the full description on the dataset page: https://huggingface.co/datasets/vluxblaring/re-tutor-protection-mechanisms.texttext-generationn<1K0 likes75 downloads22d agoHugging Face26YangyiYY /VLM-SFTimagetext-generation1M<n<10M2 likes69 downloads2y agoHugging Face27vlinhd11 /vietnamese-sft-10k Vietnamese Instruction-Following Dataset (10K) This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts. Format: JSONL (one object per line) Fields: "prompt" (instruction or user message), "response" (assistant reply) Language:… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-sft-10k.texttext-classification10K<n<100K0 likes64 downloads21d agoHugging Face28UCSC-VLAA /CIK-Bench CIK-Bench Official dataset for Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw. CIK-Bench evaluates OpenClaw — the most widely deployed personal AI agent in early 2026 — against persistent-state poisoning attacks. It implements the CIK taxonomy, a unified framework that organizes OpenClaw's persistent state into three dimensions: Capability — executable skills (SKILL.md, .sh, .py) Identity — persona, values, and behavioral configuration (SOUL.md, IDENTITY.md… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/CIK-Bench.texttext-classificationn<1K0 likes62 downloads5mo agoHugging Face29ThunderstormXXL /deepscaler-teacher-sft-vllm-official-40k DeepScaleR teacher SFT vLLM official 40k Generated run: exp_003_vllm_official_brainlab_2gpu. Summary { "num_examples": 40300, "sft_dir": "data/processed/deepscaler/teacher_sft/exp_003_vllm_official_brainlab_2gpu", "parse_rate": 0.9999751861042183, "correct_rate": 0.5728039702233251, "format_rate": 0.005955334987593052, "mean_reward": 0.42432258064534184, "deepscaler_mean_reward": 0.6266997518610422, "deepscaler_match_mean_reward":… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.texttext-generation10K<n<100K0 likes51 downloads4mo agoHugging Face30vlinhd11 /medical_sft_crawl_vi_10k_v1 ViMed-SFT: Vietnamese Medical Conversational Dataset Dataset Description ViMed-SFT is a high-quality Vietnamese medical conversational dataset designed for Supervised Fine-Tuning (SFT) of Large Language Models. The dataset contains medical Q&A conversations between users and a virtual medical assistant. Dataset Summary Attribute Value Language Vietnamese Domain Healthcare / Medical Task Conversational AI, Instruction Tuning Samples… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/medical_sft_crawl_vi_10k_v1.texttext-generation10K<n<100K0 likes51 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.