CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ccoffee20 /flatpak1 likes337k downloads9d agoHugging Face02OpenGVLab /VideoChat-Flash-Training-Data 🦜 VideoChat-Flash-Training-Data This repos contains all annotaions and most videos for training VideoChat-Flash. 📕 How to use the LongVid data? For video_dir like longvid_subset/coin_grounding_10k_zip, you need to concat this dir to a zip file as follows: cat ego4dhcap_eventunderstanding_2k_zip/* > ego4dhcap_eventunderstanding_2k.zip ✏️ Citation @article{li2024videochatflash, title={VideoChat-Flash: Hierarchical Compression for Long-Context… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat-Flash-Training-Data.video-text-to-text10K<n<100K16 likes30k downloads1y agoHugging Face03malaiwah /GLM-5.3-Flash-TR3-partsbin-v1 GLM-5.3-Flash TR3 parts bin v1 — K6 + K8 payload stores under one transform seed This dataset is the parts bin for the GLM-5.3-Flash TR3 quantization campaign (2026-08-27/28): the complete per-choice payload stores of the two published uniform quants, plus the preparation artifacts and provenance receipts that produced them. malaiwah/GLM-5.3-Flash-TR3-6bpw (uniform K6) malaiwah/GLM-5.3-Flash-TR3-8bpw (uniform K8) What a parts bin is TR3 (trellis) encoding is… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-TR3-partsbin-v1.0 likes24k downloads26d agoHugging Face04Open-Orca /FLAN🍮 The WHOLE FLAN Collection! 🍮 Overview This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets. Generated using the official seqio templating from the Google FLAN Collection GitHub repo. The data is subject to all the same licensing of the component datasets. To keep up with our continued work on OpenOrca and other exciting research, find our Discord here: https://AlignmentLab.ai Motivation This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.text100M<n<1B195 likes20k downloads3y agoHugging Face05RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face06Muennighoff /flanThis is a repreprocessed version of the FLAN dataset with any updates that have been made to the FLAN datasets since the release of the original FLAN. The script is available here. Tasks: {'aeslc_10templates', 'ag_news_subset_10templates', 'anli_r1_10templates', 'anli_r2_10templates', 'anli_r3_10templates', 'arc_challenge_10templates', 'arc_easy_10templates', 'bool_q_10templates', 'cb_10templates', 'cnn_dailymail_10templates', 'cola_10templates', 'common_gen_10templates'… See the full description on the dataset page: https://huggingface.co/datasets/Muennighoff/flan.textother1M<n<10M52 likes15k downloads4y agoHugging Face07oguzhanmeteozturk /flame-runs0 likes10k downloads2d agoHugging Face08justachetan /flat-pack-bench Flat-Pack Bench 🧩 Furniture assembly as a spatio-temporal stress test for large vision-language models. Flat-Pack Bench is a multiple-choice benchmark for evaluating fine-grained spatio-temporal understanding in real furniture assembly videos. Each question asks a model to reason about object parts, contact events, assembly order, final connectivity, or part identity across time. Project page: https://flat-pack-bench.github.io 🎯 Benchmark Tasks The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/justachetan/flat-pack-bench.imagevisual-question-answeringn<1K0 likes9.6k downloads4mo agoHugging Face09flashinfer-ai /flashinfer-trace FlashInfer Trace We provide an official dataset called FlashInfer Trace with kernels and workloads in real-world AI system deployment environments. FlashInfer-Bench can use this dataset to measure and compare the performance of kernels. It follows the FlashInfer Trace Schema. It is organized as follows: flashinfer_trace/ # Here ├── definitions/ └── workloads/ flashinfer-trace/ # On Hugging Face ├── solutions/ └── traces/ Example solutions and traces directories, featuring… See the full description on the dataset page: https://huggingface.co/datasets/flashinfer-ai/flashinfer-trace.20 likes7.7k downloads4mo agoHugging Face10flaviagiammarino /vqa-rad Dataset Card for VQA-RAD Dataset Description VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from MedPix, which is a free open-access online database of medical images. The question-answer pairs were manually generated by a team of clinicians.… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/vqa-rad.imagevisual-question-answering1K<n<10K104 likes7.7k downloads3y agoHugging Face11medalpaca /medical_meadow_medical_flashcards Dataset Card for Medical Flashcards Dataset Summary Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge, and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.textquestion-answering10K<n<100K49 likes7.4k downloads3y agoHugging Face12stepfun-ai /Step-3.5-Flash-SFT Step-3.5-Flash-SFT Step-3.5-Flash-SFT is a general-domain supervised fine-tuning release for chat models. This repository keeps the full training interface in one place: json/: canonical raw training data tokenizers/: tokenizer snapshots for Step-3.5-Flash and Qwen3, released to preserve chat-template alignment compiled/: tokenizer-specific compiled shards for StepTronOSS training Data Format Each raw shard is a JSON file whose top level is a list of examples.… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/Step-3.5-Flash-SFT.text-generation1M<n<10M347 likes7.1k downloads7mo agoHugging Face13hf-internal-testing /transformers_flash_attn_ci0 likes7k downloads11h agoHugging Face14IGNF /FLAIR-HUB FLAIR-HUB : Large-scale Multimodal Dataset for Land Cover and Crop Mapping FLAIR-HUB builds upon and includes the FLAIR#1 and FLAIR#2 datasets, expanding them into a unified, large-scale, multi-sensor land-cover resource with very-high-resolution annotations. Spanning over 2,500 km² of diverse French ecoclimates and landscapes, it features 63 billion hand-annotated pixels across 19 land-cover and 23 crop type classes. The dataset integrates complementary data sources including… See the full description on the dataset page: https://huggingface.co/datasets/IGNF/FLAIR-HUB.image-segmentation100K<n<1M34 likes6.5k downloads4mo agoHugging Face15flaviagiammarino /path-vqa Dataset Card for PathVQA Dataset Description PathVQA is a dataset of question-answer pairs on pathology images. The dataset is intended to be used for training and testing Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions. The dataset is built from two publicly-available pathology textbooks: "Textbook of Pathology" and "Basic Pathology", and a publicly-available digital library: "Pathology… See the full description on the dataset page: https://huggingface.co/datasets/flaviagiammarino/path-vqa.imagevisual-question-answering10K<n<100K75 likes6k downloads3y agoHugging Face16yaacovgg /shiur-clips-flac0 likes6k downloads2mo agoHugging Face17laion /laions_got_talent_enhanced_flash_annotations_and_long_captions18 likes5.4k downloads2y agoHugging Face18brandonmusic /GLM-5.3-Flash-BF16-Teacher-Logits GLM-5.3-Flash BF16 teacher logits This dataset contains full-vocabulary float32 teacher logits from the immutable zai-org/GLM-5.3-Flash-BF16 revision a6c167b62691b2bac901344b65cb651a70f53e43. It keeps the sealed final KLD panel qualification-only and publishes the separate non-final calibration panel under role-specific paths. Qualification-only final windows: 25 Qualification-only final prediction positions: 51175 Vocabulary size: 154880 Teacher receipt:… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits.text-generation4 likes5.3k downloads29d agoHugging Face19richardchencccc /Flash100K3 likes4.5k downloads1mo agoHugging Face20flarevault /kj0 likes4.3k downloads3mo agoHugging Face21orionweller /tulu_flan_mds_incremental-tokens0 likes4.1k downloads2y agoHugging Face22Vision-Flan /vision-flan_191-task_1k 🚀 Vision-Flan Dataset vision-flan_191-task-1k is a human-labeled visual instruction tuning dataset consisting of 191 diverse tasks and 1,000 examples for each task. It is constructed for visual instruction tuning and for building large-scale vision-language models. Paper or blog for more information: https://github.com/VT-NLP/MultiInstruct/ https://vision-flan.github.io/ Paper coming soon 😊 Citation Paper coming soon 😊. If you use Vision-Flan, please use the… See the full description on the dataset page: https://huggingface.co/datasets/Vision-Flan/vision-flan_191-task_1k.imagevisual-question-answering100K<n<1M22 likes3.6k downloads3y agoHugging Face23orionweller /tulu_flan_mds_incremental0 likes3.4k downloads2y agoHugging Face24chiayewken /flan-v2 Dataset Card for "flan-v2" More Information needed text10M<n<100M4 likes3.3k downloads3y agoHugging Face250xAIT /sinhala-flantext10M<n<100M3 likes3k downloads2y agoHugging Face26ByteDance-Seed /Multi-SWE-bench-flash 👋 Overview This repository contains the Multi-SWE-bench dataset, introduced in Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving, to address the lack of multilingual benchmarks for evaluating LLMs in real-world code issue resolution. Unlike existing Python-centric benchmarks (e.g., SWE-bench), this framework spans 7 languages (Java, TypeScript, JavaScript, Go, Rust, C, and C++) with 1,632 high-quality instances, curated from 2,456 candidates by 68 expert annotators… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/Multi-SWE-bench-flash.text-generation3 likes3k downloads9mo agoHugging Face27flagrantia /character_select_stand_alone_apphttps://github.com/mirabarukaso/character_select_stand_alone_app 7 likes3k downloads13d agoHugging Face28FlagEval /EmbSpatial-Bench Introduction Disclaimer: This dataset is organized and adapted from Phineas476/EmbSpatial-Bench. The original data was image format and has been converted here into a more accessible and easy-to-use format. EmbSpatial-Bench is a benchmark for evaluating embodied spatial understanding of LVLMs. The benchmark is automatically derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective. The constructed benchmark comprises a total of 3,640 QA pairs… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/EmbSpatial-Bench.image1K<n<10K6 likes2.9k downloads1y agoHugging Face29laion /laions_got_talent_enhanced_just_flash_annotations0 likes2.7k downloads2y agoHugging Face30ChanceFocus /flare-finqa Dataset Card for "flare-finqa" More Information needed text1K<n<10K3 likes2.6k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.