CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes2.9k downloads5mo agoHugging Face02General-Medical-AI /GMAI-Reasoning10K GMAI-Reasoning10K Medical Reasoning dataset used in GMAI-VL-R1 Data description GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI. Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.imagevisual-question-answering10K<n<100K6 likes1.4k downloads1y agoHugging Face03MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes557 downloads10mo agoHugging Face04LLaMAX /BenchMAX_General_Translation Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_General_Translation is a dataset of BenchMAX, which evaluates the translation capability on the general domain. We collect parallel test data from Flore-200, TED-talk, and WMT24. Usage Run the following commands to generate… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_General_Translation.texttranslation100K<n<1M0 likes417 downloads1y agoHugging Face05arcadia-impact /reward-projection-goal-generalisation-vlmtabular1K<n<10K0 likes358 downloads2mo agoHugging Face06ajibawa-2023 /General-Stories-CollectionGeneral Stories Collection A great synthetic datasets consists of around 1.3 million stories especially meant for General audience. You can directly use these datasets for training large models. Total 10 datasets are available for download. You can use any one or all the json files for training purpose. These datasets are in "prompt" and "text" format. Total token length is also available. Thanks for your love & support. texttext-generation1M<n<10M41 likes349 downloads3y agoHugging Face07Antix5 /general-product-token-quality-datasettext1K<n<10K0 likes342 downloads29d agoHugging Face08ericrcwu /week1-general-20b-dolma2-v1 Week-One General 20B Dolma2 This is a deterministic, pretokenized 20-billion-token baseline corpus for controlled language-model architecture and training experiments. It contains nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered extension of the previous view. It also includes a dataset-only 370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count. The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.texttext-generationn<1K0 likes312 downloads2mo agoHugging Face09teknium /GPTeacher-General-InstructGPTeacher General-Instruct dataset is GPT-4 Generated self-instruct dataset. There are multiple versions, with more or less similarity reductions. The dedupe only dataset contains 18194 entries, with less the more similarity is reduced. Format is identical to alpaca's, with a varyiable mix of Instruction/Input/Response, and Instruction/NullInput/Response fields. Learn more on github here:https://github.com/teknium1/GPTeacher text10K<n<100K45 likes275 downloads3y agoHugging Face10placeholderlabs /Kimi-K2.5-Reasoning-General-Sharded Kimi-K2.5-Reasoning-General-Sharded Byte-preserving sequential 100 MB JSONL shards of selected files from Jackrong/Kimi-K2.5-Reasoning-1M-Cleaned. All credit for data generation and upstream curation belongs to the source authors. See the upstream dataset card for attribution, source descriptions and license terms. Included files: General-Distillation.jsonl. No filtering, shuffling, normalization, tokenization or truncation was performed. Complete records and all original fields… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/Kimi-K2.5-Reasoning-General-Sharded.text100K<n<1M0 likes227 downloads16d agoHugging Face11jiaxin-wen /generalization-dynamics-evals Generalization Dynamics — Main Eval Suite Prepared test sets for the 6 main evaluation families from Generalization dynamics across fine-tuning (Table 1). Use with the unified runner: https://github.com/jiaxin-wen/FT-generalization/tree/main/release from huggingface_hub import snapshot_download root = snapshot_download( repo_id="jiaxin-wen/generalization-dynamics-evals", repo_type="dataset") Or browse a single task (the dataset viewer shows all configs): from datasets… See the full description on the dataset page: https://huggingface.co/datasets/jiaxin-wen/generalization-dynamics-evals.texttext-classification10K<n<100K0 likes183 downloads4mo agoHugging Face12OysterCoreAI /SFT-General-Japanese-60K SFT-General-Japanese-60K Welcome to this dataset! 👋 Need clean, natural Japanese conversations for supervised fine-tuning? You are in the right place. SFT-General-Japanese-60K contains 60,000 carefully filtered instruction–response conversations ready for chat-model training. It combines the practical breadth of open Japanese SFT data with transparent gates for safety, recency, formatting, language consistency, and redundancy—so you can focus on training rather… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Japanese-60K.texttext-generation10K<n<100K0 likes155 downloads6d agoHugging Face13OysterCoreAI /SFT-General-Spanish-50K SFT-General-Spanish-50K now available🎇 Una buena conversación no necesita hacer ruido: necesita entender la pregunta, ordenar lo importante y dejar a la otra persona con un siguiente paso claro. SFT-General-Spanish-50K reúne 50.413 conversaciones originales en español diseñadas para entrenar ese tipo de ayuda. Resumen Registros: 50.413 (no se redondeó a 50K). Idioma: español contemporáneo, registro general y neutro. Formato: JSONL de mensajes estilo chat; tres… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Spanish-50K.text10K<n<100K0 likes150 downloads6d agoHugging Face14HPAI-BSC /Aloe-Beta-General-Collection Aloe-Beta-Medical-Collection Collection of curated general datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including: Coding, math, data analysis, STEM, etc. Function calling Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.textquestion-answering10K<n<100K2 likes141 downloads10mo agoHugging Face15luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes114 downloads3mo agoHugging Face16AdvancedDataIntelligence /glm5.2-general-distill Teacher-generated instruction/response pairs used to distill small, local student models (the ADI / Advanced Data Intelligence series) from the frontier teacher glm-5.2. How it was built Teacher: glm-5.2 (served via Ollama Cloud as glm-5.2:cloud), queried with thinking/reasoning disabled so every record is a single clean final answer. Seed prompts: databricks/databricks-dolly-15k, filtered to remove items that require an attached context passage — the closed_qa… See the full description on the dataset page: https://huggingface.co/datasets/AdvancedDataIntelligence/glm5.2-general-distill.texttext-generation1K<n<10K3 likes109 downloads3mo agoHugging Face17tkdonda /gujarati-general-purpose-instruction Gujarati General-Purpose Instruction Dataset (GGJI v1) Dataset Summary GGJI v1 (Gujarati General-Purpose Instruction v1) is a large-scale, high-quality supervised fine-tuning (SFT) dataset designed to train instruction-following language models in Gujarati. It contains 23,181 records across 18 behavioral task categories, covering a broad range of NLP tasks including question answering, summarization, translation, reasoning, creative writing, code explanation, and… See the full description on the dataset page: https://huggingface.co/datasets/tkdonda/gujarati-general-purpose-instruction.texttext-generation10K<n<100K0 likes100 downloads2mo agoHugging Face18Bisilivan /dataset-ohada-droit-commercial-general-echantillon Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon Description Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.texttext-generationn<1K1 likes87 downloads2mo agoHugging Face19agentjudge-anon /GeneralAgentBench GeneralAgentBench GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review. Each task provides a natural-language instruction plus a list of verification checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.documentother1K<n<10K0 likes82 downloads2mo agoHugging Face20BNNT /mozi_general_instructions_3mSources are listed below: Chinese General Instruction 2000k BELLE https://huggingface.co/datasets/BelleGroup/train_2M_CN English generic instruction 52k alpaca-gpt4 https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM Chinese generic dialog instructions 800k BELLE https://huggingface.co/datasets/BelleGroup/multiturn_chat_0.8M English Universal Dialog Instruction 94k sharegpt_vicuna https://huggingface.co/datasets/jeffwan/sharegpt_vicuna Chinese-English-Japanese Universal Command 49k… See the full description on the dataset page: https://huggingface.co/datasets/BNNT/mozi_general_instructions_3m.text1M<n<10M3 likes77 downloads3y agoHugging Face21Phoebe-cyt /GeneralProbe GeneralProbe: A Multi-Domain LLM Capability Probe Dataset Overview This dataset is designed to provide a broad probe of language model capabilities across multiple specialized domains. We collect tens of thousands of samples from publicly available datasets on Hugging Face, covering six domains, and we divide them into six subset: Healthcare Legal Finance Math Science Code The collected tasks are primarily multiple-choice questions and multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/Phoebe-cyt/GeneralProbe.text10K<n<100K0 likes57 downloads23d agoHugging Face22MaatAI /histoire-general-afrique-global-adaption This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. Svngoku/Histoire-General-Afrique-Global This dataset contains French-language text excerpts detailing the political, social, and economic history of Africa from the 16th to the 18th centuries. The content covers specific regions such as the Lower Guinea Coast and the Zambezi, discussing topics like ethnic migrations, kingdom formations, and trade dynamics. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/histoire-general-afrique-global-adaption.textquestion-answering1K<n<10K1 likes56 downloads5mo agoHugging Face23laiba-laiba /llm_general_texttext10M<n<100M0 likes54 downloads2mo agoHugging Face24folkopinion /general-qa-swedishtext10K<n<100K4 likes53 downloads3y agoHugging Face25ajirs /weird-generalization-final-dataset Weird Generalization Final Dataset Clean handoff bundle for the two strongest weird-generalization tasks: 3_1_old_bird_names 3_2_german_city_names This folder intentionally keeps only the data, evaluation materials, and final shareable plots needed to inspect or reuse these tasks. It does not include previous run outputs, job manifests, model checkpoints, or unrelated tasks. Layout datasets/ 3_1_old_bird_names/ train/ test/ original_full/… See the full description on the dataset page: https://huggingface.co/datasets/ajirs/weird-generalization-final-dataset.texttext-generation1K<n<10K0 likes52 downloads4mo agoHugging Face26cs-552-2026-databand /general_knowledge_dataset General Knowledge SFT Dataset This dataset contains the exact train and validation data used for the general knowledge LoRA SFT model in the MNLP project Specialize and Merge: Post Training Qwen3-1.7B for Multi Skill Reasoning. The dataset has two splits. Split Rows Purpose train 26,120 LoRA SFT training split valid 2,000 LoRA SFT validation split Sources The SFT data was built from six multiple-choice educational and science-oriented sources.… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_dataset.textquestion-answering10K<n<100K0 likes50 downloads4mo agoHugging Face27rlundqvist /vea-generalization-benchmark VEA-Generalization Benchmark A diagnostic set of matched response pairs to test whether a Reward Model's dispreference for verbalized evaluation-awareness (VEA) is broad (it penalizes any "I might be being tested" signal) or narrow (it mainly fires on the specific "Wood Labs" cue seen in training). Companion to rlundqvist/ifeval-obf-rl-preferences and the paper "LLM Judges Disprefer Evaluation Awareness." The idea Each item is a matched pair: an identical model… See the full description on the dataset page: https://huggingface.co/datasets/rlundqvist/vea-generalization-benchmark.texttext-classificationn<1K0 likes49 downloads26d agoHugging Face28generaleoley /manim-codegentext1K<n<10K11 likes48 downloads3y agoHugging Face29ademchaoua /GeneralTextCorpus Mixed Content Dataset Description:This dataset contains a diverse collection of text from multiple domains, including general knowledge, cooking, articles, and more. Each entry typically includes text content along with metadata such as source, title, and language. The dataset is structured to support research, analysis, or training of NLP models on varied textual content. Data Structure:Each item typically contains: id: Unique identifier text: Main text content meta: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/ademchaoua/GeneralTextCorpus.texttext-generation10K<n<100K0 likes48 downloads9mo agoHugging Face30visv-Bro /repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes48 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.