CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V ├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.imagequestion-answering100K<n<1M98 likes813 downloads1y agoHugging Face02SZLHOLDINGS /alloy-sovereign-eval-runs Alloy Sovereign Eval Runs · the honest first measured run Append-only measured eval runs produced by routing SZL's K-Verify Benchmark v1 through the live Alloy governed-inference stack on SZL's own sovereign metal (provider: sovereign, zero cloud, zero spend). Each row is one inference: its verdict, latency, NVML-measured energy, and a signed receipt id that is re-checkable against the live Alloy receipt chain. Built and maintained by SZL Holdings. Apache-2.0.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/alloy-sovereign-eval-runs.textquestion-answeringn<1K0 likes339 downloads2mo agoHugging Face03lodestones /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/ALLaVA-4V.imagequestion-answering1M<n<10M0 likes246 downloads2y agoHugging Face04agentlans /allenai-WildChat AllenAI WildChat Combined Dataset This unofficial repository provides the AllenAI WildChat Combined Dataset, which merges the WildChat-4.8M and WildChat-1M collections of human–ChatGPT conversations. WildChat-1M contains 1 million chats, of which 25.53% are from GPT‑4 and the remainder from GPT‑3.5. These conversations cover a wide range of complex interactions, including code-switching, ambiguity, and political topics. WildChat-4.8M originally comprised 4.8 million conversations.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-WildChat.texttext-generation1M<n<10M3 likes232 downloads10mo agoHugging Face05lyon-nlp /alloprofThis is a re-edit from the Alloprof dataset (which can be found here : https://huggingface.co/datasets/antoinelb7/alloprof). For more information about the data source and the features, please refer to the original dataset card made by the authors, along with their paper available here : https://arxiv.org/abs/2302.07738 This re-edition of the dataset is a preprocessed version of the original, in a more ready-to-use format. Essentially, the texts have been cleaned, and data not usable for… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/alloprof.texttext-classification10K<n<100K4 likes174 downloads2y agoHugging Face06annnli /TOFU-C-All TOFU: Task of Fictitious Unlearning 🍢 The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set. Quick Links Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/annnli/TOFU-C-All.textquestion-answering10K<n<100K0 likes126 downloads2y agoHugging Face07allenai /tulu-v2-sft-mixture-olmo-2048 Dataset Card for Tulu V2 Mix (2048 OLMo version) Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. This is a modified version of the Tulu V2 Mix used to train OLMo-Instruct. The two primary differences are: long conversations are resplit into 2048-token chunks, and the hardcoded subset has been replaced with similar examples about… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-2048.textquestion-answering100K<n<1M5 likes123 downloads2y agoHugging Face08LARK-Lab /EnvFactory-SFT-ALL EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL ## Overview EnvFactory-SFT-ALL is the complete supervised fine-tuning (SFT) dataset containing 26,500 tool-use trajectories synthesized using the EnvFactory framework. This dataset includes all generated trajectories before filtering. The dataset contains multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-ALL.texttext-generation10K<n<100K0 likes116 downloads4mo agoHugging Face09FreedomIntelligence /ALLaVA-4V-Chinese ALLaVA-4V for Chinese This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.imagequestion-answering100K<n<1M16 likes87 downloads2y agoHugging Face10Gyikoo /TOFU-C-All TOFU: Task of Fictitious Unlearning 🍢 The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set. Quick Links Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/Gyikoo/TOFU-C-All.textquestion-answering10K<n<100K0 likes85 downloads2y agoHugging Face11allenai /openscilm_queries Literature Synthesis Queries This dataset contains 50k real-world literature synthesis queries from our public demo. Dataset Summary This dataset contains real-world literature synthesis questions collected from users of a scientific question-answering system. Each entry includes: The user’s query (in natural language) (query) The subject of the question (e.g., computer science, medicine, engineering) (subject) The query intent (e.g., Literature Understanding, Paper… See the full description on the dataset page: https://huggingface.co/datasets/allenai/openscilm_queries.textquestion-answering10K<n<100K5 likes68 downloads1y agoHugging Face12allenai /tulu-v2-sft-mixture-olmo-4096 Dataset Card for Tulu V2 Mix (4096 OLMo version) Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. This is a modified version of the Tulu V2 Mix used to train newer (after April 2024) OLMo-SFT/Instruct variants (e.g. this model, or this one). The only difference is that the hardcoded subset (dataset='hard_coded') has been replaced… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-4096.textquestion-answering100K<n<1M0 likes65 downloads2y agoHugging Face13haiderkamal23 /allaM-offsec-arabic-chat-v2 Arabic Offensive Security Chat Dataset v2 High-quality category-aware bilingual Arabic/English dataset for offensive security assistants. What's New in v2 ✅ Category-aware responses: Different response structures for web vulns, DeFi, reconnaissance tools, social engineering, etc. ✅ No generic templates: Each category has specialized analysis framework ✅ No verbatim copying: Responses analyze and transform the input, not repeat it ✅ Semantic accuracy: Tools (nmap… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat-v2.textquestion-answering10K<n<100K0 likes56 downloads10mo agoHugging Face14allenai /tulu-v2-sft-long-mixtureThis is a recreation of the tulu-v2-sft-mixture, without splitting ShareGPT dataset into chunks of max 4096 tokens. This might be interesting to people who are doing long-context finetuning. Please refer to the original tulu-v2-sft-mixture for the details of this dataset mixture. License We are releasing this dataset under the terms of ODC-BY. By using this, you are also bound by the Common Crawl terms of use in respect of the content contained in the dataset. texttext-generation100K<n<1M7 likes54 downloads3y agoHugging Face15FreedomIntelligence /ALLaVA-4V-Arabic ALLaVA-4V for Arabic This is the Arabic version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Arabic through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of Hong… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Arabic.imagequestion-answering100K<n<1M4 likes49 downloads2y agoHugging Face16haiderkamal23 /allaM-offsec-arabic-chat Arabic Offensive Security Chat Dataset Bilingual Arabic/English dataset for training offensive security assistants. Dataset Details Training examples: 18,412 Validation examples: 2,000 Total: 20,412 Languages: Arabic (primary) + English (technical terms) Format: ChatML (messages field) Source: Filtered and processed from WNT3D/Ultimate-Offensive-Red-Team Intended Use This dataset is designed for fine-tuning models to assist with: Vulnerability analysis and… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat.textquestion-answering10K<n<100K0 likes32 downloads10mo agoHugging Face17PratikDhonde /letterboxd-all-movie-data Letterboxd Film Dataset This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings. Dataset Summary Total Films: 847,209 File Size: ~1.12 GB (1,120,572,122 bytes) Format: JSONL (JSON Lines) Language: Primarily English, with some multilingual content Data Structure Each line contains a JSON object with the following fields: { "url":… See the full description on the dataset page: https://huggingface.co/datasets/PratikDhonde/letterboxd-all-movie-data.imagetext-classification100K<n<1M1 likes28 downloads6mo agoHugging Face18lihaoxin2020 /evidence-subagent-sft-gpt54-single-all-jina-v1 GPT-5.4 Evidence Subagent SFT with Jina-refreshed Browse Outputs This dataset contains synthetic SFT conversations for training a small evidence execution subagent for deep-research systems. Each row is a single delegated evidence-gathering subtask derived from a full DR-Tulu trajectory. GPT-5.4 synthesized the delegated subtask, grouped original tool events into one or more batch tool-call turns, and wrote a structured cited evidence report. Tool outputs are reconstructed from… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/evidence-subagent-sft-gpt54-single-all-jina-v1.textquestion-answering10K<n<100K1 likes28 downloads5mo agoHugging Face19AllyArc /allyarc_oai_format Dataset Card for AllyArc/allyarc_oai_format This dataset card provides a structured overview of the AllyArc/allyarc_oai_format dataset, designed for training conversational AI models tailored for educational purposes, with a special focus on supporting students with diverse learning needs, including those in Special Educational Needs (SEN) education. Dataset Details Dataset Description The AllyArc/allyarc_oai_format dataset is comprised of conversational… See the full description on the dataset page: https://huggingface.co/datasets/AllyArc/allyarc_oai_format.textquestion-answering1K<n<10K0 likes22 downloads2y agoHugging Face20xzitao /All_university Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/xzitao/All_university.textquestion-answering100K<n<1M0 likes20 downloads1y agoHugging Face21YuYuanzi /pregnancy_all 🤰 孕期健康与临床知识库 (Pregnancy All) 这是一个专注于妇幼健康、产科临床及孕期护理的中文结构化数据集。数据涵盖了从备孕、孕期管理(早中晚三期)、分娩、产褥期护理到新生儿保健的全周期知识。 📋 数据集描述 本数据集整合了多个权威来源的信息,旨在为医疗 AI、智能问诊机器人及医学教育提供高质量语料。 数据来源与内容 数据主要包含以下几类信息: 临床指南:妊娠期糖尿病 (GDM)、高血压、贫血等并发症的诊疗规范。 孕期周报:按孕周划分的胎儿发育情况与母体变化指南。 法律法规:涉及母婴保健法、产假政策等国家政策文件。 中医保胎:包含中医辨证、药膳食疗(如砂仁鲫鱼汤)、穴位按摩等传统医学知识。 用药安全:孕期禁用与慎用药物清单。 数据格式 文件采用 JSONL 格式,每一行是一个独立的 JSON 对象。 主要字段包括: id: 唯一标识符 question: 问题或主题 answer: 详细解答或内容 source: 数据来源(如“国家卫健委”、“中医文献集”等)… See the full description on the dataset page: https://huggingface.co/datasets/YuYuanzi/pregnancy_all.textquestion-answering1K<n<10K1 likes20 downloads6mo agoHugging Face22cantonesesra /Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K Yue_WizardLMEvolved_AllAspectQA_Small_1.5K A specialized collection of high-quality question-answer pairs in Cantonese (粵語) inspired by the WizardLM evolution methodology, covering diverse and complex topics. Overview Yue_WizardLMEvolved_AllAspectQA_Small_1.5K is a curated dataset of 1,500 evolved question-answer pairs in Cantonese. This dataset applies the WizardLM evolution philosophy to generate in-depth, nuanced responses to complex questions in Cantonese. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K.texttext-generation1K<n<10K0 likes19 downloads1y agoHugging Face23cantonesesra /Cantonese_AllAspectQA_11K Cantonese_AllAspectQA_11K A comprehensive Question-Answer dataset in Cantonese (粵語) covering a wide range of conversational topics and aspects. Overview Cantonese_AllAspectQA_11K is a curated collection of 11,000 question-answer pairs in Cantonese, designed to facilitate the development, training, and evaluation of Cantonese language models and conversational AI systems. The dataset captures authentic Cantonese speech patterns, colloquialisms, and cultural nuances across… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_AllAspectQA_11K.texttext-generation10K<n<100K3 likes16 downloads1y agoHugging Face24anonymous-submission-678 /backtrader-mcq-base-pool-all-strategies Backtrader MCQ Benchmark This dataset contains multiple-choice questions for evaluating whether a model can reason about trading-strategy behavior using the Backtrader backtesting framework. Each question provides a complete backtest configuration and asks for a single answer choice in the format <<< X >>>, where X is one of A, B, C, or D. The primary evaluation file used in the paper is: backtrader_mcq_balanced_30_all_strategies.jsonl The larger supporting pool is:… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-678/backtrader-mcq-base-pool-all-strategies.textquestion-answeringn<1K0 likes8 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.