datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V
├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.alloy-sovereign-eval-runs
Alloy Sovereign Eval Runs · the honest first measured run
Append-only measured eval runs produced by routing SZL's
K-Verify Benchmark v1
through the live Alloy governed-inference stack on SZL's own sovereign
metal (provider: sovereign, zero cloud, zero spend). Each row is one
inference: its verdict, latency, NVML-measured energy, and a
signed receipt id that is re-checkable against the live Alloy receipt chain.
Built and maintained by SZL Holdings. Apache-2.0.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/alloy-sovereign-eval-runs.ALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/ALLaVA-4V.allenai-WildChat
AllenAI WildChat Combined Dataset
This unofficial repository provides the AllenAI WildChat Combined Dataset, which merges the WildChat-4.8M and WildChat-1M collections of human–ChatGPT conversations.
WildChat-1M contains 1 million chats, of which 25.53% are from GPT‑4 and the remainder from GPT‑3.5. These conversations cover a wide range of complex interactions, including code-switching, ambiguity, and political topics.
WildChat-4.8M originally comprised 4.8 million conversations.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-WildChat.alloprofThis is a re-edit from the Alloprof dataset (which can be found here : https://huggingface.co/datasets/antoinelb7/alloprof).
For more information about the data source and the features, please refer to the original dataset card made by the authors, along with their paper available here : https://arxiv.org/abs/2302.07738
This re-edition of the dataset is a preprocessed version of the original, in a more ready-to-use format. Essentially, the texts have been cleaned, and data not usable for… See the full description on the dataset page: https://huggingface.co/datasets/lyon-nlp/alloprof.TOFU-C-All
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/annnli/TOFU-C-All.tulu-v2-sft-mixture-olmo-2048
Dataset Card for Tulu V2 Mix (2048 OLMo version)
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This is a modified version of the Tulu V2 Mix used to train OLMo-Instruct.
The two primary differences are: long conversations are resplit into 2048-token chunks, and the hardcoded subset has been replaced with similar examples about… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-2048.EnvFactory-SFT-ALL
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL
## Overview
EnvFactory-SFT-ALL is the complete supervised fine-tuning (SFT) dataset containing 26,500 tool-use trajectories synthesized using the EnvFactory framework. This dataset includes all generated trajectories before filtering.
The dataset contains multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-ALL.ALLaVA-4V-Chinese
ALLaVA-4V for Chinese
This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR.
The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V.
Citation
If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.TOFU-C-All
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/Gyikoo/TOFU-C-All.openscilm_queries
Literature Synthesis Queries
This dataset contains 50k real-world literature synthesis queries from our public demo.
Dataset Summary
This dataset contains real-world literature synthesis questions collected from users of a scientific question-answering system.
Each entry includes:
The user’s query (in natural language) (query)
The subject of the question (e.g., computer science, medicine, engineering) (subject)
The query intent (e.g., Literature Understanding, Paper… See the full description on the dataset page: https://huggingface.co/datasets/allenai/openscilm_queries.tulu-v2-sft-mixture-olmo-4096
Dataset Card for Tulu V2 Mix (4096 OLMo version)
Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This is a modified version of the Tulu V2 Mix used to train newer (after April 2024) OLMo-SFT/Instruct variants (e.g. this model, or this one).
The only difference is that the hardcoded subset (dataset='hard_coded') has been replaced… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-4096.allaM-offsec-arabic-chat-v2
Arabic Offensive Security Chat Dataset v2
High-quality category-aware bilingual Arabic/English dataset for offensive security assistants.
What's New in v2
✅ Category-aware responses: Different response structures for web vulns, DeFi, reconnaissance tools, social engineering, etc.
✅ No generic templates: Each category has specialized analysis framework
✅ No verbatim copying: Responses analyze and transform the input, not repeat it
✅ Semantic accuracy: Tools (nmap… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat-v2.tulu-v2-sft-long-mixtureThis is a recreation of the tulu-v2-sft-mixture, without splitting ShareGPT dataset into chunks of max 4096 tokens. This might be interesting to people who are doing long-context finetuning.
Please refer to the original tulu-v2-sft-mixture for the details of this dataset mixture.
License
We are releasing this dataset under the terms of ODC-BY. By using this, you are also bound by the Common Crawl terms of use in respect of the content contained in the dataset.
ALLaVA-4V-Arabic
ALLaVA-4V for Arabic
This is the Arabic version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Arabic through ChatGPT and instructed ChatGPT not to translate content related to OCR.
The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V.
Citation
If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of Hong… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Arabic.allaM-offsec-arabic-chat
Arabic Offensive Security Chat Dataset
Bilingual Arabic/English dataset for training offensive security assistants.
Dataset Details
Training examples: 18,412
Validation examples: 2,000
Total: 20,412
Languages: Arabic (primary) + English (technical terms)
Format: ChatML (messages field)
Source: Filtered and processed from WNT3D/Ultimate-Offensive-Red-Team
Intended Use
This dataset is designed for fine-tuning models to assist with:
Vulnerability analysis and… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat.letterboxd-all-movie-data
Letterboxd Film Dataset
This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings.
Dataset Summary
Total Films: 847,209
File Size: ~1.12 GB (1,120,572,122 bytes)
Format: JSONL (JSON Lines)
Language: Primarily English, with some multilingual content
Data Structure
Each line contains a JSON object with the following fields:
{
"url":… See the full description on the dataset page: https://huggingface.co/datasets/PratikDhonde/letterboxd-all-movie-data.evidence-subagent-sft-gpt54-single-all-jina-v1
GPT-5.4 Evidence Subagent SFT with Jina-refreshed Browse Outputs
This dataset contains synthetic SFT conversations for training a small evidence
execution subagent for deep-research systems.
Each row is a single delegated evidence-gathering subtask derived from a full
DR-Tulu trajectory. GPT-5.4 synthesized the delegated subtask, grouped original
tool events into one or more batch tool-call turns, and wrote a structured
cited evidence report. Tool outputs are reconstructed from… See the full description on the dataset page: https://huggingface.co/datasets/lihaoxin2020/evidence-subagent-sft-gpt54-single-all-jina-v1.allyarc_oai_format
Dataset Card for AllyArc/allyarc_oai_format
This dataset card provides a structured overview of the AllyArc/allyarc_oai_format dataset, designed for training conversational AI models tailored for educational purposes, with a special focus on supporting students with diverse learning needs, including those in Special Educational Needs (SEN) education.
Dataset Details
Dataset Description
The AllyArc/allyarc_oai_format dataset is comprised of conversational… See the full description on the dataset page: https://huggingface.co/datasets/AllyArc/allyarc_oai_format.All_university
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/xzitao/All_university.pregnancy_all
🤰 孕期健康与临床知识库 (Pregnancy All)
这是一个专注于妇幼健康、产科临床及孕期护理的中文结构化数据集。数据涵盖了从备孕、孕期管理(早中晚三期)、分娩、产褥期护理到新生儿保健的全周期知识。
📋 数据集描述
本数据集整合了多个权威来源的信息,旨在为医疗 AI、智能问诊机器人及医学教育提供高质量语料。
数据来源与内容
数据主要包含以下几类信息:
临床指南:妊娠期糖尿病 (GDM)、高血压、贫血等并发症的诊疗规范。
孕期周报:按孕周划分的胎儿发育情况与母体变化指南。
法律法规:涉及母婴保健法、产假政策等国家政策文件。
中医保胎:包含中医辨证、药膳食疗(如砂仁鲫鱼汤)、穴位按摩等传统医学知识。
用药安全:孕期禁用与慎用药物清单。
数据格式
文件采用 JSONL 格式,每一行是一个独立的 JSON 对象。
主要字段包括:
id: 唯一标识符
question: 问题或主题
answer: 详细解答或内容
source: 数据来源(如“国家卫健委”、“中医文献集”等)… See the full description on the dataset page: https://huggingface.co/datasets/YuYuanzi/pregnancy_all.Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K
Yue_WizardLMEvolved_AllAspectQA_Small_1.5K
A specialized collection of high-quality question-answer pairs in Cantonese (粵語) inspired by the WizardLM evolution methodology, covering diverse and complex topics.
Overview
Yue_WizardLMEvolved_AllAspectQA_Small_1.5K is a curated dataset of 1,500 evolved question-answer pairs in Cantonese. This dataset applies the WizardLM evolution philosophy to generate in-depth, nuanced responses to complex questions in Cantonese. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_WizardLMEvolved_AllAspectQA_Small_1.5K.Cantonese_AllAspectQA_11K
Cantonese_AllAspectQA_11K
A comprehensive Question-Answer dataset in Cantonese (粵語) covering a wide range of conversational topics and aspects.
Overview
Cantonese_AllAspectQA_11K is a curated collection of 11,000 question-answer pairs in Cantonese, designed to facilitate the development, training, and evaluation of Cantonese language models and conversational AI systems. The dataset captures authentic Cantonese speech patterns, colloquialisms, and cultural nuances across… See the full description on the dataset page: https://huggingface.co/datasets/cantonesesra/Cantonese_AllAspectQA_11K.backtrader-mcq-base-pool-all-strategies
Backtrader MCQ Benchmark
This dataset contains multiple-choice questions for evaluating whether a model can reason about trading-strategy behavior using the Backtrader backtesting framework. Each question provides a complete backtest configuration and asks for a single answer choice in the format <<< X >>>, where X is one of A, B, C, or D.
The primary evaluation file used in the paper is:
backtrader_mcq_balanced_30_all_strategies.jsonl
The larger supporting pool is:… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-678/backtrader-mcq-base-pool-all-strategies.
