datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CoT-Reasoning-Instruct
reasoning-0.01 subset
synthetic dataset of reasoning chains for a wide variety of tasks.
we leverage data like this across multiple reasoning experiments/projects.
stay tuned for reasoning models and more data.
Thanks to Hive Digital Technologies (https://x.com/HIVEDigitalTech) for their compute support in this project and beyond.
medical-qa-shared-task-v1-toy
Dataset Card for "medical-qa-shared-task-v1-toy"
More Information needed
lmsys-chat-1m
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
This dataset contains one million real-world conversations with 25 state-of-the-art LLMs.
It is collected from 210K unique IP addresses in the wild on the Vicuna demo and Chatbot Arena website from April to August 2023.
Each sample includes a conversation ID, model name, conversation text in OpenAI API JSON format, detected language tag, and OpenAI moderation API tag.
User consent is obtained through the "Terms of… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/lmsys-chat-1m.AraSeg-2026-Shared-Task-NP
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NP.AraSeg-2026-Shared-Task-PA
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-PA.AraSeg-2026-Shared-Task-NoPnx-PA
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NoPnx-PA.AraSeg-2026-Shared-Task-NoPnx-NP
Arabic Sentence Segmentation Shared Task 2026
For details about the shared task, evaluation scripts, leaderboard, and submission guidelines, visit:
https://www.araseg.aramlab.ai/
Dataset Summary
AraSeg is the first comprehensive benchmark for Arabic sentence segmentation.
The corpus is designed to support research on sentence segmentation in Modern Standard Arabic (MSA), particularly in settings where punctuation is inconsistent, missing, or noisy.
AraSeg contains… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AraSeg-2026-Shared-Task-NoPnx-NP.aya_collection_telugumedical-qa-shared-task-v1-all
Dataset Card for "medical-qa-shared-task-v1-all"
More Information needed
Organic-Chemistry-VLM-PostTraining
Dataset Card for "Chemistry_text_to_image"
More Information needed
MeetingBank-QA-Summarization-Instruct
Dataset Card for MeetingBank-QA-Summary
This dataset is introduced in LLMLingua-2 (Pan et al., 2024) and is designed to assess the performance of compressed meeting transcripts on downstream tasks such as question answering (QA) and summarization.
It includes 862 meeting transcripts from the test set of meeting transcripts introduced in MeetingBank (Hu et al, 2023) as the context, togeter with QA pairs and summaries that were generated by GPT-4 for each context transcripts.… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/MeetingBank-QA-Summarization-Instruct.sharedtask_nlpai4health_filtereduber_text-Vision-QA
Dataset Card for "uber_text_qa"
More Information needed
LID201_Devanagari_Script_Languages_Identificationmedical-qa-shared-task-v1-half
Dataset Card for "medical-qa-shared-task-v1-half"
More Information needed
BAREC-Shared-Task-2026-sent
BAREC Shared Task 2026
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2026, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2026-sent.civil_comments_Safety
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/civil_comments_Safety.Sujet-Vision-QA
Dataset Description 📊🔍
The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering.
Key Features:
🖼️ 9,801 unique financial document images
❓ 107,050 question-answer pairs
🇬🇧 English language
📄 Diverse financial document… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/Sujet-Vision-QA.SNLI-NLI
Dataset Card for SNLI
Dataset Summary
The SNLI corpus (version 1.0) is a collection of 570k human-written English sentence pairs manually labeled for balanced classification with the labels entailment, contradiction, and neutral, supporting the task of natural language inference (NLI), also known as recognizing textual entailment (RTE).
Supported Tasks and Leaderboards
Natural Language Inference (NLI), also known as Recognizing Textual Entailment… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/SNLI-NLI.sharedtask_nlpai4health_information_extraction_filteredenglish-to-telugu-mtBAREC-Shared-Task-2026-doc
BAREC Shared Task 2026
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2026, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2026-doc.edition_0612_lavita-medical-qa-shared-task-v1-toy-readymade
edition_0612_lavita-medical-qa-shared-task-v1-toy-readymade
A Readymade by TheFactoryX
Original Dataset
lavita/medical-qa-shared-task-v1-toy
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0612_lavita-medical-qa-shared-task-v1-toy-readymade.edition_1117_lavita-medical-qa-shared-task-v1-toy-readymade
edition_1117_lavita-medical-qa-shared-task-v1-toy-readymade
A Readymade by TheFactoryX
Original Dataset
lavita/medical-qa-shared-task-v1-toy
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_1117_lavita-medical-qa-shared-task-v1-toy-readymade.Wiki2018_Devanagari_Script_Language_Identificationedition_0926_lavita-medical-qa-shared-task-v1-toy-readymade
edition_0926_lavita-medical-qa-shared-task-v1-toy-readymade
A Readymade by TheFactoryX
Original Dataset
lavita/medical-qa-shared-task-v1-toy
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0926_lavita-medical-qa-shared-task-v1-toy-readymade.edition_0954_lavita-medical-qa-shared-task-v1-toy-readymade
edition_0954_lavita-medical-qa-shared-task-v1-toy-readymade
A Readymade by TheFactoryX
Original Dataset
lavita/medical-qa-shared-task-v1-toy
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0954_lavita-medical-qa-shared-task-v1-toy-readymade.edition_0218_lavita-medical-qa-shared-task-v1-toy-readymade
edition_0218_lavita-medical-qa-shared-task-v1-toy-readymade
A Readymade by TheFactoryX
Original Dataset
lavita/medical-qa-shared-task-v1-toy
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0218_lavita-medical-qa-shared-task-v1-toy-readymade.edition_0257_lavita-medical-qa-shared-task-v1-toy-readymade
edition_0257_lavita-medical-qa-shared-task-v1-toy-readymade
A Readymade by TheFactoryX
Original Dataset
lavita/medical-qa-shared-task-v1-toy
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0257_lavita-medical-qa-shared-task-v1-toy-readymade.edition_0805_lavita-medical-qa-shared-task-v1-toy-readymade
edition_0805_lavita-medical-qa-shared-task-v1-toy-readymade
A Readymade by TheFactoryX
Original Dataset
lavita/medical-qa-shared-task-v1-toy
Process
This dataset is a "readymade" - inspired by Marcel Duchamp's concept of taking everyday objects and recontextualizing them as art.
What we did:
Selected the original dataset from Hugging Face
Shuffled each column independently
Destroyed all row-wise relationships
Preserved structure, removed meaning
The result:… See the full description on the dataset page: https://huggingface.co/datasets/TheFactoryX/edition_0805_lavita-medical-qa-shared-task-v1-toy-readymade.
