datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IMF-Reports
IMF Technical Assistance Reports — Recommendation Process Corpus
A page-grounded research corpus of 780 IMF technical-assistance report
records. It contains source PDFs, layout-aware Markdown, page-level text, extracted
visuals, metadata, observations, recommendations, and labeled links between observations
and recommendations.
Required acknowledgement
All research, publications, datasets, models, applications, or other work derived from
this corpus should… See the full description on the dataset page: https://huggingface.co/datasets/FrenchCastle/IMF-Reports.corral-QAs-reports
Corral – QA Reports
Model completions for question-answer evaluations probing factual knowledge and reasoning across Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the model completions and reports for the question-answer evaluations used to test the factual knowledge and reasoning ability of models across Corral… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports.McKinsey-Reportsmeta-llama/synthetic-data-kit
https://github.com/meta-llama/synthetic-data-kit
McKinsey reports
https://www.mckinsey.com/featured-insights/insights-store
Synthetic_PenTest_ReportsThe full CJ Jones' synthetic dataset catalog is available at:
https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
📄 100 Samples of Synthetic Automated Penetration Test Reports
This dataset contains 100+ realistic, synthetic penetration testing reportsstructured to simulate professional internal security assessments. Each record models the full flow of a pentest engagement, including:
Reconnaissance / Discovery Phase
Vulnerability Assessment… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_PenTest_Reports.corral-QAs-topic_reports
Corral – QA Topic Reports
Averaged QA results for factual-knowledge and reasoning evaluations across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments.
The… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-topic_reports.Chameleon-Radiology-Reportspubmed_case_reports
PubMed Case Reports
A collection of 13,989 full-text case reports from the PubMed Central (PMC) Open Access subset, spanning 2005–2025. Each article includes structured metadata, abstract, full body text, and section-level annotations. This dataset is designed for medical NLP, clinical reasoning, and biomedical text mining.
Dataset Description
Summary
This dataset comprises case reports published in peer-reviewed medical journals, sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/awinml/pubmed_case_reports.french-senate-session-reports
🏛️ French Senate Session Reports Dataset
A dataset of parliamentary debates and sessions reports from the French Senate.508,647,861 tokens of high-quality French text transcribed manually from Senate Sessions
Description
This dataset consists of all session reports from the French Senate debates, crawled from the official website senat.fr. It provides high-quality text data of parliamentary discussions, covering a wide range of political, economic, and social topics… See the full description on the dataset page: https://huggingface.co/datasets/TheJeanneCompany/french-senate-session-reports.
