datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
syntheticDocQA_artificial_intelligence_test
Dataset Description
This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications.
It includes documents about the Artificial Intelligence.
Data Collection
Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('artificial intelligence'). From these documents, we randomly sampled 1000 pages.
We associated these with 100 questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/vidore/syntheticDocQA_artificial_intelligence_test.solar-MID-descriptorsdocqa_artificial_intelligence_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence_beir.IndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.solar-MFI-imagesIndustryCorpus2_artificial_intelligence_machine_learning
IndustryCorpus2: Artificial Intelligence
This repository contains the IndustryCorpus2: Artificial Intelligence domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_artificial_intelligence_machine_learning.africa-synth-telecom-artificial-intelligence-automation-nigeria
Africa Synth Telecom Artificial Intelligence Automation Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-telecom-artificial-intelligence-automation-nigeria.syntheticDocQA_artificial_intelligence_test_tesseractDoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence
DoD Public Affairs Use of Artificial Intelligence
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 150 document-grounded question-and-answer records based on DoD Instruction 5400.19, “Public Affairs Use of Artificial Intelligence,” effective July 28, 2025.
The source establishes Department of Defense policy, responsibilities, and procedures for the appropriate use of artificial-intelligence capabilities in… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence.artificial-intelligence-wikipedia-datasetArtificial-Generic-Intelligence
Artificial-Generic-Intelligence Dataset
This dataset provides question-answer pairs designed to push AI development beyond generic knowledge recall towards generating truly effective advice, which ultimately requires accurate predictions and long term feedback loops (not RLHF).
The Problem: The Rise of "Artificial Generic Intelligence"
In recent years, Artificial Intelligence has vaulted from being an abstract concept to a mainstream phenomenon. o3 and other Large… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/Artificial-Generic-Intelligence.Artificial-intelligence-dataset-for-IR-systems
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
information-retrieval
semantic-search
Languages
English
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Adel-Elwan/Artificial-intelligence-dataset-for-IR-systems.artificial-intelligence-multi-source-datasetsyntheticDocQA_artificial_intelligence_test_ocr_chunksyntheticDocQA_artificial_intelligence_test_captioningdocqa_artificial_intelligence
Creation
This dataset is build upon the corresponding dataset from the ViDoRe Benchmark. For more information regarding the filtering please read our paper or this discussion on github.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence.africa-synth-telecom-artificial-intelligence-automation-nigeria
Artificial Intelligence Automation Dataset (TsFile)
This dataset is an Apache TsFile conversion of electricsheepafrica/africa-synth-telecom-artificial-intelligence-automation-nigeria, a synthetic Nigerian telecom dataset about self-optimizing network operations with AI recommendations and outcomes.
Source Dataset
Original dataset: electricsheepafrica/africa-synth-telecom-artificial-intelligence-automation-nigeria
Dataset title: Artificial Intelligence Automation… See the full description on the dataset page: https://huggingface.co/datasets/THULab/africa-synth-telecom-artificial-intelligence-automation-nigeria.algemap
Algemap
Algemap (Algorithmically generated math problems) is a dataset of computer-generated, logical text for the purpose of LLM training. Rather than use an LLM to generate the synthetic data, Algemap more straightforwardly substitutes varying numbers, phrasing, and identifiers into pre-specified problem templates.
The code for generating the dataset as well as other information is available on GitHub here.
schemaforge-artificial-intelligence-news-4
www.wired.com
Auto-refined by SchemaForge
Metadata
Topic: Artificial Intelligence News
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
Meetily lets you transcribe and summarize meetings without a subscription
AI barons are ready to give away their fortunes
Scientists used AI to create 16 new viruses
China's most powerful AI model has escaped containment
ICE's DNA collection increases
SpaceX's rocket crashes into the moon… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-artificial-intelligence-news-4.schemaforge-artificial-intelligence-22
rss.nytimes.com
Auto-refined by SchemaForge
Metadata
Topic: Artificial Intelligence
Quality Score: 0.85
Source: Autonomous web scraper
Extracted Facts
Waymo is growing faster than ever.
Waymo's vehicles are encountering new and unexpected situations.
Waymo is deploying more driverless cars to 15 U.S. cities.
Spotify will label A.I. artists and avoid promoting them.
Brad Lightcap, a Top OpenAI Executive, Steps Down.
Meta Unveils an Open Version of… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-artificial-intelligence-22.huberman_lab_Marc_Andreessen_How_Risk_Taking_Innovation__Artificial_Intelligence_Transform_Humanschemaforge-artificial-intelligence-in-hea-16
www.statnews.com
Auto-refined by SchemaForge
Metadata
Topic: Artificial Intelligence in Health
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
STAT covers AI use in health care and medical science
AI may diminish physician autonomy
Federal regulators hold closed-door meetings on clinical AI
Schrödinger CEO Ramy Farid changed his approach to AI
Federation of State Medical Boards licenses AI to practice medicine
AI scribes are… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-artificial-intelligence-in-hea-16.schemaforge-artificial-intelligence-in-hea-14
www.statnews.com
Auto-refined by SchemaForge
Metadata
Topic: Artificial Intelligence in Health
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
AI use in healthcare
STAT+ subscribers
weekly newsletter
Harvard
MIT
California startup
federal regulators
clinical AI
Schrödinger CEO
AI licensure
medical education
clinical LLMs
OpenEvidence
Doximity
clinical chatbots
Nabla CEO
market share
AI scribe environment
Medicare test
prior… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-artificial-intelligence-in-hea-14.syntheticDocQA_artificial_intelligence_test
Dataset Description
This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications.
It includes documents about the Artificial Intelligence.
Data Collection
Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('artificial intelligence'). From these documents, we randomly sampled 1000 pages.
We associated these with 100 questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/Madhu348/syntheticDocQA_artificial_intelligence_test.english-artificial-intelligence-ethics-30asia-owid-artificial-intelligence-patents-submitted
Artificial Intelligence Patents Submitted | Asia (Our World in Data)
🌏 71 observations · 18 Asia countries · 2016–2021 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 71 observations of Artificial Intelligence Patents Submitted data across 18 Asia countries, spanning 2016–2021.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Artificial Intelligence Patents Submitted… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-artificial-intelligence-patents-submitted.africa-owid-annual-scholarly-publications-on-artificial-intelligence
Annual Scholarly Publications On Artificial Intelligence | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-annual-scholarly-publications-on-artificial-intelligence.asia-owid-private-investment-in-artificial-intelligence-cset
Private Investment In Artificial Intelligence Cset | Asia (Our World in Data)
🌏 260 observations · 36 Asia countries · 2016–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 260 observations of Private Investment In Artificial Intelligence Cset data across 36 Asia countries, spanning 2016–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Private Investment In Artificial… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-private-investment-in-artificial-intelligence-cset.
