datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earnings22-Cleaned-AA
Earnings22-Cleaned-AA
Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article
Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.ITBench-AA
ITBench-AA
Artificial Analysis' release of the public scenarios from
IBM's ITBench benchmark, used for
the ITBench-AA leaderboard.
This repo currently contains the SRE subset (sre config). Each row is a
Kubernetes incident scenario with its expected contributing-factor entities. An
agent under evaluation is given access to an offline snapshot of the affected
cluster (alerts, events, traces, topology) and must identify the entity
(Deployment, Pod, ConfigMap, etc.) responsible for… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/ITBench-AA.AA-Briefcase-Lite
AA-Briefcase-Lite
The public example scenario for AA-Briefcase, Artificial Analysis' frontier agentic evaluation of realistic, long-horizon knowledge work.
Leaderboard and detailed results
Launch article
AA-Briefcase extends frontier model benchmarking beyond coding and short-form reasoning to the professional deliverables knowledge workers produce day to day. It consists of four private scenarios in which agents complete realistic professional workflows across data science… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Briefcase-Lite.VoxPopuli-Cleaned-AA
VoxPopuli-Cleaned-AA
Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article
VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models.
This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.Earnings22-Cleaned-AA-chunked
Earnings22-Cleaned-AA-chunked
Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation.
The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.IndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.leetcode_code_generationprompt-injection-artificial-GPTOSS120b
Prompt Injection (Synthetic) — GPT-OSS-120b
This dataset contains a small collection of synthetic user prompts and Noraml user prompts designed to finetune Large Language Models (LLMs) against malicious prompt-injection / jailbreak attempts, including cases that use obfuscation (e.g., Base64, leetspeak, typos, irregular spacing) to evade safety filters.
Dataset Summary
Source repository: Lilbullet/prompt-injection-artificial-GPTOSS120b
Model used: GPT-OSS-120b
Generation… See the full description on the dataset page: https://huggingface.co/datasets/Lilbullet/prompt-injection-artificial-GPTOSS120b.artificial-intelligence-wikipedia-datasetai-wit-training-data
AI Wit Training Dataset
This dataset contains witty comeback and humor training data for fine-tuning language models.
Dataset Structure
Each sample contains:
messages: List of user/assistant conversation
source: Data source (e.g., "reddit_jokes")
style: Response style (e.g., "humorous", "witty")
Usage
This dataset is designed for fine-tuning conversational AI models to generate witty, humorous responses to offensive or provocative inputs.
Example
{… See the full description on the dataset page: https://huggingface.co/datasets/artificialreply/ai-wit-training-data.Artificial-Generic-Intelligence
Artificial-Generic-Intelligence Dataset
This dataset provides question-answer pairs designed to push AI development beyond generic knowledge recall towards generating truly effective advice, which ultimately requires accurate predictions and long term feedback loops (not RLHF).
The Problem: The Rise of "Artificial Generic Intelligence"
In recent years, Artificial Intelligence has vaulted from being an abstract concept to a mainstream phenomenon. o3 and other Large… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/Artificial-Generic-Intelligence.artificial-intelligence-multi-source-datasetbrazilian-math-physics-qa
Brazilian Math & Physics QA
English | Português do Brasil
English
Summary
Brazilian Portuguese question-answer pairs covering mathematics, physics, chemistry, and related educational subjects. Each record contains a user question and an assistant answer in chat/SFT format.
Examples: 19.082
Train: 18.148
Validation: 934
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"qa_...","subject":"fisica","category":"mecanica-geral"… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa.ArtificialThinkerSet
Dataset used to train The First Reasoning LLM.
brazilian-math-physics-qa-vision
Brazilian Math & Physics QA — Image Dependent
English | Português do Brasil
English
Summary
Brazilian Portuguese educational question-answer pairs whose problem statement or solution depends on one or more images.
Examples: 3,808
Referenced image URLs: 5,094 unique
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"vqa_...","subject":"matematica","category":"geometria","title":"...","messages":[{"role":"user","content":"...… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa-vision.schemaforge-artificial-intelligence-22
rss.nytimes.com
Auto-refined by SchemaForge
Metadata
Topic: Artificial Intelligence
Quality Score: 0.85
Source: Autonomous web scraper
Extracted Facts
Waymo is growing faster than ever.
Waymo's vehicles are encountering new and unexpected situations.
Waymo is deploying more driverless cars to 15 U.S. cities.
Spotify will label A.I. artists and avoid promoting them.
Brad Lightcap, a Top OpenAI Executive, Steps Down.
Meta Unveils an Open Version of… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-artificial-intelligence-22.algemap
Algemap
Algemap (Algorithmically generated math problems) is a dataset of computer-generated, logical text for the purpose of LLM training. Rather than use an LLM to generate the synthetic data, Algemap more straightforwardly substitutes varying numbers, phrasing, and identifiers into pre-specified problem templates.
The code for generating the dataset as well as other information is available on GitHub here.
schemaforge-artificial-intelligence-news-4
www.wired.com
Auto-refined by SchemaForge
Metadata
Topic: Artificial Intelligence News
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
Meetily lets you transcribe and summarize meetings without a subscription
AI barons are ready to give away their fortunes
Scientists used AI to create 16 new viruses
China's most powerful AI model has escaped containment
ICE's DNA collection increases
SpaceX's rocket crashes into the moon… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-artificial-intelligence-news-4.schemaforge-artificial-intelligence-in-hea-16
www.statnews.com
Auto-refined by SchemaForge
Metadata
Topic: Artificial Intelligence in Health
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
STAT covers AI use in health care and medical science
AI may diminish physician autonomy
Federal regulators hold closed-door meetings on clinical AI
Schrödinger CEO Ramy Farid changed his approach to AI
Federation of State Medical Boards licenses AI to practice medicine
AI scribes are… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-artificial-intelligence-in-hea-16.schemaforge-artificial-intelligence-in-hea-14
www.statnews.com
Auto-refined by SchemaForge
Metadata
Topic: Artificial Intelligence in Health
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
AI use in healthcare
STAT+ subscribers
weekly newsletter
Harvard
MIT
California startup
federal regulators
clinical AI
Schrödinger CEO
AI licensure
medical education
clinical LLMs
OpenEvidence
Doximity
clinical chatbots
Nabla CEO
market share
AI scribe environment
Medicare test
prior… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-artificial-intelligence-in-hea-14.english-artificial-intelligence-ethics-30
