datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AA-LCR
Artificial Analysis Long Context Reasoning (AA-LCR) Dataset
AA-LCR includes 100 hard text-based questions that require reasoning across multiple real-world documents, with each document set averaging ~100k input tokens. Questions are designed such that answers cannot be directly retrieved from documents and must instead be reasoned from multiple information sources.
New in Version 1.1 (September 2026)
Sixteen corrected answer keys. Each one was re-verified… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-LCR.AA-Omniscience-Public
Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies.
Leaderboard and detailed results
Paper
Introduction
We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.Artificial-intelligence-dataset-for-IR-systems
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
information-retrieval
semantic-search
Languages
English
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Adel-Elwan/Artificial-intelligence-dataset-for-IR-systems.ArtificialIntelligenceEthics
ArtificialIntelligenceEthics
tags: AI ethics, classification, text analysis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'ArtificialIntelligenceEthics' dataset contains a collection of text passages discussing various ethical issues related to artificial intelligence. Each passage has been preprocessed to remove any personally identifiable information, ensuring privacy and compliance with data protection regulations. The… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ArtificialIntelligenceEthics.ArtificiallyNoisySpeechTranscriptionsThis dataset contains transcriptions of speech files derived from the Norwegian language corpus provided by Språkbanken, specifically the nb_samtale subset. These transcriptions have been subjected to controlled noise addition to simulate various acoustic environments.
URI: https://huggingface.co/datasets/Sprakbanken/nb_samtale/viewer/annotations/train?f[duration][min]=24.6432&f[duration][imax]=27.368
Original Audio properties:
Duration : 24 secounds - 27 secounds
Format: WAV
Number of… See the full description on the dataset page: https://huggingface.co/datasets/kjetMol/ArtificiallyNoisySpeechTranscriptions.Artificial_Intelligence
[!NOTE]
Dataset origin: https://www.eurotermbank.com/collections/1035
ted
