datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fiqa
FiQA Dataset for RAG Evaluation
The FiQA (Financial Opinion Mining and Question Answering) dataset reformatted specifically for evaluating Retrieval-Augmented Generation (RAG) systems. This dataset contains financial domain questions with ground truth answers and retrieved contexts, making it ideal for testing RAG pipelines on domain-specific content.
Recommended Usage: ragas_eval_v3
The ragas_eval_v3 configuration is the primary and recommended way to use this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/vibrantlabsai/fiqa.TTCW-Based-Review
TTCW Creative Writing Evaluation Dataset
If you use this dataset in your research, please cite our paper — it helps support ongoing academic work. Citation details are at the bottom of this page.
Dataset Description
Summary
A supervised fine-tuning (SFT) dataset for training LLMs to act as creative writing evaluators. Each example contains a creative story and four message-format columns representing different evaluation objectives — from… See the full description on the dataset page: https://huggingface.co/datasets/VibrantVista/TTCW-Based-Review.amnesty_qa
Amnesty QA Dataset
A grounded question-answering dataset for evaluating RAG (Retrieval-Augmented Generation) systems, created from reports collected from Amnesty International.
This dataset is designed for testing and evaluating RAG pipelines with real-world human rights content.
Dataset Structure
Each sample contains:
user_input: The question to be answered
reference: Ground truth answer for evaluation
response: Generated answer from the system
retrieved_contexts: List… See the full description on the dataset page: https://huggingface.co/datasets/vibrantlabsai/amnesty_qa.hindi-wikipediaenterprise-worlds
Enterprise-Worlds
Executable enterprise environments for measuring agents on operational work — the layer that turns
an agent's output into a durable outcome. Each world ships persistent state, typed tools, a written
policy, and a simulated colleague who discloses information only when asked. Reward is read from the
final state of the world, not from the transcript.
This repository hosts the datasets. The environment and evaluator live at
vibrantlabsai/Enterprise-Worlds.… See the full description on the dataset page: https://huggingface.co/datasets/vibrantlabsai/enterprise-worlds.WikiEval
WikiEval
Dataset for to do correlation analysis of difference metrics proposed in Ragas
This dataset was generated from 50 pages from Wikipedia with edits post 2022.
Column description
question: a question that can be answered from the given Wikipedia page (source).
source: The source Wikipedia page from which the question and context are generated.
grounded_answer: answer grounded on context_v1
ungrounded_answer: answer generated without context_v1
poor_answer: answer… See the full description on the dataset page: https://huggingface.co/datasets/vibrantlabsai/WikiEval.style-judge-dataset
Style Judge Dataset
A pairwise dataset for learning a continuous style-similarity function while controlling for topic, introduced in Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning (arXiv:2512.05747).
Dataset Summary
Total rows: 156k (default subset)
Splits: train 130k, validation 13k, test 13k
Format: Arrow
Columns
sentence1 (string): original chunk text
sentence2 (string): refilled chunk text
score (float): calibrated… See the full description on the dataset page: https://huggingface.co/datasets/VibrantVista/style-judge-dataset.ragas-wikiqa
Dataset Card for "ragas-wikiqa"
More Information needed
story-style-SFT-dataset
Citation
BibTeX:
@misc{liu2025capturingclassicauthorialstyle,
title={Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning},
author={Jinlong Liu and Mohammed Bahja and Venelin Kovatchev and Mark Lee},
year={2025},
eprint={2512.05747},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2512.05747},
}
wiki-eval
Dataset Card for "wiki-eval"
More Information needed
tau2-infinity-dag
tau2-infinity
An adaptive benchmark for evaluating LLM tool-use agents on airline customer service tasks. Generated using EnvScaler by VibrantLabs.
Overview
Each task requires an agent to transform an initial database state S_0 into a golden final state S* by executing a sequence of tool calls (flight searches, bookings, cancellations, updates, etc.). Tasks were adaptively generated to target specific difficulty levels against a calibration model.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/vibrantlabsai/tau2-infinity-dag.grpo-style-training
GRPO Style Training Dataset
This dataset contains style-specific training and testing data for classic authors, used in the research paper "Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning".
The data is designed to train models to mimic specific literary styles using Group Relative Policy Optimization (GRPO).
Dataset Structure
The data is organized by author folders. Inside each author's folder, there are separate train and test… See the full description on the dataset page: https://huggingface.co/datasets/VibrantVista/grpo-style-training.ragas-webgpt
Dataset Card for "ragas-webgpt"
More Information needed
masakhaner2
MasakhaNER2
...
annotations_creators:
expert-generated
language:
bm
bbj
ee
fon
ha
ig
rw
lg
luo
mos
ny
pcm
sn
sw
tn
tw
wo
xh
yo
zu
language_creators:
expert-generated
license:
afl-3.0
multilinguality:
multilingual
pretty_name: masakhaner2.0
size_categories:
1K<n<10K
source_datasets:
original
tags:
ner
masakhaner
masakhane
task_categories:
token-classification
task_ids:
named-entity-recognition
Dataset Card for [Dataset Name]
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/Vibrant8849/masakhaner2.enterprise-ops-gym-plusyann-lecun-wisdomeli5-test
Dataset Card for "eli5-test"
More Information needed
earning_report_summaryphysics_metrics_alignment
Dataset Card for "physics_metrics_alingment"
More Information needed
qrecc_conversational_embeddings
Dataset Card for "qrecc_conversational_embeddings"
More Information needed
ELI5Sample_non_english_corpusdiabetes_assistant_datasettau2-infinity-wg
tau2-infinity-wg
Airline customer-service tasks mined by world-gen adversarial failure-search — each task is an artifact on which a target model diverges from an oracle model. Companion to vibrantlabsai/tau2-infinity; consumed by the tau2_infinity_wg Prime Intellect RL environment.
Overview
Unlike tau2-infinity (adaptively generated toward a target difficulty band), tasks here are discovered by hypothesis-disproof: generate a hypothesized failure mode, confirm the oracle… See the full description on the dataset page: https://huggingface.co/datasets/vibrantlabsai/tau2-infinity-wg.aspect_critic_answer_correctnessVibrantValentineVistas
VibrantValentineVistas
tags: coloring page, vibrant, Valentine's Day
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'VibrantValentineVistas' dataset is designed for ML practitioners seeking a high-quality, detailed outline of a coloring page. It features a whimsical tomte gnome, a folk figure from Scandinavian folklore, adorned with a vibrant, colorful background filled with Valentine's Day themed elements such as hearts… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/VibrantValentineVistas.VibrantIMplanbench_vibrantlabsai_ragas
