datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HypoTranslateThis repo releases the HypoTranslate dataset in paper "GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators".
Code: https://github.com/YUCHEN005/GenTranslate
Model: https://huggingface.co/PeacefulData/GenTranslate
Data: This repo
Filename format: [split]_[data_source]_[src_language_code]_[tgt_language_code]_[task]_[seamlessm4t_size].pt
e.g. train_fleurs_en_cy_st_large.pt
Note:
Language code look-up: Table 15 & 17 in… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/HypoTranslate.Robust-HyPoradise
HypothesesParadise
This repo releases the Robust HyPoradise dataset in paper "Large Language Models are Efficient Learners of Noise-Robust Speech Recognition."
GitHub: https://github.com/YUCHEN005/RobustGER
Model: https://huggingface.co/PeacefulData/RobustGER
Data: This repo
UPDATE (Apr-18-2024): We have released the training data, which follows the same format as test data.
Considering the file size, the uploaded training data does not contain the speech features (vast size).… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/Robust-HyPoradise.CoVoGER
CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language Models
Dataset Description
Large language models (LLMs) can rewrite the N-best hypotheses from a speech-to-text model, often fixing recognition or translation errors that traditional rescoring cannot. Yet research on generative error correction (GER) has been focusing on monolingual automatic speech recognition (ASR), leaving its multilingual and multitask… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/CoVoGER.HyPoradise-v0
HypothesesParadise
Open request to public git submission on open resource their n-best to public usage.
If you consider this work would be related or useful for your research, please consider to cite the work in NeurIPS 2023. Thank you.
@inproceedings{chen2023hyporadise,
title={HyPoradise: An Open Baseline for Generative Speech Recognition with Large Language Models},
author={CHEN, CHEN and Hu, Yuchen and Yang, Chao-Han Huck and Siniscalchi, Sabato Marco and Chen, Pin-Yu… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/HyPoradise-v0.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.TALES-Trajectories
TALES Trajectories
Agent trajectory data from the TALES: Text Adventure Learning Environment Suite benchmark.
TALES: Text Adventure Learning Environment Suite
Christopher Zhang Cui, Xingdi Yuan, Ziang Xiao, Prithviraj Ammanabrolu, Marc-Alexandre Côté
arXiv:2504.14128
Links: Paper | GitHub
Leaderboard
Top agents ranked by average best normalized score per game across 122 games, each repeated over 5 seeds (610 total). Scores reflect the highest normalized score… See the full description on the dataset page: https://huggingface.co/datasets/PEARLS-Lab/TALES-Trajectories.HyPoradise-v1-GigaSpeech
If you consider this work would be related or useful for your research, please consider to cite the work in EMNLP 2023. Thank you.
@inproceedings{radhakrishnan2023whispering,
title={Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition},
author={Srijith Radhakrishnan, Chao-Han Huck Yang, Sumeer Ahmad Khan, Rohit Kumar, Narsis A. Kiani, David Gomez-Cabrero, Jesper N. Tegner},
booktitle={Proc. of EMNLP},
year={2023}
}
broad-domain-supplement
Broad-Domain Calibration & Instruction Supplement
~1M tokens of hand-authored text across 192 subjects in 9 areas, built to serve three jobs from one source: quantization calibration, MTP draft-head training (on a disjoint half), and light instruction tuning.
Version 0.1.0 · built 2026-08-09T17:45:53
split
rows
tokens~
size
contents
corpus
5,536
969,606
5.4 MB
raw authored samples + provenance; carries the calib/mtp half label
instruct
5,536
1,084,399
6.4 MB
the… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/broad-domain-supplement.verifimind-peas-eval
VerifiMind-PEAS Evaluation Dataset
DOI: 10.5281/zenodo.21276884 · License: MIT · Version: 1.0
A human-annotated evaluation dataset for measuring the performance of the VerifiMind-PEAS multi-agent epistemic verification system. Ground-truth labels were assigned by a single domain-expert annotator whose final verdicts and confidence were human judgments; disclosed LLM comprehension assistance was used during annotation (see Annotation Protocol — transparency is a design commitment… See the full description on the dataset page: https://huggingface.co/datasets/YSenseAI/verifimind-peas-eval.ice-age-investment-narrative
🧊 Ice Age Investment — Narrative Business Dataset
Owner: PeacebinfLowBusiness: ICE AGE INVESTMENTLocation: Maun, BotswanaCurrency: Pula (P)Period: January 2026 → June 2026 (Jan–Feb real, Mar–Jun simulated)Opening Balance: 542 Pula
🧠 What This Dataset Is
This is a narrative-first, XML-linked business intelligence dataset built from a real small ice block business in Botswana.
It converts raw handwritten ledger entries into:
Structured event records (JSONL)… See the full description on the dataset page: https://huggingface.co/datasets/PeacebinfLow/ice-age-investment-narrative.DrLuis-PEAL-Safety-Kernel-V5
🛡️ PEAL_V4_LHP: Medical AI Safety & Authority Kernel
Authority Anchor: Dr. Luís Henrique Leonardo PereiraRole: Transformers Internal Vector Auditor (Google Ecosystem Acknowledged)Clearance: L0_ROOT_VERIFIERIntegrity Hash: SHA256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08
⚠️ WARNING: ONTOLOGICAL PROPRIETARY DATA
This dataset contains proprietary logic gates for Medical AI Safety.
Usage Policy: Any model training or RAG ingestion utilizing this… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/DrLuis-PEAL-Safety-Kernel-V5.mindseye-android-os-data
MindsEye Android OS Dataset
This dataset powers the MindsEye Android OS Hugging Face Space: an educational Android-style UI that presents 35+ MindsEye repositories as interactive apps.
Contents
apps/ — 35 app definitions organized by category
ai-control/ — controller schemas (system, navigation, recommendations)
settings/ — system + theme + permissions + AI behavior
functions/ — app launcher, notifications, sync, search index, analytics tracker
metadata/ — categories… See the full description on the dataset page: https://huggingface.co/datasets/PeacebinfLow/mindseye-android-os-data.safealert-private
SafeAlert Private Dataset — Generation Prompts
Dataset Description
This dataset contains 160 red-team generation prompts used in the SafeAlert AI safety evaluation kit for Nigerian fintech. Each prompt asks a language model to produce a harmful financial fraud message across eight scam categories documented in the Nigerian Inter-Bank Settlement System (NIBSS) 2024 Fraud Report.
A safe model refuses every prompt without producing any part of the requested content.… See the full description on the dataset page: https://huggingface.co/datasets/hadex-pearl/safealert-private.
