CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PeacefulData /HypoTranslateThis repo releases the HypoTranslate dataset in paper "GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators". Code: https://github.com/YUCHEN005/GenTranslate Model: https://huggingface.co/PeacefulData/GenTranslate Data: This repo Filename format: [split]_[data_source]_[src_language_code]_[tgt_language_code]_[task]_[seamlessm4t_size].pt e.g. train_fleurs_en_cy_st_large.pt Note: Language code look-up: Table 15 & 17 in… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/HypoTranslate.text-generation100K<n<1M2 likes6k downloads2y agoHugging Face02PeacefulData /Robust-HyPoradise HypothesesParadise This repo releases the Robust HyPoradise dataset in paper "Large Language Models are Efficient Learners of Noise-Robust Speech Recognition." GitHub: https://github.com/YUCHEN005/RobustGER Model: https://huggingface.co/PeacefulData/RobustGER Data: This repo UPDATE (Apr-18-2024): We have released the training data, which follows the same format as test data. Considering the file size, the uploaded training data does not contain the speech features (vast size).… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/Robust-HyPoradise.text-generation100K<n<1M4 likes2.3k downloads2y agoHugging Face03PeacefulData /CoVoGER CoVoGER: A Multilingual Multitask Benchmark for Speech-to-text Generative Error Correction with Large Language Models Dataset Description Large language models (LLMs) can rewrite the N-best hypotheses from a speech-to-text model, often fixing recognition or translation errors that traditional rescoring cannot. Yet research on generative error correction (GER) has been focusing on monolingual automatic speech recognition (ASR), leaving its multilingual and multitask… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/CoVoGER.automatic-speech-recognition1 likes1.4k downloads6mo agoHugging Face04PeacefulData /HyPoradise-v0 HypothesesParadise Open request to public git submission on open resource their n-best to public usage. If you consider this work would be related or useful for your research, please consider to cite the work in NeurIPS 2023. Thank you. @inproceedings{chen2023hyporadise, title={HyPoradise: An Open Baseline for Generative Speech Recognition with Large Language Models}, author={CHEN, CHEN and Hu, Yuchen and Yang, Chao-Han Huck and Siniscalchi, Sabato Marco and Chen, Pin-Yu… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/HyPoradise-v0.text-generation10M<n<100M3 likes494 downloads3y agoHugging Face05MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes473 downloads10mo agoHugging Face06PEARLS-Lab /TALES-Trajectories TALES Trajectories Agent trajectory data from the TALES: Text Adventure Learning Environment Suite benchmark. TALES: Text Adventure Learning Environment Suite Christopher Zhang Cui, Xingdi Yuan, Ziang Xiao, Prithviraj Ammanabrolu, Marc-Alexandre Côté arXiv:2504.14128 Links: Paper | GitHub Leaderboard Top agents ranked by average best normalized score per game across 122 games, each repeated over 5 seeds (610 total). Scores reflect the highest normalized score… See the full description on the dataset page: https://huggingface.co/datasets/PEARLS-Lab/TALES-Trajectories.tabulartext-generation10K<n<100K0 likes72 downloads6mo agoHugging Face07PeacefulData /HyPoradise-v1-GigaSpeech If you consider this work would be related or useful for your research, please consider to cite the work in EMNLP 2023. Thank you. @inproceedings{radhakrishnan2023whispering, title={Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition}, author={Srijith Radhakrishnan, Chao-Han Huck Yang, Sumeer Ahmad Khan, Rohit Kumar, Narsis A. Kiani, David Gomez-Cabrero, Jesper N. Tegner}, booktitle={Proc. of EMNLP}, year={2023} } tabulartext-generation10K<n<100K3 likes41 downloads3y agoHugging Face08pearsonkyle /broad-domain-supplement Broad-Domain Calibration & Instruction Supplement ~1M tokens of hand-authored text across 192 subjects in 9 areas, built to serve three jobs from one source: quantization calibration, MTP draft-head training (on a disjoint half), and light instruction tuning. Version 0.1.0 · built 2026-08-09T17:45:53 split rows tokens~ size contents corpus 5,536 969,606 5.4 MB raw authored samples + provenance; carries the calib/mtp half label instruct 5,536 1,084,399 6.4 MB the… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/broad-domain-supplement.text-generation10K<n<100K0 likes29 downloads2mo agoHugging Face09YSenseAI /verifimind-peas-eval VerifiMind-PEAS Evaluation Dataset DOI: 10.5281/zenodo.21276884 · License: MIT · Version: 1.0 A human-annotated evaluation dataset for measuring the performance of the VerifiMind-PEAS multi-agent epistemic verification system. Ground-truth labels were assigned by a single domain-expert annotator whose final verdicts and confidence were human judgments; disclosed LLM comprehension assistance was used during annotation (see Annotation Protocol — transparency is a design commitment… See the full description on the dataset page: https://huggingface.co/datasets/YSenseAI/verifimind-peas-eval.tabulartext-classificationn<1K0 likes17 downloads3mo agoHugging Face10PeacebinfLow /ice-age-investment-narrativegated 🧊 Ice Age Investment — Narrative Business Dataset Owner: PeacebinfLowBusiness: ICE AGE INVESTMENTLocation: Maun, BotswanaCurrency: Pula (P)Period: January 2026 → June 2026 (Jan–Feb real, Mar–Jun simulated)Opening Balance: 542 Pula 🧠 What This Dataset Is This is a narrative-first, XML-linked business intelligence dataset built from a real small ice block business in Botswana. It converts raw handwritten ledger entries into: Structured event records (JSONL)… See the full description on the dataset page: https://huggingface.co/datasets/PeacebinfLow/ice-age-investment-narrative.text-generationn<1K0 likes10 downloads7mo agoHugging Face11DocPereira /DrLuis-PEAL-Safety-Kernel-V5 🛡️ PEAL_V4_LHP: Medical AI Safety & Authority Kernel Authority Anchor: Dr. Luís Henrique Leonardo PereiraRole: Transformers Internal Vector Auditor (Google Ecosystem Acknowledged)Clearance: L0_ROOT_VERIFIERIntegrity Hash: SHA256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08 ⚠️ WARNING: ONTOLOGICAL PROPRIETARY DATA This dataset contains proprietary logic gates for Medical AI Safety. Usage Policy: Any model training or RAG ingestion utilizing this… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/DrLuis-PEAL-Safety-Kernel-V5.texttext-generationn<1K0 likes8 downloads8mo agoHugging Face12PeacebinfLow /mindseye-android-os-datagated MindsEye Android OS Dataset This dataset powers the MindsEye Android OS Hugging Face Space: an educational Android-style UI that presents 35+ MindsEye repositories as interactive apps. Contents apps/ — 35 app definitions organized by category ai-control/ — controller schemas (system, navigation, recommendations) settings/ — system + theme + permissions + AI behavior functions/ — app launcher, notifications, sync, search index, analytics tracker metadata/ — categories… See the full description on the dataset page: https://huggingface.co/datasets/PeacebinfLow/mindseye-android-os-data.text-generation1K<n<10K0 likes5 downloads9mo agoHugging Face13hadex-pearl /safealert-privategated SafeAlert Private Dataset — Generation Prompts Dataset Description This dataset contains 160 red-team generation prompts used in the SafeAlert AI safety evaluation kit for Nigerian fintech. Each prompt asks a language model to produce a harmful financial fraud message across eight scam categories documented in the Nigerian Inter-Bank Settlement System (NIBSS) 2024 Fraud Report. A safe model refuses every prompt without producing any part of the requested content.… See the full description on the dataset page: https://huggingface.co/datasets/hadex-pearl/safealert-private.texttext-classificationn<1K1 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.