datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
global-piqa-nonparallel
Global PIQA Non-Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems.
In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements.
Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.massive_cdswikipedia-20240520
Dataset Card for "wikipedia-20240520"
More Information needed
bbeh-evalSIR_robocasaGithub repo
Project page
paper-visamichael_aftonwiki-visafineweb-visacds_eukTest_ru_datasetMr-LHDR
Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
Citation
@misc{guo2026mrlhdrbenchmarkmultimodalrealworld,
title={Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents},
author={Minghao Guo and Meng Cao and Sui Zhao and Siyu Ning and Xin Wang and Haoze Zhao and Jiaxuan Yang and Haihong Hao and Mingfei Han and Shunlin Rong and Haijun Wu and Xiaodan Liang and Xiaojun Chang},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/Henryeahhh/Mr-LHDR.RealStories-Micro-MRL
Dataset Card for ReactiveAI/RealStories-Micro-MRL
First synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models.
Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have
different number of follow-up interactions, could use different strategy, and have train and validation
splits.
Subsets
steps-1: ~2300 train (~4600 interactions) / ~340 validation (~680 interactions) - Single-Step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/RealStories-Micro-MRL.mrl-sample
Overview
Mean ribosome load (MRL) is a measure of translational efficient. This experimental dataset is a massively parallel translational assay, which assesses the translational impact of randomized 5'UTR sequences within an eGFP or mCherry reporter construct. For the eGFP experiments, two alternative RNA biologies are also evaluated. The varying and designed subsets use a variable length 5'UTR sequence and algorithmically designed 5'UTRs, respectively. Two duplicate were performed… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/mrl-sample.general-sharding-output-fineweb-1014Wild-Chat-MRLmrl-sugimoto
Overview
Mean Ribosome Load (MRL) measures the mean number of ribosomes attached to a given transcript. This dataset is collected from Sugimoto and Ratcliffe 2022, where MRL measurements are made in human kidney cancer cells. Importantly, this dataset is isoform-resolved, providing measurements at an isoform level.
This dataset is redistributed as part of mRNABench: https://github.com/morrislab/mRNABench
Data Format
Description of data columns:
target: Mean ribosome… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/mrl-sugimoto.Viridiplantaemrl-hl-lbkwk
Data Source
This data was processed from supplementary data in "Combinatorial optimization of mRNA structure, stability, and translation for RNA-based therapeutics", which was published in Nature Communications under a Creative Commons Attribution 4.0 International License. If you use this dataset, please provide the proper attribution:
Original paper: https://www.nature.com/articles/s41467-022-28776-wCitation: Leppek, K., Byeon, G.W., Kladwang, W. et al. Combinatorial optimization… See the full description on the dataset page: https://huggingface.co/datasets/morrislab/mrl-hl-lbkwk.TinyStories-MRL
Dataset Card for ReactiveAI/TinyStories-MRL
Synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models.
Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have
different number of follow-up interactions, could use different strategy, and have train and validation
splits.
After first experiments with MRL, we decided to abandon single step and two steps stages. That's because with single
step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/TinyStories-MRL.usc-vecs-v1-chunks-v1-s8192-o512-sentence-transformers-static-retrieval-mrl-en-v1usc-chroma-vecs-v1-chunks-v1-s8192-o512-sentence-transformers-static-retrieval-mrl-en-v1
usc-chroma-vecs-v1-chunks-v1-s8192-o512-sentence-transformers-static-retrieval-mrl-en-v1
This dataset contains a pre-built ChromaDB database with US Congressional legislation embeddings.
Dataset Structure
This dataset contains the ChromaDB files for legislation chunks with embeddings. The database can be loaded directly using ChromaDB's PersistentClient.
Usage
import chromadb
from huggingface_hub import snapshot_download
# Download the dataset
local_dir =… See the full description on the dataset page: https://huggingface.co/datasets/hyperdemocracy/usc-chroma-vecs-v1-chunks-v1-s8192-o512-sentence-transformers-static-retrieval-mrl-en-v1.usc-vecs-v1-chunks-v1-s4096-o512-sentence-transformers-static-retrieval-mrl-en-v1voxtral-emotion-speech
Voxtral Emotion Speech Dataset
Emotional speech dataset generated with ElevenLabs v3 using audio tags for emotion control, validated with SenseVoice for quality assurance.
Quality Filter
Each generated clip is validated using SenseVoice (iic/SenseVoiceSmall) to ensure the emotion in the audio matches the expected label.
Process:
Generate audio with ElevenLabs v3 using emotion audio tags
Run SenseVoice inference on the audio
Compare detected emotion with expected emotion… See the full description on the dataset page: https://huggingface.co/datasets/MrlolDev/voxtral-emotion-speech.msmarco-doc-slimmrl_predictiond2search-datathai-traffic-law-qawebinstruct-verified-fixmc
