datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Benetech_PlotQa_DVQA_combined_matcha_completeChartQA_Benetech_PlotQa_DVQA_combined_matcha_completematcha_stir
matcha_stir
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
details_matchaaaaa__chaifighter-20Bnewsqa_200_11064_v2.0.0
Private NewsQA RAG Evaluation Dataset
Private, human-reviewed evaluation source data for the NewsQA RAG project.
This repository is not a prebuilt retrieval index. Chunk, BM25, Chroma, and
ground-truth chunk mappings must be rebuilt from the pinned release.
Version: v1.0.0
Evaluation articles: 200
Distractor articles: 10864
Source questions: 1340
Redistribution rights for the upstream NewsQA-derived text must be verified
before changing this repository from private to public.
TextMining-Phase-1-Index
phase1-indexes
Retrieval indexes built by the NewsQA RAG Phase 1 pipeline, packaged so a later
session can attach them instead of spending GPU time rebuilding.
Generated 2026-09-06T11:52:10+00:00 from /kaggle/input/notebooks/meowluvmatcha/newsqa-rag-phase-1-retrieval-tournament-kaggl/newsqa_phase1/indexes/round1.
Contents
round1/ — dense: BAAI/bge-large-en-v1.5, dense: BAAI/bge-small-en-v1.5, dense: all-MiniLM-L6-v2, dense: intfloat/e5-base-v2, sparse:… See the full description on the dataset page: https://huggingface.co/datasets/MatchaMacchiato/TextMining-Phase-1-Index.Benetech_PlotQa_DVQA_combined_matchaMatCha
Dataset Description
Materials characterization plays a key role in understanding the processing–microstructure–property relationships that guide material design and optimization. While multimodal large language models (MLLMs) have shown promise in generative and predictive tasks, their ability to interpret real-world characterization imaging data remains underexplored.
MatCha is the first benchmark designed specifically for materials characterization image understanding. It… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/MatCha.improved-matcha-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 40,
"total_frames": 27404,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/improved-matcha-dataset.matcha-making-final-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 175,
"total_frames": 102869,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:175"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/matcha-making-final-dataset.matcha-makingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 82,
"total_frames": 54058,
"total_tasks": 4,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:82"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/matcha-making.s14llamatest-matcha-stirringThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 4,
"total_frames": 5123,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/test-matcha-stirring.official-matcha-making-repoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 10,
"total_frames": 16194,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/official-matcha-making-repo.matcha-hail-maryThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 5,
"total_frames": 3476,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/matcha-hail-mary.MATCHA
MATCHA: A Multi-Disease Chinese Mental Health Dataset with Cross-Lingual Analyses
1. Overview / 项目简介
MATCHA (Multi-Disorder mental health daTa in CHinA) is a large-scale, community-driven Chinese Weibo dataset designed for multi-disorder mental health research. It is the first Chinese multi-disease dataset that links community posts from Weibo Super-Topics (超话) with user-level historical timelines.
MATCHA(Multi-Disorder mental health daTa in… See the full description on the dataset page: https://huggingface.co/datasets/emiliememe/MATCHA.matcha-dataset-xxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 127,
"total_frames": 93810,
"total_tasks": 5,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:127"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/matcha-dataset-xx.matcha-hail-mary-againThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 101,
"total_frames": 53140,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/matcha-hail-mary-again.2026-thesis-datamatch_attck_datasethomemade-matcha-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 30,
"total_frames": 53218,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/SIGRoboticsUIUC/homemade-matcha-dataset.match_annotations_hf_dataset_220kmatcha-making-final-dataset-v2.1matcharm_whisk_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 67,
"total_frames": 19926,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:67"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dhanasrimanthini/matcharm_whisk_v1.S1_ABmatcha-dataset-xx-v2.1matcharm_whisk_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 20,
"total_frames": 6036,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dhanasrimanthini/matcharm_whisk_v2.img2gps3kaa2dna-dataset
Amino Acid → DNA Translation & Codon optimization Dataset
Description
This dataset contains randomly generated Amino acids sequences and their reverse translated DNA seqs with codon optimization.
It is suitable for Seq2Seq learning and translation model training.
Dataset Structure
Fields
src: input DNA sequence (str)
tgt: translated amino acid sequence (str)
Splits
train, validation, test
eval_matcharm_smolVLA_whisk_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dhanasrimanthini/eval_matcharm_smolVLA_whisk_v1.
