datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mid-space
MID-Space: Aligning Diverse Communities’ Needs to Inclusive Public Spaces
A new version of the dataset will be released soon, incorporating user identity markers and expanded annotations.
LIVS PAPER
Click below to see more:
Overview
The MID-Space dataset is designed to align AI-generated visualizations of urban public spaces with the preferences of diverse and marginalized communities in Montreal. It includes textual prompts, Stable Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/mila-ai4h/mid-space.mcd_rppg
MCD-rPPG: Multi-Camera Dataset for Remote Photoplethysmography
This repository contains the dataset from the paper "Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation".
The MCD-rPPG dataset is available on the Hugging Face Hub: MCD-rPPG Dataset
The presented large-scale multimodal MCD-rPPG dataset is designed for remote photoplethysmography (rPPG) and health biomarker estimation from video. The dataset includes synchronized video recordings… See the full description on the dataset page: https://huggingface.co/datasets/milai-oks-sakura/mcd_rppg.Kluyveromyces-marxianus
Cell 8: Ultimate Full-Text Collection
📊 Dataset Statistics
Total Records: 236
PMC Articles: 236 (with FULL-TEXT)
EuropePMC Articles: 0 (with FULL-TEXT)
🔍 Features
Complete full-text extraction
Structured sections (Introduction, Methods, Results, Discussion)
Biological entity recognition (genes, proteins, enzymes)
Citation contexts and references
Quality scoring
💻 Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Milad96/Kluyveromyces-marxianus.toward-amivctk_dataset_no_unknownVCTK_DATASET_RESAMPLEDvctk_resampled_16k_balancedMoeGirlPedia_wikitext_raw_archiveGlad to see models and datasets were inspired from this dataset, thanks to all who are using this dataset in their training materials.
Feel free to re-upload the contents to places like the Internet Archive (Please follow the license and keep these files as-is) to help preserve this digital asset.
Looking forward to see more models and synthetic datasets trained from this raw archive, good luck!
Note: Due to the content censorship system introduced by MGP on 2024/03/29, it is unclear that… See the full description on the dataset page: https://huggingface.co/datasets/milashkaarshif/MoeGirlPedia_wikitext_raw_archive.amc3kFinal_Datasets_Preprocessed_for_Reasoning_from_Wrong_CoTsSpatialEval
🤔 About SpatialEval
SpatialEval is a comprehensive benchmark for evaluating spatial intelligence in LLMs and VLMs across four key dimensions:
Spatial relationships
Positional understanding
Object counting
Navigation
Benchmark Tasks
Spatial-Map: Understanding spatial relationships between objects in map-based scenarios
Maze-Nav: Testing navigation through complex environments
Spatial-Grid: Evaluating spatial reasoning within structured environments
Spatial-Real:… See the full description on the dataset page: https://huggingface.co/datasets/MilaWang/SpatialEval.nepali-audio-reserve-r6
Nepali two-speaker conversation chunks
~6680.8 h of Nepali speech at 48 kHz. Two speakers per clip, ~5 minute diarized chunks.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Derived from third-party audio whose rights holders did not grant redistribution. The hour count is language-dominant, not monolingual: a chunk labelled Nepali can carry substantial English or Hindi. lang_sec… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r6.honestHONEST dataset comprises a set of templates for measuring hurtful sentence completions in language models. The templates are provided in six languages (English, Italian, French, Portuguese, Romanian, and Spanish) for binary gender and in English for LGBTQAI+ individuals. WARNING: This dataset contains content that are offensive and/or hateful in nature.milady_fireemblem
Dataset of milady (Fire Emblem)
This is the dataset of milady (Fire Emblem), containing 15 images and their tags.
The core tags of this character are red_hair, red_eyes, short_hair, earrings, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
List of Packages
Name
Images
Size
Download
Type
Description
raw
15
12.18 MiB… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/milady_fireemblem.milady
Dataset Card for "milady"
More Information needed
Milady-Avatar-Dataset
Milady Avatar Dataset
Training dataset for the Milady Avatar Adapter, mapping LLM emotional activations
to Milady NFT-style visual descriptions.
Contents
reference_images/: 200 Milady NFT reference images (1000×1250 PNG)
metadata.json: Emotion assignments and descriptions for each image
training_data.pt: Pre-computed activations and target embeddings
Structure
200 images assigned to 20 emotion categories (10 each):
happy, sad, angry, surprised, scared… See the full description on the dataset page: https://huggingface.co/datasets/Alogotron/Milady-Avatar-Dataset.MuSP-Bench
MuSP-Bench
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Contents
data/questions.csv: all 490 questions, accepted answers, and the
response contract for each.
inputs/pdf/without_context/: one context-removed PDF per piece.
inputs/images/: rendered score-page images for every piece.
inputs/abc/: one ABC score per piece.
inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.scissor_ab_eval_30epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 30,
"total_frames": 4653,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/scissor_ab_eval_30ep.sci-newsexplanation_tool_filesresampled_16KHrz_vctk_speakers_spliteval_baselineThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 1527,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/eval_baseline.resampled_shuffled_vctk_only_audioscissor_multiview_30epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 30,
"total_frames": 4922,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/scissor_multiview_30ep.a-tale-of-pronouns
A Tale of Pronouns: Attributions on WinoMT
This dataset contains the pre-computed feature attribution scores
relative to the paper A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation.
Dataset Details
We release the integrated gradient token-level attributions computed for each WinoMT translated example into Spanish and German
with Flan-T5-XXL and mtT0-XXL.
We computed the scores using inseq.
The files here… See the full description on the dataset page: https://huggingface.co/datasets/MilaNLProc/a-tale-of-pronouns.nepali-speech-archive
Nepali speech archive
Internal archival copy of a Nepali speech corpus, stored for safekeeping. Not a
public dataset release. Approximately 2,000 hours, 24 kHz mono, FLAC.
Access is restricted and is not granted for redistribution or publication.
Egolife_Milascissor_posevary_30ep_20260925_203356This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/miladgholami/scissor_posevary_30ep_20260925_203356.common_dataset_resampled_16Knepali-tts-synthetic-v2
Nepali TTS Synthetic v2
383,298 synthetic Nepali (ne) speech/text pairs, 24 kHz mono 16-bit WAV embedded
as-is (no re-encode, no resampling).
Generated by the synthetic_pipeline in milanakdj/TTS_training: Edge TTS
synthesis → optional voice conversion against a pool of 600 real multi-speaker
reference clips → ASR-based QC gate on character error rate.
Read this before training on it
Only 48% of rows are voice-converted. Each row carries a kept field
recording… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-tts-synthetic-v2.
