datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.RACER-Mini
RACER-Mini
RACER (Rationale-Aware Captioning of Edge-Case Driving Scenarios) is a reasoning caption dataset designed for training vision-language-action (VLA) models in autonomous driving.
This repository provides approximately 1,000 samples, as a small subset of the RACER dataset. Each sample consists of a temporal sequence of front camera images, the ego vehicle’s future trajectory, and a corresponding reasoning caption.
For details, please refer to our techblog RACER:… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/RACER-Mini.ZImage-Turbo-200k-multires-aspectbucketed
ZImage-Turbo WebDataset
Generated images from the ZImage-Turbo model using DiffusionDB prompts.
Generation Details
Hardware: 8x NVIDIA RTX 3090 GPUs
Generation Time: ~2 days
Estimated Cost: ~$70 (cloud compute)
Dataset Statistics
Total Samples: 211,081
Total Shards: 216
Samples per Shard: ~1000
Shard Naming Convention
Tarballs are named: {base_resolution}-{aspect_ratio}-{shard_num:04d}-of-{total_shards:04d}.tar
For example:… See the full description on the dataset page: https://huggingface.co/datasets/RareConcepts/ZImage-Turbo-200k-multires-aspectbucketed.urdu-turn-detection-audio-v2
🗣️ Urdu Turn Detection (Audio Dataset V2)
This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech.
It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection.
🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/0x3/Easy-Turn-Trainset.TurbSR3D
TurbSR3D
Processed data release used by the accompanying codebase.
Source
Derived from public resources available through JHTDB / SciServer.
JHTDB: https://turbulence.idies.jhu.edu/
SciServer: https://www.sciserver.org/
Related Repositories
Sample dataset: https://huggingface.co/datasets/heyhenry03/TurbSR3D-Sample
Checkpoints: https://huggingface.co/heyhenry03/TurbSR3D-checkpoints
MelTrim
MelTrim 数据集
欢迎来到 MelTrim 数据集!该存储库托管了用于训练和评估 MelTrim 模型的数据集,其中包含了多个知名的情感和多模态对话数据集。
简介
MelTrim 是一个专注于情感识别和多模态情感分析的模型。为了让模型能够更好地理解和处理复杂的人类情感,我们整合了以下三个高质量的数据集:
MEAD (Multi-view Emotional Audio-visual Dataset)
MELD (Multimodal EmotionLines Dataset)
M3ED (Multi-modal Multi-scene Multi-label Emotional Dialogue Database)
这些数据集提供了丰富的音频、视频和文本信息,涵盖了多种情感、场景和语言,为训练强大的多模态情感模型奠定了坚实的基础。
数据集详情
MEAD (Multi-view Emotional Audio-visual Dataset)
MEAD… See the full description on the dataset page: https://huggingface.co/datasets/turturtur250/MelTrim.librispeech_asr_for_speaker_turnurdu-turn-detection-audio-v2
🗣️ Urdu Turn Detection (Audio Dataset V2)
This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech.
It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection.
🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/urdu-turn-detection-audio-v2.Multi-turn-editingsdxl-turbo-fp-baselineturboquant-art-5kturboquant-art-1kturtle-imageturboquant-art-50k
