datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Drag-to-Live-Dataset
Drag-to-Live Dataset ☁️
This dataset contains processed time-lapse videos of clouds and their motion trajectories extracted using CoTracker.
It is designed for training Latent Trajectory Guidance models (like Wan-Move adaptation) on Edge Devices.
Dataset Structure
videos/: Resized (256x256) and clipped video segments (8-16 frames).
tracks/: .npz files containing dense point trajectories extracted by CoTracker.
index.jsonl: Metadata linking videos to their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/namin72/Drag-to-Live-Dataset.solana-clawd-live-datadataset-antarctique-live-25012026text-gpt-live-dataset
Text GPT-Live Training Dataset
The three supervised fine-tuning stages behind
Text GPT-Live, a text
interaction model that predicts one action from a chronological stream of
typed-text and tool events.
Write-up: Can you train a GPT-Live for $50?
· Code: github.com/huyxdang/Text-GPT-Live
Contents
Train the stages in this order, each resuming from the previous checkpoint at
a lower learning rate:
Order
File
Rows
Role
1
stage1_train.jsonl
6,013
Core… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/text-gpt-live-dataset.live_audit_test_datalive_stream_dataset_huggingface"""
_HOMEPAGE = "https://github.com/freeziyou/live_stream_dataset"
_LICENSE = "Creative Commons Attribution 4.0 International"
_TRAIN_DOWNLOAD_URL = "https://raw.githubusercontent.com/freeziyou/live_stream_dataset/main/train.csv"
# _TRAIN_DOWNLOAD_URL = "https://gitee.com/didi233/test_date_gitee/raw/master/train.csv"
_TEST_DOWNLOAD_URL = "https://raw.githubusercontent.com/freeziyou/live_stream_dataset/main/test.csv"
# _TEST_DOWNLOAD_URL = "https://gitee.com/didi233/test_date_gitee/raw/master/test.csv"
class live_stream_dataset_huggingface(datasets.GeneratorBasedBuilder):entheogen-live-datasetlivedatacourseChinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales
ID
King-TTS-271
Duration
4.24 hours
Language
Chinese
URL
https://dataoceanai.com/datasets/tts/chinese-female-speech-synthesis-corpus-live-streaming-for-sales/
Chinese_Male_Speech_Synthesis_Corpus_Live_Streaming_for_Sales
ID
King-TTS-272
Duration
4.32 hours
Language
Chinese
URL
https://dataoceanai.com/datasets/tts/chinese-male-speech-synthesis-corpus-live-streaming-for-sales/
Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles
ID
King-TTS-241
Duration
8.56 hours
Speakers
100 People
Labeling Details
Pronunciation, Rhythm, Breath sounds marked with {hx}
Language
Chinese
Description
Two styles: Deep and uplifting; covers a variety of product categories including food, clothing, beauty, personal care, electronics, and home goods.
URL… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales_with_Multi_Styles.Live_dataset
Live_dataset
Live_dataset is a dataset created using Russian-language YouTube videos. It currently contains text from 290 videos. This dataset is currently being expanded and will be used for training and further training the Vexion-LM models.
Important!! Since this dataset is based on videos, they may contain small advertisements and imperfect text. Since these are spoken by real people, it's impossible to eliminate "dirty" language.
Currently, the dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/DZER-Studios/Live_dataset.milvus-es-live-dataLIVE-World-Datasethumanoid-live-datalivedatacourselivedatamy_custom_live_datalive_reportlivedatacourselivedatacourselive-cell-segmentation-dataset
Light-microscopy cell segmentation dataset
Transmitted-light fields -- phase contrast, brightfield, DIC and quantitative phase -- with
uint16 instance masks, from 14 public datasets, curated for spaCR and used to fine-tune
live-cell-segmentation-cpsam.
11,007 fields: 6,778 train, 2,030 valid,
2,199 test. Split by ACQUISITION -- a well, a dish, a time-lapse or a
z-stack is never divided between sets -- so the test set is fields the model never saw
anything of.
Layout… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/live-cell-segmentation-dataset.
