CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LanguageBind /Open-Sora-Plan-v1.0.0 Open-Sora-Dataset Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪 We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution. If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.text1K<n<10K66 likes924 downloads2y agoHugging Face02Mizukiluke /ureader-instruction-1.0image10K<n<100K15 likes282 downloads3y agoHugging Face03laion /majestrino-1.00-16xk5-sae-features Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5) Top-2000 activating audio samples for each feature in the Majestrino 1.00 SAE. Overview Metric Value SAE Architecture 16x expansion, k=5, d_model=768 Total Features 12,288 Alive Features 10,684 Audio per Feature Up to 2,000 highest-activating Audio Format Opus (24 kbps OGG container) Total TAR Files 1069 Source Dataset laion/majestrino-data File Structure Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.audioaudio-classification10M<n<100M0 likes272 downloads6mo agoHugging Face04Hemgg /Open-Sora-Plan-v1.0.0Open Sora plan collected 40,258 high-quality, watermark-free videos from open-source websites under the CC0 license. About 60% of the videos are in landscape format, with a total duration of approximately 274 hours, 5 minutes, and 13 seconds. The dataset is divided into three main sources: Mixkit: Videos: 1,234 Total duration: 6h 19m 32s Total frames: 570,815 Resolution and aspect ratio distributions (less than 1% not listed). Pexels: Videos: 7,408 Total duration: 48h 49m 24s Total… See the full description on the dataset page: https://huggingface.co/datasets/Hemgg/Open-Sora-Plan-v1.0.0.text100K<n<1M0 likes266 downloads11mo agoHugging Face05Aalto-Speech-Synthesis /stortinget_speech_corpus_v1.0 Dataset Card for Stortinget Speech Corpus V1.0 Overview This is the WebDataset version of the Stortinget Speech Corpus V1.0, originally created by the National Library of Norway. We re-organize it into WebDataset format for better usability. The Stortinget Speech Corpus (SSC) is a 5000+ hours speech dataset for weak supervision ASR created from audio andaligned proceedings text from Stortinget, the Norwegian Parliament. For more information, please refer to the original… See the full description on the dataset page: https://huggingface.co/datasets/Aalto-Speech-Synthesis/stortinget_speech_corpus_v1.0.audioautomatic-speech-recognition100K<n<1M0 likes105 downloads5mo agoHugging Face06Caesarrr /objaverse-1.0-renderingsimage1M<n<10M0 likes25 downloads6mo agoHugging Face07kelvinzhaozg /droid_1.0.1-ara-cachegatedtext1M<n<10M0 likes1 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.