CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LanguageBind /Open-Sora-Plan-v1.1.0 Annotation We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973 Pexels Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.text100K<n<1M46 likes104k downloads2y agoHugging Face02hasankursun /github-code-2025-language-split 📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models. Processing Steps To create this dataset, we performed the following processing on the source data: Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.text100M<n<1B13 likes26k downloads10mo agoHugging Face03CohereLabs /aya_collection_language_split This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages. Dataset Summary The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.tabular100M<n<1B122 likes22k downloads1y agoHugging Face04akoksal /muri-it-language-split MURI-IT: Multilingual Instruction Tuning Dataset for 200 Languages via Multilingual Reverse Instructions MURI-IT is a large-scale multilingual instruction tuning dataset containing 2.2 million instruction-output pairs across 200 languages. It is designed to address the challenges of instruction tuning in low-resource languages with Multilingual Reverse Instructions (MURI), which ensures that the output is human-written, high-quality, and authentic to the cultural and linguistic… See the full description on the dataset page: https://huggingface.co/datasets/akoksal/muri-it-language-split.texttext-generation1M<n<10M6 likes11k downloads2y agoHugging Face05DesmondYMTang2024 /Language-Grounded_Sparse_Encoder_Training Language-Grounded Sparse Encoder (LanSE) — Training Data This repository hosts the AI-generated images and human annotation datasets accompanying the paper: Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.textimage-classification100K<n<1M1 likes7k downloads15d agoHugging Face06livebench /language Dataset Card for "livebench/language" LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties: LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/language.textn<1K0 likes6.1k downloads1y agoHugging Face07papluca /language-identification Dataset Card for Language Identification dataset Dataset Summary The Language Identification dataset is a collection of 90k samples consisting of text passages and corresponding language label. This dataset was created by collecting data from 3 sources: Multilingual Amazon Reviews Corpus, XNLI, and STSb Multi MT. Supported Tasks and Leaderboards The dataset can be used to train a model for language identification, which is a multi-class text classification… See the full description on the dataset page: https://huggingface.co/datasets/papluca/language-identification.texttext-classification10K<n<100K70 likes3.3k downloads4y agoHugging Face08LanguageBind /UniWorld-V1 The Geneval-style dataset is sourced from BLIP3o-60k. This dataset is presented in the paper: UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation More details can be found in UniWorld-V1 Data preparation Download the data from LanguageBind/UniWorld-V1. The dataset consists of two parts: source images and annotation JSON files. Prepare a data.txt file in the following format: The first column is the root path to the image. The second… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/UniWorld-V1.image1K<n<10K25 likes2.7k downloads1y agoHugging Face09candradhipa /Language-Detectiontext10K<n<100K0 likes1.3k downloads2y agoHugging Face10sakthivinash /Language_Detection Language_Detection - Multilingual Text Classification Dataset This dataset is a collection of multilingual text samples designed for training and predicting languages in Artificial Intelligence (AI), Machine Learning (ML), Deep Learning (DL), and Data Science (DS) applications. It contains labeled data that associates text samples with their respective languages, enabling language detection and classification tasks. Dataset Overview The dataset consists of two columns:… See the full description on the dataset page: https://huggingface.co/datasets/sakthivinash/Language_Detection.text10K<n<100K0 likes1.2k downloads2y agoHugging Face11Aletheia-ng /low_resource_languages_pretrain_data5text100M<n<1B0 likes1.2k downloads11mo agoHugging Face12LanguageBind /Open-Sora-Plan-v1.0.0 Open-Sora-Dataset Welcome to the Open-Sora-DataSet project! As part of the Open-Sora-Plan project, we specifically talk about the collection and processing of data sets. To build a high-quality video dataset for the open-source world, we started this project. 💪 We warmly welcome you to join us! Let's contribute to the open-source world together! Thank you for your support and contribution. If you like our project, please give us a star ⭐ on GitHub for latest update.… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.0.0.text1K<n<10K66 likes1.1k downloads2y agoHugging Face13Aletheia-ng /low_resource_languages_pretrain_data2text100M<n<1B0 likes1.1k downloads1y agoHugging Face14Benji-fish /ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes1k downloads1mo agoHugging Face15oxe-auge /language_table_inpainting language_table robot-removal inpainting dataset This dataset contains robot-removal inpainting results for language_table. Each episode provides: inpainting.mp4: the robot visually removed via inpainting mask.mp4: the robot mask video used for inpainting original_episode.mp4: the original (unmodified) episode video language_instructions_{split}_all.txt: tab-separated mapping from episode_id to instruction Relation to OXE-AugE This release is produced as part of… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/language_table_inpainting.textn<1K0 likes934 downloads9mo agoHugging Face16BeardedMonster /low_resource_languages_pretrain_data8text100M<n<1B0 likes891 downloads9mo agoHugging Face17jpbello /common_language_preprocessed Dataset Card for "common_language_preprocessed" More Information needed text10K<n<100K0 likes890 downloads3y agoHugging Face18Mike0307 /language-detection Dataset Card for "language-detection" More Information needed text10K<n<100K2 likes708 downloads3y agoHugging Face19Aletheia-ng /low_resource_languages_pretraintext100M<n<1B1 likes699 downloads1y agoHugging Face20AIStudioDelta /Eurovoc_2025_by_language 🇪🇺 🏷️ EuroVoc dataset (by language) This is the EuropeanParliament/Eurovoc_2025 dataset, but split up by language, not by period. The original is split up into periods (1996-03 through 2025-11), with documents in different languages mixed together. For ease of training this dataset splits the data by language instead, with documents in different periods put together. License This dataset is redistributed under the original European Union Public License 1.2. When… See the full description on the dataset page: https://huggingface.co/datasets/AIStudioDelta/Eurovoc_2025_by_language.texttext-generation1M<n<10M1 likes693 downloads10mo agoHugging Face21BrunoHays /mixed_multilingual_commonvoice_all_languages_100kBuild from mozilla commonvoice 13 using the script commited in this repo. Used to teach a model to ignore languages that are not french audio100K<n<1M0 likes672 downloads2y agoHugging Face22simoneteglia /europarl_for_language_detection_10ktext100K<n<1M0 likes649 downloads3y agoHugging Face23zomi-language-corpora /raw-text-corpus 📝 Zomi Raw Text Corpus (Community-Contributed) The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks. This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately. 📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.texttext-generationn<1K0 likes633 downloads4mo agoHugging Face24language-and-voice-lab /samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.audioautomatic-speech-recognition10K<n<100K9 likes569 downloads3y agoHugging Face25LanguageShades /BiasShadesgatedInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab! Dataset Card for BiasShades Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators. Dataset Details Version: 1.0 License: SHADES 1 Montreal Data License Dataset Description 728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.imagetext-classificationn<1K26 likes554 downloads3mo agoHugging Face26Aletheia-ng /low_resource_languages_pretrain_data4text100M<n<1B0 likes498 downloads11mo agoHugging Face27Sakonii /nepalitext-language-model-dataset Dataset Card for "nepalitext-language-model-dataset" Dataset Summary "NepaliText" language modeling dataset is a collection of over 13 million Nepali text sequences (phrases/sentences/paragraphs) extracted by combining the datasets: OSCAR , cc100 and a set of scraped Nepali articles on Wikipedia. Supported Tasks and Leaderboards This dataset is intended to pre-train language models and word representations on Nepali Language. Languages The data is… See the full description on the dataset page: https://huggingface.co/datasets/Sakonii/nepalitext-language-model-dataset.texttext-generation10M<n<100M8 likes483 downloads1y agoHugging Face28sirgecko /language_detection_traintext100K<n<1M0 likes468 downloads2y agoHugging Face29rmems /neuromorphic-event-language-bridge Neuromorphic Event-Language Bridge Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated event-language payload is now published under data/raw/. It is available for… See the full description on the dataset page: https://huggingface.co/datasets/rmems/neuromorphic-event-language-bridge.textn<1K0 likes466 downloads14d agoHugging Face30jbross-ibm-research /Marco-Bench-MIF-languagestext10K<n<100K0 likes461 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.