CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zinderud /risale-sohbet-turkish-2audio1K<n<10K0 likes3.6k downloads1y agoHugging Face02risaleinur /risale-i-nur-sohbet Risale-i Nur Sohbet Prof. Dr. Şener Dilek’ten izin alındı. Türkçe Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri kümelerine karıştırılmaz. Kapsam 2095 sohbet, 954.66 saat 16 kHz mono FLAC ses Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.audioautomatic-speech-recognition1M<n<10M1 likes1.2k downloads20d agoHugging Face03risaleinur /risale-nur-audio Risale-i Nur Audio–Text Corpus Gerçek insan okumalarını, aynı satırdaki kaynak metinle birlikte sunan açık bir ses–metin veri kümesidir. Yeni varsayılan audio-text yapılandırması 15 kitaptan 91.792 oynatılabilir klip ve 203,02 saat ses içerir. Metinler kanonik kaynaktan değiştirilmeden alınır ve her kayıt byte-exact section_id alıntılarıyla bağlanır. An open speech corpus pairing human readings with their source text in the same row. The default audio-text config contains 91,792… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-audio.audioautomatic-speech-recognition100K<n<1M1 likes401 downloads13h agoHugging Face04risaleinur /risalei-nur-text-audio Risale-i Nur Text–Audio Kaynak · Source: RNK Neşriyat — yazılı izinle · used with written permission. Her satırda gerçek insan okuması ile o sesin kanonik metni birlikte bulunur. Sesler dış bağlantı değildir: WAV baytları Parquet dosyalarının içindedir. Kaynak sitesi veya başka bir ses sunucusu gerekmez. Each row pairs a human reading with its canonical transcript. Audio is stored as WAV bytes inside the Parquet files; no source website or external audio server is required.… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risalei-nur-text-audio.audioautomatic-speech-recognition10K<n<100K1 likes378 downloads13h agoHugging Face05GEM /RiSAWOZRiSAWOZ contains 11.2K human-to-human (H2H) multiturn semantically annotated dialogues, with more than 150K utterances spanning over 12 domains, which is larger than all previous annotated H2H conversational datasets.Both single- and multi-domain dialogues are constructed, accounting for 65% and 35%, respectively.text10K<n<100K8 likes332 downloads4y agoHugging Face06risaleinur /risale-nur-grounded-multipool Risale-i Nur Grounded Multi-Pool LLM Dataset TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim çalışmaları için çok görünümlü bir veri seti. EN. A multi-view dataset built from 15 canonical Risale-i Nur books for grounded generation, SFT, preference learning, evaluation, continued pretraining, and retrieval. v2.10.0 · 199 configs · 463 config/split views · 527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.tabulartext-generation100K<n<1M3 likes276 downloads18d agoHugging Face07risashinoda /BioVITA Citation @inproceedings{shinoda2026biovita, title = {BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment}, author = {Risa Shinoda and Kaede Shiohara and Nakamasa Inoue and Kuniaki Saito and Hiroaki Santo and Fumio Okura}, booktitle = {CVPR}, year = {2026}, } image1M<n<10M2 likes265 downloads4mo agoHugging Face08risaleinur /risale-nur-multilingual Risale-i Nur Multilingual Corpus Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir. Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.tabulartranslation100K<n<1M2 likes255 downloads1mo agoHugging Face09risashinoda /AgroBenchgated AgroBench: Vision-Language Model Benchmark in Agriculture Authors: Risa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka, Masaki Onishi, Yoshitaka Ushiku (ICCV'25) Citation If you use our dataset, please cite our paper. @InProceedings{Shinoda_2025_ICCV, author = {Shinoda, Risa and Inoue, Nakamasa and Kataoka, Hirokatsu and Onishi, Masaki and Ushiku, Yoshitaka}, title = {AgroBench: Vision-Language Model Benchmark in Agriculture}, booktitle = {Proceedings of the… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/AgroBench.image1K<n<10K19 likes169 downloads9mo agoHugging Face10zinderud /risale-sohbet-turkish YouTube Transkripsiyon Veri Seti Veri Yapısı audio/: MP3 dosyaları transcripts/: Metin transkripsiyonları srt/: Altyazı dosyaları metadata/: Video bilgileri database.json: Tüm videoların indeksi Güncelleme Tarihi 2025-03-21 audion<1K1 likes123 downloads1y agoHugging Face11risashinoda /footprint_yologated AnimalClue YOLO Datasets AnimalClue: Recognizing Animals by their Traces 📌 ICCV 2025 Highlight This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format. Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/footprint_yolo.image1K<n<10K2 likes86 downloads5mo agoHugging Face12zinderud /risalehttps://github.com/zinderud/HuginRisale textn<1K0 likes26 downloads2y agoHugging Face13risashinoda /HalCap-Benchgated HalCap-Bench HalCap-Bench dataset. Columns model image_source image_name image_type sentence_index caption annotation error_type error_words agreement_ratio fleiss_Pi n_correct n_incorrect n_unknown image_url image_path_in_repo Notes Notes For COCO/CC12M items, the image is referenced by image_url. For SD/Imagen/data_generation items, the image file is stored under images/ and referenced by image_path_in_repo. image100K<n<1M0 likes15 downloads7mo agoHugging Face14ambile-official /Shah_Jo_Risalo_labeld Shah Abdul Latif Bhittai’s Poetry Dataset – “Shah Jo Risalo” Developed by: Abdul Majid Bhurgri Institute of Language Engineering (AMBILE), HyderabadUnder the administrative control of the Culture, Tourism, Antiquities & Archives Department, Government of Sindh. Dataset Overview: The "Shah Jo Risalo" Dataset is a rich linguistic and literary resource comprising 43,779 Sindhi poetic verses extracted from the 30 traditional Surs of Shah Abdul Latif Bhittai’s… See the full description on the dataset page: https://huggingface.co/datasets/ambile-official/Shah_Jo_Risalo_labeld.tabularfeature-extraction10K<n<100K1 likes14 downloads1y agoHugging Face15risashinoda /bone_yologated AnimalClue YOLO Datasets AnimalClue: Recognizing Animals by their Traces 📌 ICCV 2025 Highlight This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format. Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/bone_yolo.image0 likes13 downloads5mo agoHugging Face16ambile-official /AMBILE_Shah_Jo_Risalo_Labeled AMBILE Shah Jo Risalo Developed by:Abdul Majid Bhurgri Institute of Language Engineering (AMBILE), HyderabadUnder the administrative control of the Culture, Tourism, Antiquities & Archives Department, Government of Sindh Dataset Overview The "Shah Jo Risalo" dataset serves as a comprehensive linguistic and literary resource, encompassing 4,767 Sindhi poetic verses drawn from the 30 traditional Surs (sections) of the esteemed magnum opus of Shah Abdul Latif Bhittai. Each… See the full description on the dataset page: https://huggingface.co/datasets/ambile-official/AMBILE_Shah_Jo_Risalo_Labeled.tabulartext-classification1K<n<10K0 likes13 downloads1y agoHugging Face17risashinoda /feces_yologated AnimalClue YOLO Datasets AnimalClue: Recognizing Animals by their Traces 📌 ICCV 2025 Highlight This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format. Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/feces_yolo.image1K<n<10K0 likes12 downloads5mo agoHugging Face18DeepPavlov /RISAWOZtext10K<n<100K0 likes11 downloads1y agoHugging Face19risashinoda /egg_yologated AnimalClue YOLO Datasets AnimalClue: Recognizing Animals by their Traces 📌 ICCV 2025 Highlight This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format. Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/egg_yolo.image0 likes10 downloads5mo agoHugging Face20risan-raja-iitm /urbansound8K(card and dataset copied from https://www.kaggle.com/datasets/chrisfilo/urbansound8k) This dataset contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. The classes are drawn from the urban sound taxonomy. For a detailed description of the dataset and how it was compiled please refer to our paper.All excerpts are taken from field recordings… See the full description on the dataset page: https://huggingface.co/datasets/risan-raja-iitm/urbansound8K.audioaudio-classification1K<n<10K0 likes10 downloads1mo agoHugging Face21risangpanggalih /betawi-v0Synthetic Betawi Language dataset, generated by GPT-4o: Betawi v0 (Alpha) Betawi v0 is a synthetic dataset created using GPT-4o, consisting of over 1,000 instruction-output pairs across a range of topics, structured in JSON format. It follows the Alpaca dataset format and is designed for fine-tuning large language models (LLMs) to enhance LLMs understanding of Bahasa Betawi. Version Alpha .hf-sanitized.hf-sanitized-uZg6qHEzHPlvxjDu8mVLI h1 { font-size: 36px; color: #000000;… See the full description on the dataset page: https://huggingface.co/datasets/risangpanggalih/betawi-v0.texttext-generation1K<n<10K1 likes3 downloads2y agoHugging Face22Risalat /bengali-speech-chunksgatedaudio100K<n<1M1 likes3 downloads8mo agoHugging Face23risalstr /faq_datasettextn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.