datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
risale-sohbet-turkish-2risale-i-nur-sohbet
Risale-i Nur Sohbet
Prof. Dr. Şener Dilek’ten izin alındı.
Türkçe
Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde
birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış
sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri
kümelerine karıştırılmaz.
Kapsam
2095 sohbet, 954.66 saat 16 kHz mono FLAC ses
Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste
seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.padma-meghna-riverbank-erosionrisale-nur-audio
Risale-i Nur Audio–Text Corpus
Gerçek insan okumalarını, aynı satırdaki kaynak metinle birlikte sunan açık bir
ses–metin veri kümesidir. Yeni varsayılan audio-text yapılandırması 15 kitaptan
91.792 oynatılabilir klip ve 203,02 saat ses içerir. Metinler kanonik kaynaktan
değiştirilmeden alınır ve her kayıt byte-exact section_id alıntılarıyla bağlanır.
An open speech corpus pairing human readings with their source text in the same
row. The default audio-text config contains 91,792… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-audio.risalei-nur-text-audio
Risale-i Nur Text–Audio
Kaynak · Source: RNK Neşriyat — yazılı izinle · used with written permission.
Her satırda gerçek insan okuması ile o sesin kanonik metni birlikte bulunur.
Sesler dış bağlantı değildir: WAV baytları Parquet dosyalarının içindedir.
Kaynak sitesi veya başka bir ses sunucusu gerekmez.
Each row pairs a human reading with its canonical transcript. Audio is stored
as WAV bytes inside the Parquet files; no source website or external audio
server is required.… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risalei-nur-text-audio.RiSAWOZRiSAWOZ contains 11.2K human-to-human (H2H) multiturn semantically annotated dialogues, with more than 150K utterances spanning over 12 domains, which is larger than all previous annotated H2H conversational datasets.Both single- and multi-domain dialogues are constructed, accounting for 65% and 35%, respectively.risale-nur-grounded-multipool
Risale-i Nur Grounded Multi-Pool LLM Dataset
TR. 15 kanonik Risale-i Nur kitabından hazırlanan; kaynak
bağlı üretim, SFT, tercih, değerlendirme, sürekli ön eğitim ve erişim
çalışmaları için çok görünümlü bir veri seti.
EN. A multi-view dataset built from 15 canonical Risale-i
Nur books for grounded generation, SFT, preference learning, evaluation,
continued pretraining, and retrieval.
v2.10.0 · 199 configs · 463 config/split views ·
527,196 rows across configured views… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-grounded-multipool.BioVITA
Citation
@inproceedings{shinoda2026biovita,
title = {BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment},
author = {Risa Shinoda and Kaede Shiohara and Nakamasa Inoue and Kuniaki Saito and Hiroaki Santo and Fumio Okura},
booktitle = {CVPR},
year = {2026},
}
risale-nur-multilingual
Risale-i Nur Multilingual Corpus
Bediüzzaman Said Nursî'nin Risale-i Nur külliyatının 27 dilde çok dilli korpusu — her eser başlıklara göre bölümlere (section) ayrılmış, bölümler diller arasında hizalanmış ve konu (topic) hiyerarşisiyle etiketlenmiştir.
Güncel release: v2.10.0 · 20 config/lane · 163,820 config-split satırı. Alt başlıklardaki eski v2.x etiketleri lane'in ilk eklendiği sürümü gösterir; güncel release sürümü değildir. Deterministik projeksiyonlar duplicate_of ile… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-multilingual.AgroBench
AgroBench: Vision-Language Model Benchmark in Agriculture
Authors: Risa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka, Masaki Onishi, Yoshitaka Ushiku (ICCV'25)
Citation
If you use our dataset, please cite our paper.
@InProceedings{Shinoda_2025_ICCV,
author = {Shinoda, Risa and Inoue, Nakamasa and Kataoka, Hirokatsu and Onishi, Masaki and Ushiku, Yoshitaka},
title = {AgroBench: Vision-Language Model Benchmark in Agriculture},
booktitle = {Proceedings of the… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/AgroBench.YunXiaoHe-RISA-Avalon-Eval
云小鹤 0.3.9.8:RISA 与 Avalon 公开评测
这份成果集记录云小鹤在两类公开任务中的完整成绩:跨数据集表格建模,以及有状态、长时程的供应链决策。这里可以看到版本、能力、评测口径、结果、图和证据哈希;云小鹤的私有工作系统、RISA 与 Avalon 的实现没有公开。
一眼看懂结果
评测
结果
说明
TabArena 30 数据集
macro ROC-AUC 0.816690
较既有同集结果 +3.60 个百分点,24/30 数据集胜出
五任务同模型对照
Token -62.7%
accuracy 保留 98.9%
SupplyChainBench
54.810907
16/16 局、576/576 动作全部完成
冻结公开榜位置
完整模型点估计第 1
若加入冻结快照;较 Muse Spark 1.2 分数高 6.7%
使用的云小鹤能力
云小鹤版本:0.3.9.8
主模型:DeepSeek Pro,non-fast
云小鹤 -… See the full description on the dataset page: https://huggingface.co/datasets/HanyueShen/YunXiaoHe-RISA-Avalon-Eval.RiSAWOZrisale-sohbet-turkish
YouTube Transkripsiyon Veri Seti
Veri Yapısı
audio/: MP3 dosyaları
transcripts/: Metin transkripsiyonları
srt/: Altyazı dosyaları
metadata/: Video bilgileri
database.json: Tüm videoların indeksi
Güncelleme Tarihi
2025-03-21
footprint_yolo
AnimalClue YOLO Datasets
AnimalClue: Recognizing Animals by their Traces
📌 ICCV 2025 Highlight
This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format.
Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/footprint_yolo.animalclap-dataset
AnimalCLAP
AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait InferenceICASSP 2026
AuthorsRisa Shinoda, Kaede Shiohara, Nakamasa Inoue, Hiroaki Santo, Fumio Okura
Overview
This dataset contains 701,020 animal sound recordings collected from:
iNaturalist
Xeno-Canto
Splits
HF Split
Original Split
Description
train
train
Training data (URL only)
validation
test
Validation data (URL only)
test
zero_shot… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/animalclap-dataset.risalehttps://github.com/zinderud/HuginRisale
HalCap-Bench
HalCap-Bench
HalCap-Bench dataset.
Columns
model
image_source
image_name
image_type
sentence_index
caption
annotation
error_type
error_words
agreement_ratio
fleiss_Pi
n_correct
n_incorrect
n_unknown
image_url
image_path_in_repo
Notes
Notes
For COCO/CC12M items, the image is referenced by image_url.
For SD/Imagen/data_generation items, the image file is stored under images/ and referenced by image_path_in_repo.
matoba_risa_theidolmastercinderellagirlsu149
Dataset of Matoba Risa
This is the dataset of Matoba Risa, containing 200 images and their tags.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface organization).
Name
Images
Download
Description
raw
200
Download
Raw data with meta information.
raw-stage3
431
Download
3-stage cropped raw data with meta information.
384x512
200
Download
384x512 aligned dataset.
512x512
200
Download… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/matoba_risa_theidolmastercinderellagirlsu149.matoba_risa_idolmastercinderellagirls
Dataset of matoba_risa/的場梨沙 (THE iDOLM@STER: Cinderella Girls)
This is the dataset of matoba_risa/的場梨沙 (THE iDOLM@STER: Cinderella Girls), containing 500 images and their tags.
The core tags of this character are long_hair, black_hair, twintails, yellow_eyes, bangs, ribbon, hair_between_eyes, hair_ribbon, breasts, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is powered by DeepGHS Team(huggingface… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/matoba_risa_idolmastercinderellagirls.Shah_Jo_Risalo_labeld
Shah Abdul Latif Bhittai’s Poetry Dataset – “Shah Jo Risalo”
Developed by:
Abdul Majid Bhurgri Institute of Language Engineering (AMBILE), HyderabadUnder the administrative control of the Culture, Tourism, Antiquities & Archives Department, Government of Sindh.
Dataset Overview:
The "Shah Jo Risalo" Dataset is a rich linguistic and literary resource comprising 43,779 Sindhi poetic verses extracted from the 30 traditional Surs of Shah Abdul Latif Bhittai’s… See the full description on the dataset page: https://huggingface.co/datasets/ambile-official/Shah_Jo_Risalo_labeld.bone_yolo
AnimalClue YOLO Datasets
AnimalClue: Recognizing Animals by their Traces
📌 ICCV 2025 Highlight
This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format.
Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/bone_yolo.AMBILE_Shah_Jo_Risalo_Labeled
AMBILE Shah Jo Risalo
Developed by:Abdul Majid Bhurgri Institute of Language Engineering (AMBILE), HyderabadUnder the administrative control of the Culture, Tourism, Antiquities & Archives Department, Government of Sindh
Dataset Overview
The "Shah Jo Risalo" dataset serves as a comprehensive linguistic and literary resource, encompassing 4,767 Sindhi poetic verses drawn from the 30 traditional Surs (sections) of the esteemed magnum opus of Shah Abdul Latif Bhittai. Each… See the full description on the dataset page: https://huggingface.co/datasets/ambile-official/AMBILE_Shah_Jo_Risalo_Labeled.feces_yolo
AnimalClue YOLO Datasets
AnimalClue: Recognizing Animals by their Traces
📌 ICCV 2025 Highlight
This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format.
Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/feces_yolo.RISAWOZfeather_yolo
AnimalClue YOLO Datasets
AnimalClue: Recognizing Animals by their Traces
📌 ICCV 2025 Highlight
This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format.
Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/feather_yolo.egg_yolo
AnimalClue YOLO Datasets
AnimalClue: Recognizing Animals by their Traces
📌 ICCV 2025 Highlight
This repository is part of the AnimalClue project, which explores the recognition of wild animals from indirect clues such as feathers, footprints, feces, eggs, and bones. These datasets are designed for object detection training using YOLO format.
Each image filename is linked to an observation ID, and any use of the image must comply with the license associated… See the full description on the dataset page: https://huggingface.co/datasets/risashinoda/egg_yolo.urbansound8K(card and dataset copied from https://www.kaggle.com/datasets/chrisfilo/urbansound8k)
This dataset contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: air_conditioner, car_horn, children_playing, dog_bark, drilling, enginge_idling, gun_shot, jackhammer, siren, and street_music. The classes are drawn from the urban sound taxonomy. For a detailed description of the dataset and how it was compiled please refer to our paper.All excerpts are taken from field recordings… See the full description on the dataset page: https://huggingface.co/datasets/risan-raja-iitm/urbansound8K.arabic-al-risala-shafii-dataset
📖 Arabic Al-Risala AI Dataset (الإمام الشافعي)
📌 Dataset Overview
This dataset contains a cleaned and structured version of "Al-Risala" by Imam Al-Shafi'i (Edited by Ahmad Muhammad Shakir), formatted specifically for training Large Language Models (LLMs), Arabic NLP, and Islamic QA Systems.
Author: Imam Al-Shafi'i (150–204 AH)
Investigator: Ahmad Muhammad Shakir
Domain: Islamic Jurisprudence & Classical Arabic NLP
Format: .parquet & .jsonl with Rich Metadata… See the full description on the dataset page: https://huggingface.co/datasets/sabateen83/arabic-al-risala-shafii-dataset.urban_sounds_8kbetawi-v0Synthetic Betawi Language dataset, generated by GPT-4o: Betawi v0 (Alpha)
Betawi v0 is a synthetic dataset created using GPT-4o, consisting of over 1,000 instruction-output pairs across a range of topics, structured in JSON format. It follows the Alpaca dataset format and is designed for fine-tuning large language models (LLMs) to enhance LLMs understanding of Bahasa Betawi.
Version Alpha
.hf-sanitized.hf-sanitized-uZg6qHEzHPlvxjDu8mVLI h1 { font-size: 36px; color: #000000;… See the full description on the dataset page: https://huggingface.co/datasets/risangpanggalih/betawi-v0.
