datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
farsi-asr-unified-cleaned
🎧 Farsi ASR Unified Dataset (Parquet Sharded Edition)
Overview
The Farsi ASR Unified Dataset is a large-scale, high-quality, and fully standardized collection of Persian (Farsi) speech-to-text data — designed specifically for modern machine learning and ASR (Automatic Speech Recognition) workflows.
This dataset consolidates audio–text pairs from multiple open sources, applies a rigorous cleaning and normalization pipeline, and stores everything efficiently in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/kiarashQ/farsi-asr-unified-cleaned.kiarashQ-aug1NightFogPedestrianDataset_KiaRayPOV
Night Fog Pedestrian Dataset (Kia Ray POV)
이 데이터셋은 유니티 퍼셉션(Unity Perception)을 활용하여 생성된 **합성 데이터셋(Synthetic Dataset)**입니다. 자율주행 모델이 안개 낀 야간 환경에서 보행자를 얼마나 정확하게 탐지하는지 테스트하기 위해 제작되었습니다.
1. 데이터 개요
시점 (POV): 기아 레이(Kia Ray) 차량의 블랙박스 위치 (지면으로부터 약 1.4m 높이)
환경 조건: 야간 (Night), 안개 (Foggy/Low Visibility)
클래스: Pedestrian (단일 클래스)
데이터 포맷: YOLOv8 (images/labels)
해상도: 1101 x 514 pixels
2. 데이터셋 구조
.
├── dataset.yaml # YOLOv8 설정 파일
├── images/ # .jpg 이미지 파일… See the full description on the dataset page: https://huggingface.co/datasets/JTSGRIT/NightFogPedestrianDataset_KiaRayPOV.GhostWordcrmsc-envsGPTMicro-Nanowire-Sintering
GPTMicro — Nanowire Sintering & Symbolic Regression Dataset
Curated data for data-driven discovery of governing equations in nanowire
sintering. It pairs raw molecular-dynamics (MD) trajectories with the ML-ready
train/validation/test splits used to learn closed-form models for the sintering
dynamics (change in flattening ddelta and rotation dtheta) and for two
effective material properties (effective diffusion coefficient D_eff and
effective relaxation/viscosity coefficient… See the full description on the dataset page: https://huggingface.co/datasets/Kiarash99/GPTMicro-Nanowire-Sintering.architecture-collection
Architecture Multimodal3 Data Notes
Dataset summary
Preparation notes and schema examples for Architecture tasks using Multimodal3 data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/Kiarasin/architecture-collection.kiara1TinyPersianStories_normalizedsessyoin_kiara_fgo
Dataset of sessyoin_kiara/殺生院キアラ/杀生院祈荒 (Fate/Grand Order)
This is the dataset of sessyoin_kiara/殺生院キアラ/杀生院祈荒 (Fate/Grand Order), containing 500 images and their tags.
The core tags of this character are yellow_eyes, breasts, long_hair, black_hair, large_breasts, facial_mark, parted_bangs, very_long_hair, multicolored_hair, pink_hair, streaked_hair, wavy_hair, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/sessyoin_kiara_fgo.chibi-animekiaraTTS
Kiara TTS
Masih belum ada training, hanya data mentahan saja
berisi CSV dan file audio WAV
data hanya 270 Audio
Semoga ada yang bisa melatih data ini, beritau saya jika kamu
bisa melatih data ini.terimakasih banya atas kerja samanya
Developed by: [Niki]
Language(s) (NLP): [Indonesian]
License: [Apache]
Tallbergwikipedia_wikimedia_fa_normalizedkiaradataset_kiaraV5
