CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Kaphathy /Dataset MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records 1. Executive Summary & Repository Overview The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.textimage-classificationn<1K2 likes17k downloads1d agoHugging Face02kapturecx /bolAIndiagated bolAIndia Human-side speech from production call recordings, cut into utterance-level chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR providers. Each row keeps the transcript, the provider's confidence, and full provenance back to the source recording. Sources One config per transcription system, so their output stays separable. config (source_id) provider model hours rows shards vendor-a vendor-a undisclosed 420.03 480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.audioautomatic-speech-recognition10M<n<100M1 likes12k downloads7m agoHugging Face03kapilrao /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.text-generation1M<n<10M1 likes4.3k downloads5mo agoHugging Face04Kaphathy /Shanghai Shanghai Eye Disease Center Ophthalmic Multimodal Dataset (上海市眼病防治中心多模态眼科数据集) A comprehensive, multi-center, longitudinal ophthalmic foundation dataset from Shanghai Eye Disease Prevention & Treatment Center. Repository Layout Kaphathy/Shanghai/ └── Topcon/ └── shards/ ├── manifest.json # O(1) Index mapping each exam_id to its shard ├── meta.tar.gz # Complete clinical JSON metadata for all 30,711 exams (3.3 MB)… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Shanghai.0 likes3.4k downloads9d agoHugging Face05amazon /kaputtgated Kaputt: A Large-Scale Dataset for Visual Defect Detection Abstract We present a novel large-scale dataset for defect detection in a logistics setting. Recent work on industrial anomaly detection has primarily focused on manufacturing scenarios with highly controlled poses and a limited number of object categories. Existing benchmarks like MVTec-AD (Bergmann et al., 2021) and VisA (Zou et al., 2022) have reached saturation, with state-of-the-art methods achieving… See the full description on the dataset page: https://huggingface.co/datasets/amazon/kaputt.image-classification100K<n<1M4 likes2.5k downloads2mo agoHugging Face06kapturecx /Chaashinigated Chaashini (चाशनी) Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker, studio-grade Indian-language speech built for training speech models (text-to-speech, speech recognition, speech language models). Every clip in the corpus has passed a strict multi-stage quality gate; the aim is purity over volume. Total: 1,368,934 clips · 2909.22 hours · 33 languages Format: mono 24 kHz FLAC (audio column) with a verbatim transcript and rich… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/Chaashini.text-to-speech1M<n<10M1 likes2.3k downloads4m agoHugging Face07kapturecx /tts-dataset-combinedgated1 likes2k downloads2mo agoHugging Face08kapilrao /SEC_filings_1994_2024 Dataset Card for SEC EDGAR Filings Master Index Dataset Details Dataset Description This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates. Curated by: Arthur (arthur@cicero.chat) Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.tabular10M<n<100M0 likes818 downloads6mo agoHugging Face09bhabha-kapil /Dartboard-Detection-Dataset Dartboard Detection Dataset A curated dartboard image dataset for computer vision tasks such as detection, recognition, localization, and model training. This dataset is used in my dartboard AI projects built with Rust and PyTorch. Anyone can use this dataset to train, test, or improve their own models for dartboard-related computer vision tasks. About This dataset contains cropped dartboard images organized in folders by capture sessions and dates. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/bhabha-kapil/Dartboard-Detection-Dataset.image10K<n<100K1 likes760 downloads6mo agoHugging Face10kapturecx /Vartalaapgated Vartalaap — full-duplex Hindi/English conversational speech Dual-channel synthetic Indian customer-support calls for training full-duplex speech-to-speech models. 68,674 calls · 1,564.4 hours · 1786 shards (last updated 2026-09-24 10:34 IST) Audio layout Each row's audio is a stereo FLAC at 24000 Hz: channel content 0 (LEFT) agent — pristine, TTS speech and silence only 1 (RIGHT) user — the caller from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/Vartalaap.audioautomatic-speech-recognition10K<n<100K0 likes643 downloads1d agoHugging Face11kaptaan45 /KapInstruct-100M KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.question-answering10K<n<100K0 likes630 downloads1mo agoHugging Face12kapturecx /ohungated ohùn — Igbo · Yorùbá · Hausa · Pidgin speech corpus ohùn (Yorùbá for voice) merges the three WaZoBiaSpeech corpora published by Africanvoice into a single repository, so all three of Nigeria's major languages can be pulled from one place. Audio is byte-identical to the sources — this repo re-registers the very same objects, it does not re-encode anything. Contents 718,336 utterances · 1,035 GB of audio across four languages. config split rows shards size… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/ohun.audioautomatic-speech-recognition10K<n<100K0 likes492 downloads22d agoHugging Face13KapoorLabs-UCL /seville_maria_data SevilleWorkflow 0 likes355 downloads1y agoHugging Face14alban-labs /Kapibara Kapibara: Albanian Multi-turn Conversation Dataset Dataset Summary Kapibara is a comprehensive Albanian language dataset designed for multi-turn conversations. It contains over 5,300 entries covering a wide range of topics including physics, biology, mathematics, chemistry, culture, and logic. The dataset is aimed at improving text generation and question-answering capabilities in the Albanian language. Supported Tasks The dataset supports the following NLP… See the full description on the dataset page: https://huggingface.co/datasets/alban-labs/Kapibara.text-generation10K<n<100K5 likes279 downloads2y agoHugging Face15kantine /kapla_tower_3_expertThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 15, "total_frames": 17925, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:15" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kantine/kapla_tower_3_expert.tabularrobotics10K<n<100K0 likes255 downloads1y agoHugging Face16zwang2 /helical_dna_theory_kappa_adaptive_ds_2000nmtext1M<n<10M0 likes199 downloads2y agoHugging Face17Thomasgudan /kapibala-sales-dialogues Kapibala Sales Dialogues A sales-conversation dataset with outcome, conversation-level and sentence-level labels 🤗 Hugging Face · Annotation details · 中文 630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset: L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.tabulartext-classification10K<n<100K2 likes197 downloads7d agoHugging Face18zwang2 /helical_dna_high_rand_new_kappa_adaptive_ds_2000nmtext1M<n<10M0 likes186 downloads2y agoHugging Face19chuvash-data /magazine-kapkan Magazine «Kapkăn» Description Illustrated literary magazine of satire and humor, published in Cheboksary (Chuvash Republic) from 1925 to 2017. Digitized span in source: 1925–1940. Completeness note Long run in print; only part is digitized here. Within 1925–1940, individual months/issues may still be missing on disk. Data layout Recommended: one folder per year 1925/ … 1940/, PDF per issue. Source National Library of the Chuvash Republic:… See the full description on the dataset page: https://huggingface.co/datasets/chuvash-data/magazine-kapkan.document0 likes186 downloads5mo agoHugging Face20zwang2 /helical_dna_theory_kappa_fixed_dstext1M<n<10M0 likes181 downloads2y agoHugging Face21finansai /kap-turkish-financial-sentiment KAP Turkish Financial Sentiment Dataset Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti. Dataset Bilgileri Özellik Değer Kayıt Sayısı 3,839 Dil Türkçe Kaynak KAP Bildirimleri Etiketleme GPT-4 (Teacher Model) Format JSONL (Chat Messages) Kullanım Alanları Türkçe finansal sentiment analizi KAP bildirimi sınıflandırma Volatilite tahmini İlişkili taraf işlemi tespiti LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/finansai/kap-turkish-financial-sentiment.texttext-classification1K<n<10K1 likes154 downloads10mo agoHugging Face22zwang2 /helical_dna_theory_kappatext1M<n<10M0 likes130 downloads2y agoHugging Face23kantine /oopsie_kaplaTowerThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/kantine/oopsie_kaplaTower.tabularrobotics100K<n<1M0 likes130 downloads1mo agoHugging Face24kapturecx /call-transcript-intent-data-v2 Call Transcript Intent Dataset Multimodal Hindi/Hinglish customer utterance dataset for loan/EMI/payment call intent classification. Dataset Summary Metric Value Total examples 139,348 Total audio duration 51.04 h Number of intents 17 Split Statistics Split Examples Duration Hours train 126,848 2755.14 min 45.92 h validation 10,000 219.03 min 3.65 h eval 2,500 88.37 min 1.47 h Class Distribution… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/call-transcript-intent-data-v2.audio100K<n<1M0 likes121 downloads1mo agoHugging Face25kantine /kapla_tower_2_expertThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 20, "total_frames": 23917, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kantine/kapla_tower_2_expert.tabularrobotics10K<n<100K0 likes99 downloads1y agoHugging Face26zwang2 /helical_dna_higher_rand_new_kappa_adaptive_ds_2000nmtext1M<n<10M0 likes96 downloads2y agoHugging Face27furkanyllmz /kap-turkish-financial-sentiment KAP Turkish Financial Sentiment Dataset Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti. Dataset Bilgileri Özellik Değer Kayıt Sayısı 3,839 Dil Türkçe Kaynak KAP Bildirimleri Etiketleme GPT-4 (Teacher Model) Format JSONL (Chat Messages) Kullanım Alanları Türkçe finansal sentiment analizi KAP bildirimi sınıflandırma Volatilite tahmini İlişkili taraf işlemi tespiti LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/furkanyllmz/kap-turkish-financial-sentiment.texttext-classification1K<n<10K0 likes88 downloads10mo agoHugging Face28KapoorLabs-Copenhagen /Xenopus_Models0 likes81 downloads1y agoHugging Face29kapilchauhan /processed_bert_dataset_QA100K<n<1M0 likes76 downloads4y agoHugging Face30kantine /kapla_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 2, "total_frames": 2392, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kantine/kapla_test.tabularrobotics1K<n<10K0 likes76 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.