datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dataset
MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records
1. Executive Summary & Repository Overview
The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.bolAIndia
bolAIndia
Human-side speech from production call recordings, cut into utterance-level
chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR
providers. Each row keeps the transcript, the provider's confidence, and full
provenance back to the source recording.
Sources
One config per transcription system, so their output stays separable.
config (source_id)
provider
model
hours
rows
shards
vendor-a
vendor-a
undisclosed
420.03
480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.Shanghai
Shanghai Eye Disease Center Ophthalmic Multimodal Dataset (上海市眼病防治中心多模态眼科数据集)
A comprehensive, multi-center, longitudinal ophthalmic foundation dataset from Shanghai Eye Disease Prevention & Treatment Center.
Repository Layout
Kaphathy/Shanghai/
└── Topcon/
└── shards/
├── manifest.json # O(1) Index mapping each exam_id to its shard
├── meta.tar.gz # Complete clinical JSON metadata for all 30,711 exams (3.3 MB)… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Shanghai.kaputt
Kaputt: A Large-Scale Dataset for Visual Defect Detection
Abstract
We present a novel large-scale dataset for defect detection in a logistics
setting. Recent work on industrial anomaly detection has primarily focused on
manufacturing scenarios with highly controlled poses and a limited number of
object categories. Existing benchmarks like MVTec-AD (Bergmann et al., 2021) and
VisA (Zou et al., 2022) have reached saturation, with state-of-the-art methods
achieving… See the full description on the dataset page: https://huggingface.co/datasets/amazon/kaputt.Chaashini
Chaashini (चाशनी)
Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker,
studio-grade Indian-language speech built for training speech models (text-to-speech, speech
recognition, speech language models). Every clip in the corpus has passed a strict multi-stage
quality gate; the aim is purity over volume.
Total: 1,368,934 clips · 2909.22 hours · 33 languages
Format: mono 24 kHz FLAC (audio column) with a verbatim transcript and rich… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/Chaashini.tts-dataset-combinedSEC_filings_1994_2024
Dataset Card for SEC EDGAR Filings Master Index
Dataset Details
Dataset Description
This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates.
Curated by: Arthur (arthur@cicero.chat)
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.Dartboard-Detection-Dataset
Dartboard Detection Dataset
A curated dartboard image dataset for computer vision tasks such as detection, recognition, localization, and model training.
This dataset is used in my dartboard AI projects built with Rust and PyTorch. Anyone can use this dataset to train, test, or improve their own models for dartboard-related computer vision tasks.
About
This dataset contains cropped dartboard images organized in folders by capture sessions and dates. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/bhabha-kapil/Dartboard-Detection-Dataset.Vartalaap
Vartalaap — full-duplex Hindi/English conversational speech
Dual-channel synthetic Indian customer-support calls for training full-duplex
speech-to-speech models.
68,674 calls · 1,564.4 hours · 1786 shards
(last updated 2026-09-24 10:34 IST)
Audio layout
Each row's audio is a stereo FLAC at 24000 Hz:
channel
content
0 (LEFT)
agent — pristine, TTS speech and silence only
1 (RIGHT)
user — the caller
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/Vartalaap.KapInstruct-100M
KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset
KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.ohun
ohùn — Igbo · Yorùbá · Hausa · Pidgin speech corpus
ohùn (Yorùbá for voice) merges the three WaZoBiaSpeech corpora published by
Africanvoice into a single repository, so all
three of Nigeria's major languages can be pulled from one place.
Audio is byte-identical to the sources — this repo re-registers the very same
objects, it does not re-encode anything.
Contents
718,336 utterances · 1,035 GB of audio across four languages.
config
split
rows
shards
size… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/ohun.seville_maria_data
SevilleWorkflow
Kapibara
Kapibara: Albanian Multi-turn Conversation Dataset
Dataset Summary
Kapibara is a comprehensive Albanian language dataset designed for multi-turn conversations. It contains over 5,300 entries covering a wide range of topics including physics, biology, mathematics, chemistry, culture, and logic. The dataset is aimed at improving text generation and question-answering capabilities in the Albanian language.
Supported Tasks
The dataset supports the following NLP… See the full description on the dataset page: https://huggingface.co/datasets/alban-labs/Kapibara.kapla_tower_3_expertThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 15,
"total_frames": 17925,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kantine/kapla_tower_3_expert.helical_dna_theory_kappa_adaptive_ds_2000nmkapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.helical_dna_high_rand_new_kappa_adaptive_ds_2000nmmagazine-kapkan
Magazine «Kapkăn»
Description
Illustrated literary magazine of satire and humor, published in Cheboksary (Chuvash Republic) from 1925 to 2017. Digitized span in source: 1925–1940.
Completeness note
Long run in print; only part is digitized here. Within 1925–1940, individual months/issues may still be missing on disk.
Data layout
Recommended: one folder per year 1925/ … 1940/, PDF per issue.
Source
National Library of the Chuvash Republic:… See the full description on the dataset page: https://huggingface.co/datasets/chuvash-data/magazine-kapkan.helical_dna_theory_kappa_fixed_dskap-turkish-financial-sentiment
KAP Turkish Financial Sentiment Dataset
Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti.
Dataset Bilgileri
Özellik
Değer
Kayıt Sayısı
3,839
Dil
Türkçe
Kaynak
KAP Bildirimleri
Etiketleme
GPT-4 (Teacher Model)
Format
JSONL (Chat Messages)
Kullanım Alanları
Türkçe finansal sentiment analizi
KAP bildirimi sınıflandırma
Volatilite tahmini
İlişkili taraf işlemi tespiti
LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/finansai/kap-turkish-financial-sentiment.helical_dna_theory_kappaoopsie_kaplaTowerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/kantine/oopsie_kaplaTower.call-transcript-intent-data-v2
Call Transcript Intent Dataset
Multimodal Hindi/Hinglish customer utterance dataset for loan/EMI/payment call intent classification.
Dataset Summary
Metric
Value
Total examples
139,348
Total audio duration
51.04 h
Number of intents
17
Split Statistics
Split
Examples
Duration
Hours
train
126,848
2755.14 min
45.92 h
validation
10,000
219.03 min
3.65 h
eval
2,500
88.37 min
1.47 h
Class Distribution… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/call-transcript-intent-data-v2.kapla_tower_2_expertThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 20,
"total_frames": 23917,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kantine/kapla_tower_2_expert.helical_dna_higher_rand_new_kappa_adaptive_ds_2000nmkap-turkish-financial-sentiment
KAP Turkish Financial Sentiment Dataset
Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti.
Dataset Bilgileri
Özellik
Değer
Kayıt Sayısı
3,839
Dil
Türkçe
Kaynak
KAP Bildirimleri
Etiketleme
GPT-4 (Teacher Model)
Format
JSONL (Chat Messages)
Kullanım Alanları
Türkçe finansal sentiment analizi
KAP bildirimi sınıflandırma
Volatilite tahmini
İlişkili taraf işlemi tespiti
LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/furkanyllmz/kap-turkish-financial-sentiment.Xenopus_Modelsprocessed_bert_dataset_QAkapla_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 2392,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kantine/kapla_test.
