kap
Datasets
All datasets matching “kap”Dataset
MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records
1. Executive Summary & Repository Overview
The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.bolAIndia
bolAIndia
Human-side speech from production call recordings, cut into utterance-level
chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR
providers. Each row keeps the transcript, the provider's confidence, and full
provenance back to the source recording.
Sources
One config per transcription system, so their output stays separable.
config (source_id)
provider
model
hours
rows
shards
vendor-a
vendor-a
undisclosed
420.03
480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.Shanghai
Shanghai Eye Disease Center Ophthalmic Multimodal Dataset (上海市眼病防治中心多模态眼科数据集)
A comprehensive, multi-center, longitudinal ophthalmic foundation dataset from Shanghai Eye Disease Prevention & Treatment Center.
Repository Layout
Kaphathy/Shanghai/
└── Topcon/
└── shards/
├── manifest.json # O(1) Index mapping each exam_id to its shard
├── meta.tar.gz # Complete clinical JSON metadata for all 30,711 exams (3.3 MB)… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Shanghai.kaputt
Kaputt: A Large-Scale Dataset for Visual Defect Detection
Abstract
We present a novel large-scale dataset for defect detection in a logistics
setting. Recent work on industrial anomaly detection has primarily focused on
manufacturing scenarios with highly controlled poses and a limited number of
object categories. Existing benchmarks like MVTec-AD (Bergmann et al., 2021) and
VisA (Zou et al., 2022) have reached saturation, with state-of-the-art methods
achieving… See the full description on the dataset page: https://huggingface.co/datasets/amazon/kaputt.Chaashini
Chaashini (चाशनी)
Chaashini — Hindi/Urdu for sugar syrup — is a continuously growing corpus of clean, single-speaker,
studio-grade Indian-language speech built for training speech models (text-to-speech, speech
recognition, speech language models). Every clip in the corpus has passed a strict multi-stage
quality gate; the aim is purity over volume.
Total: 1,363,490 clips · 2900.08 hours · 33 languages
Format: mono 24 kHz FLAC (audio column) with a verbatim transcript and rich… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/Chaashini.
