datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Dataset
MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records
1. Executive Summary & Repository Overview
The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.bolAIndia
bolAIndia
Human-side speech from production call recordings, cut into utterance-level
chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR
providers. Each row keeps the transcript, the provider's confidence, and full
provenance back to the source recording.
Sources
One config per transcription system, so their output stays separable.
config (source_id)
provider
model
hours
rows
shards
vendor-a
vendor-a
undisclosed
420.03
480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.SEC_filings_1994_2024
Dataset Card for SEC EDGAR Filings Master Index
Dataset Details
Dataset Description
This dataset contains metadata for all submissions to the Securities and Exchange Commission (SEC) through their EDGAR system from 1994 until December 14, 2024. The data is extracted from quarterly master files and includes key information about company filings such as CIK numbers, company names, form types, and filing dates.
Curated by: Arthur (arthur@cicero.chat)
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC_filings_1994_2024.ohun
ohùn — Igbo · Yorùbá · Hausa · Pidgin speech corpus
ohùn (Yorùbá for voice) merges the three WaZoBiaSpeech corpora published by
Africanvoice into a single repository, so all
three of Nigeria's major languages can be pulled from one place.
Audio is byte-identical to the sources — this repo re-registers the very same
objects, it does not re-encode anything.
Contents
718,336 utterances · 1,035 GB of audio across four languages.
config
split
rows
shards
size… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/ohun.helical_dna_theory_kappa_adaptive_ds_2000nmkapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.helical_dna_high_rand_new_kappa_adaptive_ds_2000nmhelical_dna_theory_kappa_fixed_dskap-turkish-financial-sentiment
KAP Turkish Financial Sentiment Dataset
Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti.
Dataset Bilgileri
Özellik
Değer
Kayıt Sayısı
3,839
Dil
Türkçe
Kaynak
KAP Bildirimleri
Etiketleme
GPT-4 (Teacher Model)
Format
JSONL (Chat Messages)
Kullanım Alanları
Türkçe finansal sentiment analizi
KAP bildirimi sınıflandırma
Volatilite tahmini
İlişkili taraf işlemi tespiti
LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/finansai/kap-turkish-financial-sentiment.helical_dna_theory_kappacall-transcript-intent-data-v2
Call Transcript Intent Dataset
Multimodal Hindi/Hinglish customer utterance dataset for loan/EMI/payment call intent classification.
Dataset Summary
Metric
Value
Total examples
139,348
Total audio duration
51.04 h
Number of intents
17
Split Statistics
Split
Examples
Duration
Hours
train
126,848
2755.14 min
45.92 h
validation
10,000
219.03 min
3.65 h
eval
2,500
88.37 min
1.47 h
Class Distribution… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/call-transcript-intent-data-v2.helical_dna_higher_rand_new_kappa_adaptive_ds_2000nmkap-turkish-financial-sentiment
KAP Turkish Financial Sentiment Dataset
Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti.
Dataset Bilgileri
Özellik
Değer
Kayıt Sayısı
3,839
Dil
Türkçe
Kaynak
KAP Bildirimleri
Etiketleme
GPT-4 (Teacher Model)
Format
JSONL (Chat Messages)
Kullanım Alanları
Türkçe finansal sentiment analizi
KAP bildirimi sınıflandırma
Volatilite tahmini
İlişkili taraf işlemi tespiti
LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/furkanyllmz/kap-turkish-financial-sentiment.kapampangan-dictionary-embeddings
Kapampangan Dictionary Embeddings
The first dedicated Kapampangan sentence embedding dataset. 4,971 entries from a 1730s Kapampangan-English dictionary, enriched with LLM-generated semantic metadata and pre-computed embeddings from 6 models.
Designed for semantic search, retrieval, and clustering over Kapampangan vocabulary. Includes a 130-query retrieval benchmark and evaluation results from 14 retrieval improvement experiments.
Read the origin story: From a 300-Year-Old Dictionary… See the full description on the dataset page: https://huggingface.co/datasets/keithmanaloto/kapampangan-dictionary-embeddings.kapx-index
KAPX Index (K 指数) · Fear-Price — CNN Fear & Greed ÷ VIX
Canonical definition: The KAPX Index is a daily U.S. equity fear-pricing gauge published by Fear-Price (chronicle.klay-wang.com), computed as the CNN Fear & Greed reading divided by the VIX. The K stands for kǒng (恐), the Chinese character for fear; readings, methodology, and the complete signal ledger are permanently free and verifiable via Git timestamps.
官方定义:KAPX 指数(K… See the full description on the dataset page: https://huggingface.co/datasets/klay24/kapx-index.trading-bot-backtesting
Trading bot Backtesting CSVs
Gunbot ⇄ Trading Bot backtests (pair-level candle & matched order data)Created with Gunbot on Binance spot.
Dataset structure
column
type
description
ts
int64
candle Unix ms timestamp
open/high/low/close
float
OHLC price values
volume
float
traded volume in base currency
order_type
str
buy, sell, or empty (no order)
order_rate
float
executed rate
order_amount
float
amount traded
order_id
int64
exchange order id
pnl… See the full description on the dataset page: https://huggingface.co/datasets/kapr/trading-bot-backtesting.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/kapibala-ai/kapibala-sales-dialogues.processed_bert_dataset_free_speechqwen-kapi-01-datasetnew_datasetafrica-synth-cancer-kaposi-sarcoma-east-africa-comoros
Kaposi Sarcoma - East Africa | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-kaposi-sarcoma-east-africa-comoros.kaptan-pi-datasetwan21_kapi_01-datasetenglish_kapsikihelical_dna_high_rand_new_kappa_adaptive_ds_12_6_5000samples_newTiPAI-POC-Faithfulness
TiPAI-POC: Patch-Level Faithfulness Dataset
Dataset Description
This dataset contains text-to-image generation samples with patch-level faithfulness scores and systematic failure variations.
Each base prompt has 4 variations:
v0_original: Correct prompt (chosen baseline)
v1_attribute: Wrong color/size/material
v2_object: Swapped/wrong main object
v3_spatial: Wrong spatial relation or count
Purpose: Training patch-level preference models for text-to-image alignment… See the full description on the dataset page: https://huggingface.co/datasets/kapilw25/TiPAI-POC-Faithfulness.finetuning_demoIndicTTS-Hindi-Femalevaani-snac-cleanedhelical_dna_high_rand_new_kappa_adaptive_ds_12_6_5000samples_separate
