datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset.
The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database.
The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/kapilrao/SEC-EDGAR.KapInstruct-100M
KapInstruct-100M: Curated 100-Million Token Instruction Tuning Dataset
KapInstruct-100M is a high-fidelity, 100-million-token instruction-tuning dataset engineered for Supervised Fine-Tuning (SFT) and alignment of compact language models (under 1 billion parameters). Formatted with the Qwen ChatML chat template and tokenized using Qwen/Qwen3.5-0.8B-Base, the dataset enforces strict assistant-only loss masking (masking user prompts and structural delimiters to -100)… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapInstruct-100M.Kapibara
Kapibara: Albanian Multi-turn Conversation Dataset
Dataset Summary
Kapibara is a comprehensive Albanian language dataset designed for multi-turn conversations. It contains over 5,300 entries covering a wide range of topics including physics, biology, mathematics, chemistry, culture, and logic. The dataset is aimed at improving text generation and question-answering capabilities in the Albanian language.
Supported Tasks
The dataset supports the following NLP… See the full description on the dataset page: https://huggingface.co/datasets/alban-labs/Kapibara.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.kap-turkish-financial-sentiment
KAP Turkish Financial Sentiment Dataset
Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti.
Dataset Bilgileri
Özellik
Değer
Kayıt Sayısı
3,839
Dil
Türkçe
Kaynak
KAP Bildirimleri
Etiketleme
GPT-4 (Teacher Model)
Format
JSONL (Chat Messages)
Kullanım Alanları
Türkçe finansal sentiment analizi
KAP bildirimi sınıflandırma
Volatilite tahmini
İlişkili taraf işlemi tespiti
LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/finansai/kap-turkish-financial-sentiment.kap-turkish-financial-sentiment
KAP Turkish Financial Sentiment Dataset
Türkçe KAP (Kamuyu Aydınlatma Platformu) bildirimleri için çok boyutlu finansal analiz dataseti.
Dataset Bilgileri
Özellik
Değer
Kayıt Sayısı
3,839
Dil
Türkçe
Kaynak
KAP Bildirimleri
Etiketleme
GPT-4 (Teacher Model)
Format
JSONL (Chat Messages)
Kullanım Alanları
Türkçe finansal sentiment analizi
KAP bildirimi sınıflandırma
Volatilite tahmini
İlişkili taraf işlemi tespiti
LLM fine-tuning (Qwen… See the full description on the dataset page: https://huggingface.co/datasets/furkanyllmz/kap-turkish-financial-sentiment.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/kapibala-ai/kapibala-sales-dialogues.2STARGATE
🧠 STARGATE: CIA Remote Viewing Archive (Raw PDFs + Metadata)
STARGATE is the most comprehensive open-access archive of declassified CIA documents related to psychic research, remote viewing (RV), and anomalous cognition. This dataset consolidates over 12,000 scanned PDF files, drawn from decades of classified government programs designed to investigate and operationalize extrasensory perception (ESP) in intelligence-gathering contexts.
These include, but are not limited to:… See the full description on the dataset page: https://huggingface.co/datasets/kapustin2000/2STARGATE.KapCode-1B
KapCode-1B: Curated 1-Billion Token Dataset for Compact Code Models
KapCode-1B is a high-quality, 1-billion-token curated dataset designed for Continued Pre-Training (CPT) and domain adaptation of compact Large Language Models. Engineered specifically to empower models under 1 billion parameters with robust code generation, technical comprehension, mathematical reasoning, and Fill-in-the-Middle (FIM) infilling capabilities, KapCode-1B combines multi-lingual code… See the full description on the dataset page: https://huggingface.co/datasets/kaptaan45/KapCode-1B.alpaca_kapampangan
🇵🇭 Kapampangan Alpaca Dataset
Dataset Summary
This dataset is a Kapampangan translation of the original Alpaca instruction-following dataset. It is designed to support research and development of instruction-tuned language models for low-resource Philippine languages, particularly Kapampangan.
The dataset retains the original Alpaca structure while providing high-quality translations of instructions, inputs, and outputs.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/PLTAT/alpaca_kapampangan.
