datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.Bitcoin_synthetic_data
🧠 Bitcoin Synthetic Dataset Collection (AI-generated)
A collection of synthetic Bitcoin transaction datasets enriched with generative AI explanations.
🛠️ Topics
Whale Transactions
OP_RETURN rare patterns
Each transaction includes:
Fee, size, rarity score
Semantic AI-generated description
💡 Use Cases
Training predictive models of Bitcoin activity
Network and anomaly simulation
Financial behavior studies
Temporal analysis and outlier detection
📜License: CC… See the full description on the dataset page: https://huggingface.co/datasets/syn-data/Bitcoin_synthetic_data.synthetic-data-papers
Synthetic Data Papers — FineSet
A research-paper dataset on Synthetic Data Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Synthetic Data Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/synthetic-data-papers.synthetic-dataset-1208
Synthetic Key-Value Retrieval 32K
This is a deterministic synthetic benchmark for exact key-value retrieval from
a long context. It is designed for evaluating long-context inference and KV
cache compression methods.
Context format
The context contains a one-time task description followed by an array:
You are given an array of key-value entries. Every key begins with K_ and every value begins with V_. Each entry has the format [key: value]. Given a query key, find… See the full description on the dataset page: https://huggingface.co/datasets/ollamaweights/synthetic-dataset-1208.synthetic_sae_datasetsynthetic-dataset-1208-64k
Synthetic Key-Value Retrieval 64K
This is a deterministic synthetic benchmark for exact key-value retrieval from
a long context. It is designed for evaluating long-context inference and KV
cache compression methods.
Context format
The context contains a one-time task description followed by an array:
You are given an array of key-value entries. Every key begins with K_ and every value begins with V_. Each entry has the format [key: value]. Given a query key, find… See the full description on the dataset page: https://huggingface.co/datasets/ollamaweights/synthetic-dataset-1208-64k.ccs_synthetic_ar_1M-Arabic_dataset_1M_translated_jsonl_formatEcom-Chatbot-Synthetic-Test-Datasetaitf-dfk3-synthetic-audio-datasetSynthetic_Dataset_For_MCQA
