datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Forex_Factory_Calendar
📅 Forex Factory Economic Calendar Dataset (2007-01-01 to 2025-04-07)
This dataset contains a comprehensive archive of macroeconomic calendar events sourced from Forex Factory, spanning from January 1, 2007 to April 7, 2025.Each row captures a specific event with detailed metadata including currency, event type, market impact level, reported values, and descriptive context.
📦 Dataset Summary
Total timespan: 2007-01-01 → 2025-04-07
Format: CSV (UTF-8)… See the full description on the dataset page: https://huggingface.co/datasets/karenholzkopf/Forex_Factory_Calendar.pet-health-symptoms-dataset
Pet Health Symptoms Dataset
Overview
This dataset contains 2,000 LLM-generated pet health symptoms text samples covering 5 common pet health condition categories, designed to train ML models for automated pet health classification. Each entry is labeled with:
Pet health condition (1 of 5 distinct classes)
Record type (Owner Observation or Clinical Notes)
Owner observations are expressed in everyday language (e.g., "My cat scratches constantly"), whereas clinical… See the full description on the dataset page: https://huggingface.co/datasets/karenwky/pet-health-symptoms-dataset.spiritual-development-4b59d7
spiritual-development-4b59d7
Synthetic products test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/Karen-Williams/spiritual-development-4b59d7.sagaw_karen_asrThis is the first public Sagaw Karen language ASR dataset in AI history.
Sagaw Karen ASR
This dataset contains audio recordings and aligned metadata in the Sagaw Karen language (ISO 639-3: ksw), a major Sgaw Karenic language spoken throughout southern and eastern Myanmar. The language is sometimes also referred to as Sgaw Karen or Sakaw Karen in English transliterations.
All audio segments in this dataset were sourced from publicly available news broadcasts published by PVTV… See the full description on the dataset page: https://huggingface.co/datasets/freococo/sagaw_karen_asr.karenni_language_asr_audio
RFA Karenni (Kayah) Language Voices
This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.common-variety-d04dad
common-variety-d04dad
Synthetic products test data: 45 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/KarenSmith/common-variety-d04dad.grand-mouth-c49a18
grand-mouth-c49a18
Synthetic sensors test data: 45 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Azure-Karen/grand-mouth-c49a18.KarenSo
Dataset Description
This dataset includes 50 sentences spoken in colloquial Hong Kong Cantonese (HKC), covering interrogatives and statements. Sentences are sourced from online Cantonese teaching materials and classical commercial slogans. It includes approxiametly 198 seconds of audio recorded by a female native speaker of HKC. The sampling rate is 44.1 kHz with 16-bit resolution. Transcription in Jyutping were also provided.
Issues Encountered & Solution
I… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/KarenSo.western_poe_karen_asrThis is the first public Western Poe Karen language ASR dataset in AI history.
Western Poe Karen ASR
This dataset contains audio recordings and aligned transcriptions in the Western Poe Karen language (also known in linguistic literature as Western Pwo or Delta Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in the Ayeyarwady Delta region of Myanmar. Although linguists commonly refer to this language as Western Pwo Karen, the community and this project prefer the spelling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/western_poe_karen_asr.strong-blue-dcb3f9
strong-blue-dcb3f9
Synthetic weather test data: 42 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/karen-thompson/strong-blue-dcb3f9.bluemoon_Karen_cleanedThis is a version of bluemoon-fandom-1-1-rp-cleaned further cleaned up using Karen_TheEditor 13B, in Fastchat format.
I tried to fix as many of the grammatical issues as possible and didn't drop any conversations, but there are still issues since Karen is not perfect (and only 13B).
If I detected any large deviations from the original text in the corrections, I fell back to a standard spell-checker, excluding estimated proper nouns from the spell-checker (which is also not perfect).
I intended… See the full description on the dataset page: https://huggingface.co/datasets/grimulkan/bluemoon_Karen_cleaned.eastern_poe_karen_asrThis is the first public Eastern Poe Karen language ASR dataset in AI history.
Eastern Poe Karen ASR
This dataset contains audio recordings and aligned metadata in the Eastern Poe Karen language (a regional variety of Eastern Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in Mon State and Kayin State in southeastern Myanmar. While linguistically described as Eastern Pwo Karen, the community and this project prefer the term Poe as a community-endorsed spelling.
All audio… See the full description on the dataset page: https://huggingface.co/datasets/freococo/eastern_poe_karen_asr.dialect_model_demoladino-karen-TTS
Ladino Text-to-Speech (TTS) Training Dataset
Dataset Description
This dataset contains a single-speaker speech corpus in Ladino (Judeo-Spanish) recorded by a native speaker from Istanbul. The corpus was created for training text-to-speech synthesis models for this endangered language.
Dataset Statistics
Speaker: Karen (native Ladino speaker)
Recordings: 1987 segments
Total Duration: ~3.3 hours
Sampling Rate: 16 kHz
Audio Format: WAV (16-bit, mono)
Language:… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/ladino-karen-TTS.grimulkan_bluemoon_Karen_cleaned-carded-formattedJust a simple text replace of the tags.
First Character: The Beast
Second Character: Belle
First Character Description: A mysterious and intimidating figure, resembling a beast with a cape swishing behind him. He has an imposing presence, which he uses to assert dominance over others in his castle. His personality is stern and authoritative; he is not afraid to enforce rules or punish those who disobey him. Despite this harsh exterior, The Beast also displays signs of vulnerability and… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers/grimulkan_bluemoon_Karen_cleaned-carded-formatted.ADL_2023_HW1KarenSo_CantoneseRecordings
Dataset Description
This dataset includes 50 sentences spoken in colloquial Hong Kong Cantonese (HKC), covering interrogatives and statements. Sentences are sourced from online Cantonese teaching materials and classical commercial slogans. It includes approxiametly 198 seconds of audio recorded by a female native speaker of HKC. The sampling rate is 44.1 kHz with 16-bit resolution. Transcription in Jyutping were also provided.
Issues Encountered & Solution
I… See the full description on the dataset page: https://huggingface.co/datasets/kakiso/KarenSo_CantoneseRecordings.karenTTShuberman_lab_Dr._Karen_Parker_The_Causes__Treatments_for_Autismnlp-final-project-activations-3stephuberman_lab_Dr__Karen_Parker_The_Causes__Treatments_for_AutismPython_Like_A_Prodialect_model_data
shanghai-binary dataset
Train/test splits for Shanghai vs Not-Shanghai binary classification.
Contents
data/train.parquet
data/test.parquet
Each row contains:
audio: float array (mono)
sampling_rate: 16000
dialect/label: label (Shanghai=1, else 0)
maksim-garetski-rodnae-karenne-andrei-kaliada
Роднае карэнне
Metadata
Author: Максім Гарэцкі
Title: Роднае карэнне
Narrator: Андрэй Каляда
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size:… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/maksim-garetski-rodnae-karenne-andrei-kaliada.maksim-garetski-rodnae-karenne-valer-mazynski
Роднае карэнне
Metadata
Author: Максім Гарэцкі
Title: Роднае карэнне
Narrator: Валер Мазынскі
Source Group: Аўдыёкнігі
Source: rutracker.org
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/maksim-garetski-rodnae-karenne-valer-mazynski.karen-llama3-datasetkarenTTS-snac24kMuseumDatasetdatoslogin_wireframe_json
