datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phonemizer-dicts
Phonemizer Dicts
Pre-generated IPA dictionaries for GPL-free text-to-phonemes lookup.
Files
en-us.tsv — 124K English (US) words, tab-separated word<TAB>IPA
Provenance
Generated by running espeak-ng over an English wordlist. The TSV is program output; espeak-ng source (GPL-3.0) is not redistributed here.
Regeneration
See scripts/generate-espeak-dict.py in the tts-rd-team repo.
CUHK-X_Small_Model_Track
CUHK-X — Small Model Track
Multimodal human action recognition (classification).
Given a multimodal clip, predict its action class (action_id, 0–39, 40 classes).
Repository layout
.
├── Training/
│ ├── class_mapping.csv # action_id <-> action_name (40 classes)
│ └── data/
│ └── HAR.z01 … HAR.z08 + HAR.zip # multi-volume zip
│ → HAR/data/<modality>/<action>/<user>/<trial>/<files>
└── Testing/
├── data/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/Kevin-Pal/CUHK-X_Small_Model_Track.attackdex-paldeaSingle pokemon datasets containing all the attacks (from levelling or TMs) learnable by the relative monster. All the data refer to the Paldea region and they come from the project discussed in https://medium.com/@virtualmartire/i-built-an-algorithm-that-finds-the-optimal-pokemon-team-01ea152824a9.
pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.africa-synth-cerebral-palsy-synthetic-dataset-all
African Cerebral Palsy Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cerebral-palsy-synthetic-dataset-all.pali-tripitaka-thai-script-siamrath-version
📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม)
ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก
🧾 รายการพระไตรปิฎก
📘 เล่ม ๑–๘: วินัยปิฎก
เล่ม ๑: มหาวิภงฺโค (๑)
เล่ม ๒: มหาวิภงฺโค (๒)
เล่ม ๓: ภิกฺขุนีวิภงฺโค
เล่ม ๔: มหาวคฺโค (๑)
เล่ม ๕: มหาวคฺโค (๒)
เล่ม ๖: จุลฺลวคฺโค (๑)
เล่ม ๗: จุลฺลวคฺโค (๒)
เล่ม ๘: ปริวาโร
📗 เล่ม ๙–๒๕: สุตตันตปิฎก
เล่ม ๙–๑๑: ทีฆนิกาย
เล่ม ๑๒–๑๔: มัชฌิมนิกาย
เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.palmer-penguins
Palmer Penguins
The Palmer penguins dataset by Allison Horst, Alison Hill, and Kristen Gorman was first made publicly available as an R package.
The goal of the Palmer Penguins dataset is to replace the highly overused Iris dataset for data exploration & visualization.
However, now you can use Palmer penguins on huggingface!
License
Data are available by CC-0 license in accordance with the Palmer Station LTER Data Policy and the LTER Data Access Policy for Type I data.… See the full description on the dataset page: https://huggingface.co/datasets/SIH/palmer-penguins.pallas_splitted_18cpali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tipitaka_pali_in_15_scripts
Tipitaka Pali in 15 Scripts
Dataset Summary
This dataset contains the Pali Tipitaka (The Pali Canon), Commentaries (Aṭṭhakathā), Sub-commentaries (Ṭīkā), and related texts (Anya). It covers the fundamental scriptures of Theravada Buddhism.
This repository serves as a Hugging Face mirror and processed version of the open-source XML data provided by the Vipassana Research Institute (VRI). The texts are available in various scripts (including Roman and Myanmar) and are… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_pali_in_15_scripts.plot-palette-100k
Empowering Writers with a Universe of Ideas Plot Palette DataSet HuggingFace » Plot Palette was created to fine-tune large language models for creative writing, generating diverse outputs through iterative loops and seed data. It is designed to be run on a Linux system with systemctl for managing services. Included is the service structure, specific category prompts and ~100k data entries. The dataset is available here or… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/plot-palette-100k.online_retailpali-sinhalapaluiDolma-Paloma
Dataset Card for Dolma-Paloma
Dataset Summary
This dataset was created as part of my Master's thesis research on "Leveraging Model Checkpoints for Membership Inference Attacks on Large Language Models". The aim was to create a clean setup, free from distribution shifts, to study proposed checkpoint MIA methods and their performance using checkpoints from Pythia models.
It was derived from the original Dolma and Paloma datasets by applying reservoir sampling and data… See the full description on the dataset page: https://huggingface.co/datasets/ongsici/Dolma-Paloma.IMDB-Dataset-of-50K-Movie-Reviews-Backuparabic-palestinian-levantine-sample
4FACTORS — Palestinian Levantine Conversational Sample
50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data.
What this is
Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.africa-synth-cancer-palliative-care-africa-all
Palliative Care Access - Africa | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-palliative-care-africa-all.superkart-datasetcolor-palettes
Palette Vault Color Palettes
10,000 four-color palettes, each with its colors expressed in HEX, RGB, HSL and
OKLCH, plus WCAG contrast ratios and derived tags.
Every row links to a page on Palette Vault
where the same palette can be viewed, copied or downloaded.
What this is, plainly
The palettes are generated, not collected. No designer chose them and no
human curated them. They come from a generator that works in OKLCH: a mood
preset fixes a corridor of… See the full description on the dataset page: https://huggingface.co/datasets/palettevault/color-palettes.tourism-package-prediction-app-v2bort_wikipedia
BORT Wikipedia Data
This is the data used to prepare the BORT model, described by the following paper:
Robert Gale, Alexandra C. Salem, Gerasimos Fergadiotis, and Steven Bedrick. 2023. Mixed Orthographic/Phonemic Language Modeling: Beyond Orthographically Restricted Transformers (BORT). In Proceedings of the 8th Workshop on Representation Learning for NLP (RepL4NLP-2023), pages TBD, Online. Association for Computational Linguistics. [paper] [poster]
Additional resources and… See the full description on the dataset page: https://huggingface.co/datasets/palat/bort_wikipedia.flower_smellsPaloAltoAssociates
PaloAltoAssociates
tags: PartnerList, ResellerRegistry, Alberta
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'PaloAltoAssociates' dataset contains a curated list of Palo Alto Networks partner reseller businesses located in Alberta. The dataset includes information such as business names, addresses, contact information, and the specific type of services they offer. This can be particularly useful for a machine learning… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/PaloAltoAssociates.Israeli-Palestinian-Conflict
Dataset Card Creation Guide
Dataset Summary
The Israeli-Palestinian-Conflict dataset is an English-language dataset contains manually collected claims, regarding the Israel-Palestine conflict, annotated both objectively with multi-labels to categorize the content according to common themes in such arguments, and subjectively by their level of impact on a moderately informed citizen. The primary purpose of this dataset is to support Israeli public relations efforts at… See the full description on the dataset page: https://huggingface.co/datasets/avishagnevo/Israeli-Palestinian-Conflict.pali-myanmar-parallel-corpus-1kPalestinian_Truth_arPaloAltoPartnersManitoba
PaloAltoPartnersManitoba
tags: reseller_businesses, Manitoba, classification
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'PaloAltoPartnersManitoba' dataset contains records of all registered businesses in Manitoba that have partnered with Palo Alto Networks to act as resellers. Each record includes essential details about the business, its status as a partner reseller, and additional classification information. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/PaloAltoPartnersManitoba.paleography-web-text-triage-dataset
Paleography Web Text Triage Dataset
This dataset contains 400 manually labeled Chinese text snippets for an embedding-based text classification task in Ancient Chinese paleography corpus construction.
The task is to classify short web or reference snippets into four categories:
Label
Meaning
ksd
Scholarly or explanatory discussion about paleography, oracle bones, bronze inscriptions, excavated texts, or script history.
kpt
Primary transcription material such as… See the full description on the dataset page: https://huggingface.co/datasets/Harry214/paleography-web-text-triage-dataset.Palestinian_Truth_eng
