datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Palladium-1M-Preview
💎 Palladium-1M: High-Density Information for Efficient LLM Training
Palladium-1M is a curated dataset of ~1 million high-entropy, high-sophistication documents (13.5GB), mined from the open web using a novel Physics-Based Filtration System.
Unlike standard filters that rely on heuristics or keywords, the Palladium Refinery uses Information Theory (ZSTD Compression Ratios) and Linguistic Density to mathematically distinguish "Signal" from "Noise."
The result is a dataset that trains… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/Palladium-1M-Preview.PALATE
PALATE Dataset
PALATE contains de-identified human–role-playing-agent conversations,
satisfaction annotations, frozen session-level splits, bilingual character
cards, and the scoring rubrics used by the PALATE benchmark.
Related resources:
Code: Zhuyh1139/PALATE
Five user-simulator adapters:
muset-ai/PALATE-LoRA
The dataset stores source annotations rather than ready-to-train examples.
Use the processing command in the PALATE GitHub repository to construct
role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.Palma-1.0
Palma-1.0 Dataset
Comprehensive Database of Global Palm Species
Palma-1.0 is a comprehensive exploration of palm species, including the PalmTraits 1.0 dataset enriched with data from GBIF, iNaturalist, Wikimedia Commons, and Plants of the World Online (POWO).
Dataset Overview
Palma-1.0 contains comprehensive data on 2,557 palm species across 181 genera. The dataset combines morphological traits, taxonomic classification, geographic distribution, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/kitsuiwebster/Palma-1.0.vinaya-pitaka-pali-myanmar-parallel
Vinaya Pitaka: Pali-Myanmar Parallel Dataset
Description
This dataset provides a professionally aligned, paragraph-level parallel corpus of the Vinaya Pitaka (The Code of Monastic Discipline). It features the original Pali text (presented in Myanmar script) alongside its modern Myanmar translation.
The dataset covers all five major volumes of the Vinaya:
Pārājika (ပါရာဇိကပါဠိ / ပါရာဇိကဏ်)
Pācittiya (ပါစိတ္တိယပါဠိ / ပါစိတ်)
Mahāvagga (မဟာဝဂ္ဂပါဠိ / မဟာဝါ)
Cūḷavagga… See the full description on the dataset page: https://huggingface.co/datasets/freococo/vinaya-pitaka-pali-myanmar-parallel.paloalma__ECE-TW3-JRGL-V1-details
Dataset Card for Evaluation run of paloalma/ECE-TW3-JRGL-V1
Dataset automatically created during the evaluation run of model paloalma/ECE-TW3-JRGL-V1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/paloalma__ECE-TW3-JRGL-V1-details.paloalma__TW3-JRGL-v2-details
Dataset Card for Evaluation run of paloalma/TW3-JRGL-v2
Dataset automatically created during the evaluation run of model paloalma/TW3-JRGL-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/paloalma__TW3-JRGL-v2-details.paloalma__ECE-TW3-JRGL-V5-details
Dataset Card for Evaluation run of paloalma/ECE-TW3-JRGL-V5
Dataset automatically created during the evaluation run of model paloalma/ECE-TW3-JRGL-V5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/paloalma__ECE-TW3-JRGL-V5-details.glp1-telehealth-rules-us
GLP-1 Telehealth Rules by U.S. State
Which U.S. jurisdictions (50 states + Washington, D.C.) require a live video
visit to start GLP-1 treatment by telehealth — plus each jurisdiction's medical
board, Medicaid GLP-1 coverage status, and nurse practitioner
prescriptive-authority classification (full/reduced/restricted) verified
against state nurse practice acts with statute citations.
Canonical, always-current source: https://www.pallashealth.co/glp-1/telehealth-rules
Live CSV:… See the full description on the dataset page: https://huggingface.co/datasets/pallas-health/glp1-telehealth-rules-us.paloalma__ECE-TW3-JRGL-V2-details
Dataset Card for Evaluation run of paloalma/ECE-TW3-JRGL-V2
Dataset automatically created during the evaluation run of model paloalma/ECE-TW3-JRGL-V2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/paloalma__ECE-TW3-JRGL-V2-details.paloalma__Le_Triomphant-ECE-TW3-details
Dataset Card for Evaluation run of paloalma/Le_Triomphant-ECE-TW3
Dataset automatically created during the evaluation run of model paloalma/Le_Triomphant-ECE-TW3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/paloalma__Le_Triomphant-ECE-TW3-details.
