datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_Mihaiii__Pallas-0.4
Dataset Card for Evaluation run of Mihaiii/Pallas-0.4
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.4 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.4.details_Mihaiii__Pallas-0.3
Dataset Card for Evaluation run of Mihaiii/Pallas-0.3
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.3.details_Mihaiii__Pallas-0.2
Dataset Card for Evaluation run of Mihaiii/Pallas-0.2
Dataset Summary
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.2.details_Mihaiii__Pallas-0.5-LASER-0.4
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5-LASER-0.4
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5-LASER-0.4 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5-LASER-0.4.details_Mihaiii__Pallas-0.5-LASER-0.3
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5-LASER-0.3
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5-LASER-0.3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5-LASER-0.3.Palleset
Palleset - A Synthetic Generated Dataset for Pallet Detection
This dataset contains images and ground truth labels generated using Nvidia IsaacSim. The images are taken from a custom created indoor warehouse scenario in which different types of pallets as well as objects and humans are visibile.
The dataset currently contains 3000 image (2400 training, 600 validation).
Different types of annotations are available (beyond the standard RGB images):
Depth
Semantic Segmentation
2D… See the full description on the dataset page: https://huggingface.co/datasets/eusandre95/Palleset.PALL-VLM-data
PALL-VLM-data — Dental Vision-Language Dataset
The training dataset for Harisundar/PALL-VLM,
a dental vision-language model. It contains 32,884 records over 52,461 images,
formatted as image+text conversations for LLaVA-style instruction tuning.
Curated by: Harisundar R
Used by: Harisundar/PALL-VLM · PALL on GitHub
Language: English
Layout
vlm_train/
├── images/ # 52,461 dental images
├── train.jsonl # 29,667 records
├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.details_Mihaiii__Pallas-0.5-LASER-0.6
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5-LASER-0.6
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5-LASER-0.6 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5-LASER-0.6.BraTS-2024-Complete
BraTS 2024 Complete Prepared Dataset
Brain Tumor Segmentation (Leave a like 💖 if this helped you)
Dataset Description
This is an organized and verified version of the BraTS 2024 challenge datasets, including three tumor types.
Included Datasets
Dataset
Type
Cases
Source
BraTS-GLI
Glioma
1,809
Synapse (Dec 2024)
BraTS-MEN-RT
Meningioma + RT
571
Synapse (Feb 2025)
BraTS-PED
Pediatric
348
Cancer Imaging Archive
Total: 2,728… See the full description on the dataset page: https://huggingface.co/datasets/PallabDev/BraTS-2024-Complete.pallas_splitted_18cdetails_Mihaiii__Pallas-0.5
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5.details_Mihaiii__Pallas-0.5-frankenmerge
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5-frankenmerge
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5-frankenmerge on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5-frankenmerge.details_Mihaiii__Pallas-0.5-LASER-0.2
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5-LASER-0.2
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5-LASER-0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5-LASER-0.2.details_Mihaiii__Pallas-0.5-LASER-0.5
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5-LASER-0.5
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5-LASER-0.5 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5-LASER-0.5.pai-interview-a1-forklift-pallet-transport-vizThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "isaacsim_forkliftC",
"total_episodes": 2,
"total_frames": 5497,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/geonmin-kim/pai-interview-a1-forklift-pallet-transport-viz.Palladium-1M-Preview
💎 Palladium-1M: High-Density Information for Efficient LLM Training
Palladium-1M is a curated dataset of ~1 million high-entropy, high-sophistication documents (13.5GB), mined from the open web using a novel Physics-Based Filtration System.
Unlike standard filters that rely on heuristics or keywords, the Palladium Refinery uses Information Theory (ZSTD Compression Ratios) and Linguistic Density to mathematically distinguish "Signal" from "Noise."
The result is a dataset that trains… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/Palladium-1M-Preview.pallatom-ligand-assets
LevinHarness/pallatom-ligand-assets — public mirror of third-party runtime assets
This dataset is a public mirror of third-party runtime
assets required by the Levin Harness plugin(s) listed below. It is not an
official distribution: nothing here is published under this account's own
terms, and it is not affiliated with or endorsed by any upstream project.
Ownership and licensing
Every file remains the property of its upstream authors.
Each file keeps its… See the full description on the dataset page: https://huggingface.co/datasets/LevinHarness/pallatom-ligand-assets.details_Mihaiii__Pallas-0.5-LASER-exp2-0.1
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5-LASER-exp2-0.1
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5-LASER-exp2-0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5-LASER-exp2-0.1.samromur-500h-test
samromur-500h (test splits only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the test-related
splits of palli23/samromur-500h -- the training pool is not
redistributed here. Drawn from the Samrómur Milljón corpus.
test: the full official test split (9,308 utt., 8.02h).
test_4h: first-in-order subset of… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-500h-test.lettuce-pallets
Lettuce Pallets
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
1,060
Validation
299
Test
151
Total
1,510
Classes (5)
Ready
empty_pod
germination
pod
young
Usage
With LibreYOLO
from libreyolo import LIBREYOLO
# Load a model
model = LIBREYOLO(model_path="libreyoloXnano.pt")
# Train on this… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/lettuce-pallets.samromur-21.05-test
samromur-21.05 (test splits only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the test-related
splits of palli23/samromur-21.05 -- the training pool is not
redistributed here.
test: the official test split (10000 utt., 15.9h, 34% OOV
against the paper's independently-trained 21.05 scaling set).… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-21.05-test.samromur-milljon-test
samromur-milljon (test split only)
Author: Páll Rúnarsson, Reykjavík UniversityHUMI Laboratory – Language and Voice Laboratory
Evaluation-only release accompanying "Scaling Smaller ASR Models Against
Multilingual ASR Giants" (ICASSP 2027). Contains only the official test
split of palli23/samromur-milljon-splits (5017 utt.) -- the training
pool is not redistributed here. Used in the paper as the "Miljon, raw"
zero-shot difficulty anchor (92.5% of its prompts share exact text with… See the full description on the dataset page: https://huggingface.co/datasets/palli23/samromur-milljon-test.pall
PALL — Dental Training Corpus
Open training corpus for PALL-Text, a
dental-domain Llama-3.1-8B. Contains three subsets covering the full
CPT → SFT → DPO post-training pipeline.
Developed by: Harisundar R
License: CC-BY-NC-4.0 (composite corpus; individual sources may carry additional terms)
Language: English (with some multilingual medical Q&A)
Dataset structure
Subset
Schema
Train
Val
Total
cpt
{ "text", "source" }
401,900
4,059
405,959
sft
{… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/pall.details_Mihaiii__Pallas-0.5-LASER-0.1
Dataset Card for Evaluation run of Mihaiii/Pallas-0.5-LASER-0.1
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.5-LASER-0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.5-LASER-0.1.IMDB-Dataset-of-50K-Movie-Reviews-Backupomy_pnp_pallet193cleanpalladium-stem-preview-25k
⚛️ Palladium-STEM (Preview): High-Density Scientific Corpus
"The Top 0.17% of the Open Web."
Overview
This dataset is a 25,000-document preview of the upcoming Palladium-V2 STEM Corpus. It represents the "Platinum Tier" survivors from a pool of 14.8 million scanned documents, selected for high information density, academic rigor, and reasoning capability.
The "Goldilocks" Methodology
Unlike standard web scrapes, this data was processed using a custom… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/palladium-stem-preview-25k.lettuce-pallets
Dataset Card for lettuce-pallets
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
lettuce-pallets
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/lettuce-pallets.africa-synth-cancer-palliative-care-africa-all
Palliative Care Access - Africa | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-palliative-care-africa-all.palliative-care-pain-management
Palliative Care & Pain Management (Opioid Access, WHO Ladder, Quality of Life) | Africa (World Health Organization)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/palliative-care-pain-management.
