datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PPCNet
PPCNet Dataset
Biplanar DRRs + Projection Matrices + Ground-Truth Point Clouds
A curated lumbar spine dataset for 3D point cloud reconstruction from biplanar radiographs.
Overview
This dataset provides paired biplanar DRRs, calibrated 3×4 projection matrices, and dense ground-truth point clouds for 1,037 patients with complete L1–L5 lumbar vertebrae, derived from the publicly available VerSe'19, VerSe'20, and CTSpine1K collections.… See the full description on the dataset page: https://huggingface.co/datasets/ppcnet-dataset/PPCNet.PpcPC
PpcPC
An MTEB dataset
Massive Text Embedding Benchmark
Polish Paraphrase Corpus
Task category
t2t
Domains
Fiction, Non-fiction, Web, Written, Spoken, Social, News
Reference
https://arxiv.org/pdf/2207.12759.pdf
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["PpcPC"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To… See the full description on the dataset page: https://huggingface.co/datasets/mteb/PpcPC.ppc
PPC - Polish Paraphrase Corpus
Dataset Summary
Polish Paraphrase Corpus contains 7000 manually labeled sentence pairs. The dataset was divided into training, validation and test splits. The training part includes 5000 examples, while the other parts contain 1000 examples each. The main purpose of creating such a dataset was to verify how machine learning models perform in the challenging problem of paraphrase identification, where most records contain semantically… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/ppc.pp-codex-libero-exp-localppc
Political Parliamentary Corpus (PPC)
A multilingual corpus of parliamentary speech, party manifestos and (for German) historical newspapers, exposed with one config per language. Every record follows a single unified schema, so the languages are directly comparable.
44,978,179 documents (~16.5B tokens, chars/4 estimate)
5 languages: de, en, it, pl, tr
Coverage 1803–2026
22 sources, unified schema
Languages / configs
Config
Language
Documents
~Tokens
Years… See the full description on the dataset page: https://huggingface.co/datasets/gagan3012/ppc.pp_cube
pp_cube
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
p-p-character-datasetpp_cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"observation.image": {
"dtype": "video",
"shape": [
224,
224,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height": 224… See the full description on the dataset page: https://huggingface.co/datasets/lambdavi/pp_cube.PPCC14ppc-pairclassificationppc
PPC - Polish Paraphrase Corpus
Dataset Summary
Polish Paraphrase Corpus contains 7000 manually labeled sentence pairs. The dataset was divided into training, validation and test splits. The training part includes 5000 examples, while the other parts contain 1000 examples each. The main purpose of creating such a dataset was to verify how machine learning models perform in the challenging problem of paraphrase identification, where most records contain semantically… See the full description on the dataset page: https://huggingface.co/datasets/djstrong/ppc.neupane-ppcpp_cubeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"robot_type": "so-100",
"codebase_version": "v3.0",
"total_episodes": 43,
"total_frames": 21677,
"total_tasks":1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:43"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/juxhinr/pp_cube.PPCC18africa-mauritius-number-of-complaints-received-at-the-ppc-division-by-categ-60a935ee
Number of Complaints Received At the Ppc Division by Categ | Africa (MDPA)
147 rows - 1 Africa country/area - 2012-2021 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 147 rows from MDPA, covering Number of Complaints Received At the Ppc Division by Categ. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-number-of-complaints-received-at-the-ppc-division-by-categ-60a935ee.PPCC04PPCC09africa-mauritius-number-of-complaints-received-at-the-ppc-division-by-categ-d8d26253
Number of Complaints Received At the Ppc Division by Categ | Africa (MDPA)
69 rows - 1 Africa country/area - 2012-2021 - 1 indicator - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 69 rows from MDPA, covering Number of Complaints Received At the Ppc Division by Categ. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-number-of-complaints-received-at-the-ppc-division-by-categ-d8d26253.pp-codex-libero-exp-local-rlPPC-pretrain-corpusPPCC08africa-mauritius-number-of-complaints-received-at-the-ppc-division-by-categ-fc549eb8
Number of Complaints Received At the Ppc Division by Categ | Africa (MDPA)
7 rows - 1 Africa country/area - 2012-2021 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 7 rows from MDPA, covering Number of Complaints Received At the Ppc Division by Categ. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mauritius-number-of-complaints-received-at-the-ppc-division-by-categ-fc549eb8.PPCC21ppCityTestppCityTrainPPCC06PPCC0Xflashdeal_data_PPC_historical_signalpp_classification_balanced_finalPPCC12
