ppc
Datasets
All datasets matching “ppc”PPCNet
PPCNet Dataset
Biplanar DRRs + Projection Matrices + Ground-Truth Point Clouds
A curated lumbar spine dataset for 3D point cloud reconstruction from biplanar radiographs.
Overview
This dataset provides paired biplanar DRRs, calibrated 3×4 projection matrices, and dense ground-truth point clouds for 1,037 patients with complete L1–L5 lumbar vertebrae, derived from the publicly available VerSe'19, VerSe'20, and CTSpine1K collections.… See the full description on the dataset page: https://huggingface.co/datasets/ppcnet-dataset/PPCNet.PpcPC
PpcPC
An MTEB dataset
Massive Text Embedding Benchmark
Polish Paraphrase Corpus
Task category
t2t
Domains
Fiction, Non-fiction, Web, Written, Spoken, Social, News
Reference
https://arxiv.org/pdf/2207.12759.pdf
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["PpcPC"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To… See the full description on the dataset page: https://huggingface.co/datasets/mteb/PpcPC.ppc
PPC - Polish Paraphrase Corpus
Dataset Summary
Polish Paraphrase Corpus contains 7000 manually labeled sentence pairs. The dataset was divided into training, validation and test splits. The training part includes 5000 examples, while the other parts contain 1000 examples each. The main purpose of creating such a dataset was to verify how machine learning models perform in the challenging problem of paraphrase identification, where most records contain semantically… See the full description on the dataset page: https://huggingface.co/datasets/sdadas/ppc.pp-codex-libero-exp-localppc
Political Parliamentary Corpus (PPC)
A multilingual corpus of parliamentary speech, party manifestos and (for German) historical newspapers, exposed with one config per language. Every record follows a single unified schema, so the languages are directly comparable.
44,978,179 documents (~16.5B tokens, chars/4 estimate)
5 languages: de, en, it, pl, tr
Coverage 1803–2026
22 sources, unified schema
Languages / configs
Config
Language
Documents
~Tokens
Years… See the full description on the dataset page: https://huggingface.co/datasets/gagan3012/ppc.pp_cube
pp_cube
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
