datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hhtools_parc_ms
hhtools PARC MS — terrain-aware humanoid motion clips
中文说明
Motion clips in PARC MS layout for human-humanoid-tools (hhtools): per-clip folders with a PARC MSFileData pickle and a static terrain mesh. Suitable for meshmimic / interaction-mesh retargeting (parkour, climbing, box traversal, etc.).
Clips
26,396
Total size
~10.8 GB
Skeleton
15-bone PARC humanoid (humanoid.xml topology)
Terrain
Heightfield + Wavefront OBJ per clip
Source
Converted from PARC… See the full description on the dataset page: https://huggingface.co/datasets/YaojieShen/hhtools_parc_ms.parcelstow
ParcelStow Demonstrations, Checkpoints, and Videos
Comparing Learned Policies with Their Expert Demonstrators Under Temporal Scaling
This Hugging Face repository contains 970,565 control steps from 937 successful
expert demonstrations for three contact-rich manipulation tasks. It accompanies
ParcelStow, an Isaac Lab benchmark
suite that compares the task success of learned manipulation policies with that
of their expert demonstrators across temporal scaling… See the full description on the dataset page: https://huggingface.co/datasets/cenwerem/parcelstow.parczech4speech-segmented
ParCzech4Speech (Sentence-Segmented Variant)
Dataset Summary
ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts.
This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries.
It is derived from the ParCzech 4.0 corpus and AudioPSP 24.01 audio collection.
Using WhisperX and Wav2Vec 2.0… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-segmented.PARCOMED
PARCOMED - PARTAGES Corpus of Open MEdical Documents
This document describes the first version of the commercial corpus.
Overview
The availability of French biomedical data remains a major challenge for improving the multilingual capabilities of large language models (LLMs) in the medical domain.
We introduce and release the PARCOMED corpus, a collection of French biomedical texts compiled from a wide range of sources for commercial use.
While similar datasets have been… See the full description on the dataset page: https://huggingface.co/datasets/HealthDataHub/PARCOMED.PARCOMED_research_only
PARCOMED - PARTAGES Corpus of Open MEdical Documents
This document describes the first version of the research-only corpus.
Overview
The availability of French biomedical data remains a major challenge for improving the multilingual capabilities of large language models (LLMs) in the medical domain.
We introduce and release the PARCOMED_research_only corpus, a collection of French biomedical texts compiled from a wide range of sources for research-only use.
While similar… See the full description on the dataset page: https://huggingface.co/datasets/HealthDataHub/PARCOMED_research_only.us-parcel-layer
US Parcel Layer — The Landrecords.us Nationwide Parcel Dataset
The Landrecords.us National Parcel Dataset is a comprehensive,
standardized geospatial dataset aggregating ~157 million parcel boundaries and their
associated land-ownership and taxation attributes, harmonized from thousands of local
jurisdictions across the United States into a single national schema.
Each record is a land parcel: a polygon representing the boundary of an individual tract of
land as delineated for… See the full description on the dataset page: https://huggingface.co/datasets/landrecords/us-parcel-layer.eitb_parcc
Dataset Card for EiTB-ParCC
Dataset Summary
EiTB-ParCC: Parallel Corpus of Comparable News. A Basque-Spanish parallel corpus provided by
Vicomtech (https://www.vicomtech.org), extracted from comparable news produced by the
Basque public broadcasting group Euskal Irrati Telebista.
Supported Tasks and Leaderboards
Translation.
Languages
The languages in the dataset are:
Spanish (es)
Basque (eu)
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/eitb_parcc.parc_planner_runsparczech4speech-unsegmented
ParCzech4Speech (Unsegmented Variant)
Dataset Summary
ParCzech4Speech (Unsegmented Variant) is a large-scale Czech speech dataset derived from parliamentary recordings and official transcripts.
This variant captures continuous speech segments without enforcing sentence boundaries, making it well-suited for real-world streaming ASR scenarios
and speech modeling tasks that benefit from natural discourse flow.
The dataset is created using a combination of WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-unsegmented.ingush-russianPARCOMED_research_only
PARCOMED - PARTAGES Corpus of Open MEdical Documents
This document describes the first version of the research-only corpus.
Overview
The availability of French biomedical data remains a major challenge for improving the multilingual capabilities of large language models (LLMs) in the medical domain.
We introduce and release the PARCOMED_research_only corpus, a collection of French biomedical texts compiled from a wide range of sources for research-only use.
While similar… See the full description on the dataset page: https://huggingface.co/datasets/bezhanidze/PARCOMED_research_only.Damaged_Parcel_boxesbdnb_2023-11.a_parcelle_METROPOLEbdnb_2023-11.a_rel_batiment_groupe_parcelle_01error-detection-positives
error-detection-positives
This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines positives samples from multiple domains.
Domain Breakdown
gsm8k: 50 samples
math: 53 samples
metamathqa: 93 samples
orca_math: 96 samples
Features
Each example contains:
data_source: The domain/source of the problem (gsm8k, math, metamathqa, orca_math)
question: The… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-positives.P-ARC
P-ARC CSV export (PotARCin Test2)
One CSV file in UTF-8. Each row is one of the fifty P-ARC tasks from PotARCin (t1.json through t50.json). Besides the usual train/test grids, each row includes the fifty-sample bundle from t<n>_samples_50.json as compact JSON (same structure as the file, without the extra whitespace from pretty-printing), plus the generator.py and verifier.py sources from the matching task folder.
Files
File
Description
p_arc_dataset.csv… See the full description on the dataset page: https://huggingface.co/datasets/PotARCin/P-ARC.error-detection-positives_perturbed
error-detection-positives_perturbed
This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines positives_perturbed samples from multiple domains.
Domain Breakdown
gsm8k: 48 samples
math: 42 samples
metamathqa: 72 samples
orca_math: 85 samples
Features
Each example contains:
data_source: The domain/source of the problem (gsm8k, math, metamathqa… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-positives_perturbed.parc2026-track1-texture-smoke-v1
Track 1 texture mask smoke result
Two real selected episodes were decoded and segmented on A100. This repository stores the reproducibility evidence only: masks, fixed split reference, job definitions, summary, and execution log. It does not contain the original videos or constitute the final training dataset.
error-detection-negatives
error-detection-negatives
This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines negatives samples from multiple domains.
Domain Breakdown
gsm8k: 57 samples
math: 44 samples
metamathqa: 59 samples
orca_math: 54 samples
Features
Each example contains:
data_source: The domain/source of the problem (gsm8k, math, metamathqa, orca_math)
question: The… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-negatives.ingush_proverbs
Dataset Card for "ingush_proverbs"
Source
More Information needed
africa-morocco-parc-fixe-global-des-delegations-provinciales-du-departeme-56f3242c
Parc Fixe Global Des Delegations Provinciales Du Departeme | Africa (Morocco Open Data)
61 rows - 1 Africa country/area - time not specified - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 61 rows from Morocco Open Data, covering Parc Fixe Global Des Delegations Provinciales Du Departeme. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-morocco-parc-fixe-global-des-delegations-provinciales-du-departeme-56f3242c.bdnb_2023-11.a_rel_batiment_groupe_parcelle_29bdnb_2023-11.a_rel_batiment_groupe_parcelle_METROPOLEbdnb_2023-11.a_parcelle_01eitb_parcc_with_english
EITB-parcc English (10b)
The parallel corpus EITB-parcc has been (partially) translated from spanish to english using MADLAD400-10b
Size: 2000 sentences
nassau-parcels-last-soldafrica-morocco-parc-de-la-telephonie-fixe-dbc03019
Parc De La Telephonie Fixe | Africa (Morocco Open Data)
836 rows - 1 Africa country/area - 2006-2022 - 14 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 836 rows from Morocco Open Data, covering Parc De La Telephonie Fixe. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Official statistics datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-morocco-parc-de-la-telephonie-fixe-dbc03019.bdnb_2023-11.a_parcelle_89africa-morocco-parc-de-l-internet-b7a98ed7
Parc De L Internet | Africa (Morocco Open Data)
915 rows - 1 Africa country/area - 2006-2021 - 18 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 915 rows from Morocco Open Data, covering Parc De L Internet. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Official statistics datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-morocco-parc-de-l-internet-b7a98ed7.africa-cote-d-ivoire-evolution-du-parc-de-production-d-electricite-de-la-cote-d-588ce0d6
Evolution Du Parc De Production D Electricite De La Cote D | Africa (Cote d'Ivoire DataFair)
231 rows - 1 Africa country/area - 1959-2017 - 2 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 231 rows from Cote d'Ivoire DataFair, covering Evolution Du Parc De Production D Electricite De La Cote D. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cote-d-ivoire-evolution-du-parc-de-production-d-electricite-de-la-cote-d-588ce0d6.
