datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-paragraphs
wikipedia-paragraphs
wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research.
Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID.
Dataset structure
Configurations
The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.xone-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..)
A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/xone-repository-parallel-en-id-corpus.Crop-Recommendation-Parameters
🌱 Crop Recommendation Dataset
A machine learning dataset for crop recommendation based on soil properties and environmental conditions. The dataset contains measurements of essential soil nutrients and climatic parameters, along with the crop label that is suitable for those conditions.
This dataset can be used for machine learning classification, agricultural analytics, decision-support systems, and smart farming applications.
📌 Dataset Overview
Property… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/Crop-Recommendation-Parameters.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the
exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.CDial-BiasOfficial release of CDial-Bias dataset.
Notation:
Before downloading the dataset, please be aware that: The CDial-Bias Dataset is released for research purpose only and other usages require further permission. Please ensure the usage contributes to improving the safety and fairness of AI technologies. No malicious usage is allowed.
Paper:
https://aclanthology.org/2022.findings-emnlp.262/
Github Repo:
https://github.com/para-zhou/CDial-Bias
Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/para-zhou/CDial-Bias.Romansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.kepler-lc-star-params
Kepler Stellar Lightcurves
Complete Kepler mission lightcurves for ~190,000 stars, paired with stellar parameters.
Each shard contains 500 stars.
Fields
Field
Type
Description
target_id
int
Kepler Input Catalog (KIC) ID
flux
list<float32>
Brightness measurements, normalised to median = 1.0
time
list<float64>
Timestamps (BJD - 2454833.0, days)
n_points
int
Number of measurements
duration_days
float
Observation span
Teff
float
Effective temperature (K)… See the full description on the dataset page: https://huggingface.co/datasets/giggseh/kepler-lc-star-params.Ordering_Constrained_ParaphrasesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Ordering_Constrained_Paraphrases.4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a single paragraph of dense
reasoning about the immediate continuation. Only the part after </think> — the paragraph —
becomes the VALUE. The reasoning inside the think block is dropped.
The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.ParallelKernelBench_Problems
ParallelKernelBench (benchmark)
Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels.
This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Inputs are deterministic — reproduce them with create_input_tensor(rank, world_size, problem_id, base_shape, dtype, trial) from that file; you do not need stored .pt files.
Files
Path
Description… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/ParallelKernelBench_Problems.paraspeechcaps
ParaSpeechCaps
We release ParaSpeechCaps (Paralinguistic Speech Captions), a large-scale dataset that annotates speech utterances with rich style captions
('A male speaker with a husky, raspy voice delivers happy and admiring remarks at a slow speed in a very noisy American environment. His speech is enthusiastic and confident, with occasional high-pitched inflections.').
It supports 59 style tags covering styles like pitch, rhythm, emotion, and more, spanning speaker-level… See the full description on the dataset page: https://huggingface.co/datasets/ajd12342/paraspeechcaps.paranoia
Paranoia
Naturalistic fMRI dataset: 22 subjects listened to a three-part ambiguous
social narrative (~22 minutes total) designed to elicit varying levels of
paranoid interpretation. TR = 1.0 s. 3-second fixation before each run.
This repo mirrors the fmriprep-preprocessed dataset originally distributed via
DataLad at https://gin.g-node.org/ljchang/Paranoia. fmriprep version
1.2.6-1.
Layout
derivatives/fmriprep/sub-tbXXXX/
anat/ func/ figures/
participants.tsv… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/paranoia.Reorient_Block_ParaphrasesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Reorient_Block_Paraphrases.ParallelKernelBench_Problems
ParallelKernelBench (benchmark)
Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels.
This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py.
Files
Path
Description
data/problems.parquet
One row per problem (tabular access)
reference/*.py
Reference solution() implementations
utils/input_output_tensors.py
Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.snack_transfer_clean_trimmed
ParallaxData/snack_transfer_clean_trimmed
Edit trims
Frame-aligned trim of willlyb/snack_transfer. 7 episodes retained; 3684 frames at 50 FPS. Source was not modified.
Videos, numeric rows, and supported recovery sidecars use identical retained frame intervals. Each episode starts at zero. RGB cuts may be re-encoded. Existing tracking loss and reconstructed holds remain; trimming does not make them valid training labels.
See meta/trim_provenance.json for bounds and original… See the full description on the dataset page: https://huggingface.co/datasets/ParallaxData/snack_transfer_clean_trimmed.d1_code_long_paragraphssnack_transferThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "grabette",
"total_episodes": 9,
"total_frames": 5528,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:9"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ParallaxData/snack_transfer.harvest_apples_with_agilex_piper_sim_ee_paraphrases20This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"observation.state": {
"dtype": "float32",
"fps": 25,
"shape": [
8
],
"names": [
"ee.x",
"ee.y",
"ee.z",
"ee.roll",
"ee.pitch",
"ee.yaw"… See the full description on the dataset page: https://huggingface.co/datasets/Faless/harvest_apples_with_agilex_piper_sim_ee_paraphrases20.scipar_parallel_docs
SciPar Parallel Documents
Dataset Description
This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts.
In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories.
This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences.
To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.paraphrasing-french
Attribution
MTEB-format derivative of ismailiismail/paraphrasing_french. Query = phrase; corpus = paraphrase.
ipfs_paraguay_laws_ir
Paraguay legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_paraguay_laws (revision 2e8736d28819fbb4eb711f364e8343ae8b12b4c3) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Paraguay prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_paraguay_laws_ir.door_open_clean_trimmed
ParallaxData/door_open_clean_trimmed
Edit trims
Frame-aligned trim of ParallaxData/door_open. 10 episodes retained; 4796 frames at 50 FPS. Source was not modified.
Videos, numeric rows, and supported recovery sidecars use identical retained frame intervals. Each episode starts at zero. RGB cuts may be re-encoded. Existing tracking loss and reconstructed holds remain; trimming does not make them valid training labels.
See meta/trim_provenance.json for bounds and original… See the full description on the dataset page: https://huggingface.co/datasets/ParallaxData/door_open_clean_trimmed.robocasa-100demos-5chosen-tasks_params_virtual_views_splattedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 364,
"total_frames": 92882,
"total_tasks": 96,
"total_videos": 5096,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:364"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tt1225/robocasa-100demos-5chosen-tasks_params_virtual_views_splatted.omni-solar-wind-parameters
OMNI Hourly Solar Wind Parameters
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Merged hourly near-Earth solar wind magnetic field, plasma, energetic particle parameters combined with geomagnetic and solar activity indices from NASA's OMNI dataset. The master bridge dataset for space weather analysis -- it time-aligns IMF, solar wind, and geomagnetic response in a single file.
The OMNI dataset from NASA's Goddard Space… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/omni-solar-wind-parameters.human-ai-parallel-corpus-biber
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles
@misc{reinhart2024llmswritelikehumans,
title={Do LLMs write like humans? Variation in grammatical and rhetorical styles},
author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg},
year={2024},
eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-biber.bashkir-wikipedia-parallel
Bashkir-Russian Wikipedia Parallel Corpus
Sentence-level Bashkir-Russian parallel text from Wikipedia, scored and filtered
for machine translation.
Overview
Sentence-level Bashkir-Russian parallel dataset extracted from the corresponding
Bashkir and Russian Wikipedia dumps dated 2026-08-01. Candidate pairs are scored
for semantic alignment with multilingual sentence encoders (Meta LASER3, Google
LaBSE) and the in-domain Bashkir-Russian Pair Scorer. The filtered… See the full description on the dataset page: https://huggingface.co/datasets/failed09/bashkir-wikipedia-parallel.sliding_door_clean_trimmed
ParallaxData/sliding_door_clean_trimmed
Edit trims
Frame-aligned trim of ParallaxData/sliding_door. 3 episodes retained; 1364 frames at 50 FPS. Source was not modified.
Videos, numeric rows, and supported recovery sidecars use identical retained frame intervals. Each episode starts at zero. RGB cuts may be re-encoded. Existing tracking loss and reconstructed holds remain; trimming does not make them valid training labels.
See meta/trim_provenance.json for bounds and… See the full description on the dataset page: https://huggingface.co/datasets/ParallaxData/sliding_door_clean_trimmed.wiki_paragraphs_norwegian
WIKI Paragraphs Norwegian
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.Brihat_Parashara_Hora_Shastra
