datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Open-Pixel-1T
🌌 Open-Pixel-1T (Visual Atlas)
A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training
📑 Dataset Summary
Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.hacker-news
Hacker News posts and comments
This is a dataset of all HN posts and comments, current as of November 1, 2023.
privacy-filter-openpii-masking
1. Overview
privacy-filter-openpii-masking is a Korean and English entity-detection dataset for fine-tuning token-classification models. It is derived from ai4privacy/pii-masking-openpii-1.5m, relabeled to a 29-label taxonomy, and supplemented with statically authored or contextualized financial, customer-service/VOC, security, identity, and infrastructure scenarios.
The dataset provides entity annotations rather than application-specific redaction output. masked_text replaces… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/privacy-filter-openpii-masking.libero_openpi_v30task1631_openpi_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1631_openpi_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1631_openpi_answer_generation.openpii-masking-mini-10k
OpenPII Masking Mini 10K
A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models.
Sampling Methodology
Samples were selected using proportional stratified sampling by language:
Target count per language = round(lang_proportion × 10,000) — proportional representation.
Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.pii-masking-openpii-finance
1. Overview
Korean·English financial-domain PII detection dataset for fine-tuning token-classification models, derived from ai4privacy/pii-masking-openpii-1.5m and re-labeled to an 18-class policy taxonomy, then augmented with synthetic finance/card/bank/insurance/security context. Every PII value is synthetic - either invalidated by construction or inherited from the synthetic ai4privacy corpus - so no real personal data is present (see Synthetic-Data Safety).
1.1.… See the full description on the dataset page: https://huggingface.co/datasets/ahuseynli-17683/pii-masking-openpii-finance.hacker-news-scraped-storieslego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 1000,
"total_frames": 562160,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_1000.pii-masking-openpii-1m-en
OpenPII 1M — Multilingual PII Masking Dataset
Overview
The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types.
Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification… See the full description on the dataset page: https://huggingface.co/datasets/affjljoo3581/pii-masking-openpii-1m-en.lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 1000,
"total_frames": 562160,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000.lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 100,
"total_frames": 55262,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100.hacker-news-scraped-stories-filteredlego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 100,
"total_frames": 55262,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_100.openpipe-dpo-scientific-reasoning
Openpipe Dpo Scientific Reasoning
This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe:
OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.best-hn-comment-pairs-v2persian-pii-masking-openpii-690k-clean
Persian PII-Masking Combined Corpus, Cleaned
Cleaned Persian / Iranian PII-masking token-classification data.
This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits.
Persona-clean rows: 623890
Initial-clean rows: 224956
Dataset Repo
Reza2kn/persian-pii-masking-openpii-690k-clean
Schema
Rows include:
source_text
masked_text
privacy_mask with label, start… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-clean.lego_stack_openpi_oracle_success_flat_lang_subgoal_goalimage_lerobot_v3_100open-pii-masking-en-us-30k
open-pii-masking-en-us-30k
A filtered and transformed subset of ai4privacy/open-pii-masking-500k-ai4privacy filtered for only US English examples where all unique PII labels have been removed and replaced with a [PII] mask.
This also includes a new info column with metadata about the count of [PII] masks, and a task column defaulting to 'privacy_masking.'
This has resulted in a subset of ~30k examples available within this dataset.
For full information about the original… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/open-pii-masking-en-us-30k.persian-pii-masking-openpii-690k-initial-clean
Persian PII-Masking Initial-Round Corpus, Cleaned
Cleaned Persian / Iranian PII-masking token-classification data.
This repo contains the cleaned initial-clean artifact.
Audit/pruning summary:
Raw rows: 225418
Kept rows: 224956
Hard-excluded rows: 0
Dropped rows from full exact nearest-neighbor components at cosine >= 0.95: 462
Rows with a full-dataset nearest neighbor at cosine >= 0.95 before component pruning: 867
Full-NN fraction before component pruning: 0.003846… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-initial-clean.open_pii_masking_fr_datasetbest-hn-comment-pairsarxiv-categoriesbest-hn-comment-pairs-v1openpipe-chat-complete-scientific-reasoning
Openpipe Chat Complete Scientific Reasoning
This dataset contains 100 high-quality examples for chat completion fine-tuning, formatted for OpenPipe, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe:
OpenAI Chat Format: Standard messages array with… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-chat-complete-scientific-reasoning.open-pii-masking-500k-ai4privacy-augmentedflan_combined_task1631_openpi_answer_generationopenpii_sv_combined_dataset
