CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4privacy /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.texttoken-classification1M<n<10M21 likes2.9k downloads4mo agoHugging Face02LAYEK-143 /Open-Pixel-1T 🌌 Open-Pixel-1T (Visual Atlas) A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training 📑 Dataset Summary Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.imagetext-to-image10M<n<100M8 likes2.7k downloads6mo agoHugging Face03ai4privacy /pii-masking-openpii-1m OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.texttoken-classification1M<n<10M15 likes1.5k downloads6mo agoHugging Face04ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face05OpenPipe /hacker-news Hacker News posts and comments This is a dataset of all HN posts and comments, current as of November 1, 2023. tabular10M<n<100M28 likes729 downloads2y agoHugging Face06BCCard /privacy-filter-openpii-masking 1. Overview privacy-filter-openpii-masking is a Korean and English entity-detection dataset for fine-tuning token-classification models. It is derived from ai4privacy/pii-masking-openpii-1.5m, relabeled to a 29-label taxonomy, and supplemented with statically authored or contextualized financial, customer-service/VOC, security, identity, and infrastructure scenarios. The dataset provides entity annotations rather than application-specific redaction output. masked_text replaces… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/privacy-filter-openpii-masking.texttoken-classification10K<n<100K7 likes420 downloads22d agoHugging Face07GT-111 /libero_openpi_v30tabular100K<n<1M0 likes312 downloads2mo agoHugging Face08danbhf /openpi_sim_pick_place SO-101 Pick and Place Dataset (OpenPi Format) This dataset contains 40 episodes of a simulated SO-101 robot performing pick-and-place tasks, converted to OpenPi/RLDS format for use with Physical Intelligence's Pi0/Pi0.5 models. Source Converted from LeRobot dataset: danbhf/sim_pick_place_merged_40ep Format Each episode is stored as an NPZ file containing: Key Shape Type Description observation/state (N, 6) float32 Joint positions (6 DoF)… See the full description on the dataset page: https://huggingface.co/datasets/danbhf/openpi_sim_pick_place.textn<1K0 likes303 downloads9mo agoHugging Face09ai4privacy /openpii-masking-micro-100k OpenPII Micro: Multilingual PII Masking Sample A micro-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 100,000 90,000 10,000 19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.texttoken-classification100K<n<1M0 likes133 downloads4mo agoHugging Face10Lots-of-LoRAs /task1631_openpi_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1631_openpi_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1631_openpi_answer_generation.texttext-generation1K<n<10K0 likes131 downloads2y agoHugging Face11ai4privacy /openpii-masking-nano-1k OpenPII Nano: Multilingual PII Masking Sample A nano-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 1,000 900 100 19 30 37 7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.texttoken-classification1K<n<10K3 likes107 downloads4mo agoHugging Face12ai4privacy /openpii-masking-mini-10k OpenPII Masking Mini 10K A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models. Sampling Methodology Samples were selected using proportional stratified sampling by language: Target count per language = round(lang_proportion × 10,000) — proportional representation. Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.texttoken-classification10K<n<100K3 likes91 downloads6mo agoHugging Face13ahuseynli-17683 /pii-masking-openpii-finance 1. Overview Korean·English financial-domain PII detection dataset for fine-tuning token-classification models, derived from ai4privacy/pii-masking-openpii-1.5m and re-labeled to an 18-class policy taxonomy, then augmented with synthetic finance/card/bank/insurance/security context. Every PII value is synthetic - either invalidated by construction or inherited from the synthetic ai4privacy corpus - so no real personal data is present (see Synthetic-Data Safety). 1.1.… See the full description on the dataset page: https://huggingface.co/datasets/ahuseynli-17683/pii-masking-openpii-finance.texttoken-classification10K<n<100K0 likes71 downloads2mo agoHugging Face14OpenPipe /hacker-news-scraped-storiestabular1M<n<10M1 likes55 downloads2y agoHugging Face15ajaysri /lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_1000This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 1000, "total_frames": 562160, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:1000" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_1000.tabularrobotics100K<n<1M0 likes52 downloads3mo agoHugging Face16affjljoo3581 /pii-masking-openpii-1m-en OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification… See the full description on the dataset page: https://huggingface.co/datasets/affjljoo3581/pii-masking-openpii-1m-en.texttoken-classification100K<n<1M0 likes49 downloads1mo agoHugging Face17ajaysri /lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 1000, "total_frames": 562160, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:1000" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000.tabularrobotics100K<n<1M0 likes47 downloads3mo agoHugging Face18charisreneec /OpenPIR OpenPIR: Open Predominant Instrument Recognition Dataset OpenPIR is a hand-labeled dataset for predominant instrument recognition (PIR) built from OpenMic-2018, a Creative Commons-licensed collection of 10-second music clips sourced from the Free Music Archive. It was introduced as part of the ICASSP 2026 paper: Leveraging Diffusion U-Net Features for Predominant Instrument RecognitionCharis Cochran, Yeongheon Lee, Youngmoo Kim — Drexel University / University of PennsylvaniaIEEE… See the full description on the dataset page: https://huggingface.co/datasets/charisreneec/OpenPIR.textaudio-classification1K<n<10K0 likes45 downloads5mo agoHugging Face19ajaysri /lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 100, "total_frames": 55262, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:100" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100.tabularrobotics10K<n<100K0 likes43 downloads3mo agoHugging Face20sunshk /OPENPI_DATA_HOMEtextn<1K0 likes36 downloads11mo agoHugging Face21OpenPipe /hacker-news-scraped-stories-filteredtabular100K<n<1M1 likes34 downloads2y agoHugging Face22ajaysri /lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 100, "total_frames": 55262, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:100" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_100.tabularrobotics10K<n<100K0 likes32 downloads3mo agoHugging Face23abhi26 /openpipe-dpo-scientific-reasoning Openpipe Dpo Scientific Reasoning This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis. Dataset Description This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe: OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.texttext-generationn<1K0 likes30 downloads1y agoHugging Face24OpenPipe /best-hn-comment-pairs-v2tabular10K<n<100K1 likes20 downloads2y agoHugging Face25Reza2kn /persian-pii-masking-openpii-690k-clean Persian PII-Masking Combined Corpus, Cleaned Cleaned Persian / Iranian PII-masking token-classification data. This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits. Persona-clean rows: 623890 Initial-clean rows: 224956 Dataset Repo Reza2kn/persian-pii-masking-openpii-690k-clean Schema Rows include: source_text masked_text privacy_mask with label, start… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-clean.tabulartoken-classification100K<n<1M1 likes19 downloads4mo agoHugging Face26ajaysri /lego_stack_openpi_oracle_success_flat_lang_subgoal_goalimage_lerobot_v3_100tabular10K<n<100K0 likes19 downloads3mo agoHugging Face27AdamLucek /open-pii-masking-en-us-30k open-pii-masking-en-us-30k A filtered and transformed subset of ai4privacy/open-pii-masking-500k-ai4privacy filtered for only US English examples where all unique PII labels have been removed and replaced with a [PII] mask. This also includes a new info column with metadata about the count of [PII] masks, and a task column defaulting to 'privacy_masking.' This has resulted in a subset of ~30k examples available within this dataset. For full information about the original… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/open-pii-masking-en-us-30k.texttext-generation10K<n<100K2 likes18 downloads11mo agoHugging Face28Reza2kn /persian-pii-masking-openpii-690k-initial-clean Persian PII-Masking Initial-Round Corpus, Cleaned Cleaned Persian / Iranian PII-masking token-classification data. This repo contains the cleaned initial-clean artifact. Audit/pruning summary: Raw rows: 225418 Kept rows: 224956 Hard-excluded rows: 0 Dropped rows from full exact nearest-neighbor components at cosine >= 0.95: 462 Rows with a full-dataset nearest neighbor at cosine >= 0.95 before component pruning: 867 Full-NN fraction before component pruning: 0.003846… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-initial-clean.tabulartoken-classification100K<n<1M2 likes17 downloads4mo agoHugging Face29thomaub /open_pii_masking_fr_datasettabular100K<n<1M0 likes16 downloads10mo agoHugging Face30OpenPipe /best-hn-comment-pairstabular10K<n<100K0 likes15 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.