CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LAYEK-143 /Open-Pixel-1T 🌌 Open-Pixel-1T (Visual Atlas) A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training 📑 Dataset Summary Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.imagetext-to-image10M<n<100M8 likes2.7k downloads6mo agoHugging Face02OpenPipe /hacker-news Hacker News posts and comments This is a dataset of all HN posts and comments, current as of November 1, 2023. tabular10M<n<100M28 likes729 downloads2y agoHugging Face03BCCard /privacy-filter-openpii-masking 1. Overview privacy-filter-openpii-masking is a Korean and English entity-detection dataset for fine-tuning token-classification models. It is derived from ai4privacy/pii-masking-openpii-1.5m, relabeled to a 29-label taxonomy, and supplemented with statically authored or contextualized financial, customer-service/VOC, security, identity, and infrastructure scenarios. The dataset provides entity annotations rather than application-specific redaction output. masked_text replaces… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/privacy-filter-openpii-masking.texttoken-classification10K<n<100K7 likes420 downloads22d agoHugging Face04GT-111 /libero_openpi_v30tabular100K<n<1M0 likes312 downloads2mo agoHugging Face05Lots-of-LoRAs /task1631_openpi_answer_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1631_openpi_answer_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1631_openpi_answer_generation.texttext-generation1K<n<10K0 likes131 downloads2y agoHugging Face06ai4privacy /openpii-masking-mini-10k OpenPII Masking Mini 10K A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models. Sampling Methodology Samples were selected using proportional stratified sampling by language: Target count per language = round(lang_proportion × 10,000) — proportional representation. Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.texttoken-classification10K<n<100K3 likes91 downloads6mo agoHugging Face07ahuseynli-17683 /pii-masking-openpii-finance 1. Overview Korean·English financial-domain PII detection dataset for fine-tuning token-classification models, derived from ai4privacy/pii-masking-openpii-1.5m and re-labeled to an 18-class policy taxonomy, then augmented with synthetic finance/card/bank/insurance/security context. Every PII value is synthetic - either invalidated by construction or inherited from the synthetic ai4privacy corpus - so no real personal data is present (see Synthetic-Data Safety). 1.1.… See the full description on the dataset page: https://huggingface.co/datasets/ahuseynli-17683/pii-masking-openpii-finance.texttoken-classification10K<n<100K0 likes71 downloads2mo agoHugging Face08OpenPipe /hacker-news-scraped-storiestabular1M<n<10M1 likes55 downloads2y agoHugging Face09ajaysri /lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_1000This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 1000, "total_frames": 562160, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:1000" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_1000.tabularrobotics100K<n<1M0 likes52 downloads3mo agoHugging Face10affjljoo3581 /pii-masking-openpii-1m-en OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification… See the full description on the dataset page: https://huggingface.co/datasets/affjljoo3581/pii-masking-openpii-1m-en.texttoken-classification100K<n<1M0 likes49 downloads1mo agoHugging Face11ajaysri /lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 1000, "total_frames": 562160, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:1000" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000.tabularrobotics100K<n<1M0 likes47 downloads3mo agoHugging Face12ajaysri /lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 100, "total_frames": 55262, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:100" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100.tabularrobotics10K<n<100K0 likes43 downloads3mo agoHugging Face13OpenPipe /hacker-news-scraped-stories-filteredtabular100K<n<1M1 likes34 downloads2y agoHugging Face14ajaysri /lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 100, "total_frames": 55262, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:100" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_100.tabularrobotics10K<n<100K0 likes32 downloads3mo agoHugging Face15abhi26 /openpipe-dpo-scientific-reasoning Openpipe Dpo Scientific Reasoning This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis. Dataset Description This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe: OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.texttext-generationn<1K0 likes30 downloads1y agoHugging Face16OpenPipe /best-hn-comment-pairs-v2tabular10K<n<100K1 likes20 downloads2y agoHugging Face17Reza2kn /persian-pii-masking-openpii-690k-clean Persian PII-Masking Combined Corpus, Cleaned Cleaned Persian / Iranian PII-masking token-classification data. This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits. Persona-clean rows: 623890 Initial-clean rows: 224956 Dataset Repo Reza2kn/persian-pii-masking-openpii-690k-clean Schema Rows include: source_text masked_text privacy_mask with label, start… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-clean.tabulartoken-classification100K<n<1M1 likes19 downloads4mo agoHugging Face18ajaysri /lego_stack_openpi_oracle_success_flat_lang_subgoal_goalimage_lerobot_v3_100tabular10K<n<100K0 likes19 downloads3mo agoHugging Face19AdamLucek /open-pii-masking-en-us-30k open-pii-masking-en-us-30k A filtered and transformed subset of ai4privacy/open-pii-masking-500k-ai4privacy filtered for only US English examples where all unique PII labels have been removed and replaced with a [PII] mask. This also includes a new info column with metadata about the count of [PII] masks, and a task column defaulting to 'privacy_masking.' This has resulted in a subset of ~30k examples available within this dataset. For full information about the original… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/open-pii-masking-en-us-30k.texttext-generation10K<n<100K2 likes18 downloads11mo agoHugging Face20Reza2kn /persian-pii-masking-openpii-690k-initial-clean Persian PII-Masking Initial-Round Corpus, Cleaned Cleaned Persian / Iranian PII-masking token-classification data. This repo contains the cleaned initial-clean artifact. Audit/pruning summary: Raw rows: 225418 Kept rows: 224956 Hard-excluded rows: 0 Dropped rows from full exact nearest-neighbor components at cosine >= 0.95: 462 Rows with a full-dataset nearest neighbor at cosine >= 0.95 before component pruning: 867 Full-NN fraction before component pruning: 0.003846… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-initial-clean.tabulartoken-classification100K<n<1M2 likes17 downloads4mo agoHugging Face21thomaub /open_pii_masking_fr_datasettabular100K<n<1M0 likes16 downloads10mo agoHugging Face22OpenPipe /best-hn-comment-pairstabular10K<n<100K0 likes15 downloads2y agoHugging Face23OpenPipe /arxiv-categoriestext1M<n<10M0 likes14 downloads2y agoHugging Face24OpenPipe /best-hn-comment-pairs-v1tabular10K<n<100K0 likes13 downloads2y agoHugging Face25abhi26 /openpipe-chat-complete-scientific-reasoning Openpipe Chat Complete Scientific Reasoning This dataset contains 100 high-quality examples for chat completion fine-tuning, formatted for OpenPipe, focused on scientific reasoning and analysis. Dataset Description This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe: OpenAI Chat Format: Standard messages array with… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-chat-complete-scientific-reasoning.texttext-generationn<1K0 likes8 downloads1y agoHugging Face26PuxAI /open-pii-masking-500k-ai4privacy-augmentedtext100K<n<1M0 likes8 downloads6mo agoHugging Face27supergoose /flan_combined_task1631_openpi_answer_generationtext1K<n<10K0 likes4 downloads2y agoHugging Face28cyberdavve /openpii_sv_combined_datasettext10K<n<100K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.