datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
practice-radar-behavioral-health-npi-sample
New behavioral-health organization NPIs — weekly NPPES sample
A 15-row public sample from a weekly, reproducible selection of newly enumerated Type 2 behavioral-health organizations in the U.S. Centers for Medicare & Medicaid Services National Plan and Provider Enumeration System (NPPES).
Edition at a glance
Measured period: July 6–12, 2026
New Type 2 organizations screened: 2,722
Behavioral-health organizations selected: 486
States and territories represented:… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/practice-radar-behavioral-health-npi-sample.sec-nport
SEC Form N-PORT Data Sets
Monthly portfolio holdings reported by registered investment companies and ETFs
on Form N-PORT, published by the U.S. SEC as quarterly structured data sets
and mirrored here as typed, partitioned Parquet — queryable directly from DuckDB.
Source: SEC Form N-PORT Data Sets — public domain (U.S. Government work)
https://www.sec.gov/data-research/sec-markets-data/form-n-port-data-sets
Coverage: October 2019 onward, refreshed quarterly
Format: one Parquet… See the full description on the dataset page: https://huggingface.co/datasets/trader298/sec-nport.digenai-nppe-datasetdigenai-nppe-datasetdlgenai-nppe2-datasetllama-9b-bulk-npzOMat24_train_aimd_from_PBE_1000_npt
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 1000 npt. ColabFit, 2025. https://doi.org/10.60732/25f16f85
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_jqrkc9e7cgmh_0
Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_1000_npt.gpt-5.5-agentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
gpt 5.5 Agent Traces
This directory contains raw agent trace files generated by teich. (I also dropped in some of my own personal traces)
All assistant responses were generated by openai/gpt-5.5.
JSONL files: 88
Training-ready tools
A complete configured tools schema snapshot is embedded in the… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/gpt-5.5-agent.open-npm
npm Registry - Complete Package Archive
Every npm package with full metadata, versions, dependencies, and download stats
What is it?
This dataset contains a comprehensive snapshot of the npm registry, the default package manager for Node.js and the largest software package registry in the world. npm hosts millions of packages and serves billions of downloads every week. If you have ever run npm install, you have used the registry that this dataset mirrors.
The archive… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-npm.dlgenai-nppe-datasetOMat24_train_aimd_from_PBE_3000_npt
Cite this dataset Barroso-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Zitnick, C. L., and Ulissi, Z. W. OMat24 train aimd from PBE 3000 npt. ColabFit, 2025. https://doi.org/10.60732/edd12490
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_6xvvh8yl7rfd_0
Visit the ColabFit Exchange to search… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/OMat24_train_aimd_from_PBE_3000_npt.nphard_tsp2ECtHR-NPD
ECtHR-NPD — unified current release
One current dataset, not separate v1.0/v1.1 releases. This author-approved
distribution is available for manual review. It preserves the paper's
14,575-case cohort and original partitions, with documented amount corrections.
It is not a byte-identical copy of the targets used for the original experiments.
Data
data/case_level.csv: 14,575 cases; 33 columns.
data/applicant_level.csv: 44,581 applicant source units; 14 columns.… See the full description on the dataset page: https://huggingface.co/datasets/YanyiPU716/ECtHR-NPD.NP_Solutions_v2
🔬 COINjecture NP Solutions Dataset v2
Institutional-Grade Blockchain Research Data
A comprehensive, real-time dataset of NP-complete problem solutions generated through Proof-of-Useful-Work (PoUW) blockchain consensus
Overview • Data Schema • Metrics Categories • Usage • Citation
📋 Overview
This dataset contains institutional-grade metrics from the COINjecture Network B blockchain, which implements a novel Proof-of-Useful-Work (PoUW) consensus… See the full description on the dataset page: https://huggingface.co/datasets/COINjecture/NP_Solutions_v2.fund-holdings-nport
US Fund and ETF Holdings — Form N-PORT
Every position of every US mutual fund and ETF, monthly, including the
bonds, loans, asset-backed paper and derivatives that 13F does not report at
all.
145 630 731 positions · 17 513 funds · 341 049 monthly reports ·
2019-09-30 to 2026-05-31
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
Why this and not 13F
13F is what everyone uses because it is what everyone knows about. It… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/fund-holdings-nport.finemath
📐 FineMath
What is it?
📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+) of mathematical educational content filtered from CommonCrawl. To curate this dataset, we trained a mathematical content classifier using annotations generated by LLama-3.1-70B-Instruct. We used the classifier to retain only the most educational mathematics content, focusing on clear explanations and step-by-step problem solving rather than… See the full description on the dataset page: https://huggingface.co/datasets/NP235/finemath.NP_Solutions_v4
COINjecture NP Solutions v4
Dataset Description
This dataset contains verified solutions to NP-hard computational problems from the COINjecture Network B blockchain.
Version 4 Features
ADZDB Storage: File-based Append-Delete-Zero Database for efficient block storage
Unified Streaming: All problem types in one continuous dataset
Real-time Updates: Solutions streamed as blocks are mined
Problem Types
TSP (Traveling Salesman Problem)
3SAT (Boolean… See the full description on the dataset page: https://huggingface.co/datasets/COINjecture/NP_Solutions_v4.NPRVideoEmbeddingstitans_NPC
Titans - Pytorch
Unofficial implementation of Titans in Pytorch. Will also contain some explorations into architectures beyond their simple 1-4 layer MLP for the neural memory module, if it works well to any degree.
Paper review by Yannic
Quick Colab Run
Appreciation
Eryk for sharing his early experimental results with me, positive for 2 layer MLP
Install
$ pip install titans-pytorch
Usage
import torch
from titans_pytorch import… See the full description on the dataset page: https://huggingface.co/datasets/ChipYTY/titans_NPC.NP_Solutions_v3
🔬 COINjecture NP Solutions Dataset v3
Institutional-Grade Blockchain Research Data
A comprehensive, real-time dataset of NP-complete problem solutions generated through Proof-of-Useful-Work (PoUW) blockchain consensus
Overview • Data Schema • Metrics • Pipeline• Usage • Citation
📋 Overview
This dataset contains institutional-grade metrics from the COINjecture Network B blockchain, which implements a novel Proof-of-Useful-Work (PoUW) consensus mechanism.… See the full description on the dataset page: https://huggingface.co/datasets/COINjecture/NP_Solutions_v3.World_Bank_GDP_by_Country_and_Continent_2000-2024
World Bank GDP by Country & Continent (2000–2025)
A clean, visual, non-technical analysis of World Bank GDP (current US$) organized by country and summarized to a 7-continent view. All values are expressed in USD (billions) for readability.
Why this repo exists
To provide a single place to:
Extract authoritative GDP data from the World Bank (2000–2025),
Publish & share a tidy dataset on Kaggle,
Analyze & present insights in a visual, plain-English notebook.… See the full description on the dataset page: https://huggingface.co/datasets/npaleti2002/World_Bank_GDP_by_Country_and_Continent_2000-2024.npy_file_hsrremote-worker-productivity
Key Features:
Primary Research Focus:
Age vs Productivity correlation
Years of Experience impact on remote work effectiveness
WFH Days per Week optimal balance analysis
Multiple productivity metrics (not just one score)
Dataset Highlights:
1,500 rows - Perfect size for analysis
30+ columns - Rich feature set
Realistic correlations - Built-in meaningful relationships
Clean data - No missing values, proper data types
Multiple target variables - 5 different… See the full description on the dataset page: https://huggingface.co/datasets/nprak26/remote-worker-productivity.NPM-Artifact-Explanation-Benchmark
NPM-Artifact-Explanation-Benchmark
English
NPM-Artifact-Explanation-Benchmark is a cross-category multimodal corpus and benchmark resource for Chinese cultural artifact understanding and explanation.
This release contains 28,826 cleaned artifact records derived from National Palace Museum source records' opendata (https://digitalarchive.npm.gov.tw/opendata/). Each record includes structured artifact metadata, image URLs, source record URLs, and human-written… See the full description on the dataset page: https://huggingface.co/datasets/shunanhe/NPM-Artifact-Explanation-Benchmark.top-npm-packagesTop NPM Packages Dataset
This dataset contains a snapshot of Top 5200+ popular node packages hosted on Node Package Manager
The dataset was scraped in October-2024.
We aim to use this dataset to perform analysis and identify trends and get a bird's eye view of nodejs ecosystem.
Mantainers:
Nishritha Damera
eleusis-calibrated-rules
Eleusis Calibrated Rules — 100-turn reward calibration
A calibrated rule dataset for the single-player Eleusis inductive-reasoning
environment. It extends the 26-rule Hugging Face benchmark with controlled
static, transition, conditional, periodic, chunk, higher-order history, global
history, and compositional rule families.
Source benchmark: Hugging Face Eleusis.
Dataset version: v2.1-frontier-calibrated-100turn-20260812Protocol: eleusis-100-v11
The structural, GPT Sol… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-calibrated-rules.nevada-federal-contractors
Nevada Federal Contractors - FedComp Index (September 2026)
Classified dataset of 779 federal contractors registered in Nevada as of September 2026, covering five years of USASpending base contract award data. Each contractor is assigned a Posture Class (1-4) based on two axes: base contract volume and base contract frequency.
This dataset is updated weekly as new award data becomes available.
Source
All data is sourced from USASpending.gov and SAM.gov.… See the full description on the dataset page: https://huggingface.co/datasets/npetro6/nevada-federal-contractors.nptqlzdcsx
nptqlzdcsx
This dataset was generated using a phospho dev kit.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
air-hockey-amdThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 21,
"total_frames": 36046,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/npaka/air-hockey-amd.textq-german
TextQ-German
TextQ investigates how people perceive the quality of machine-generated German text and how these subjective judgments can be modeled automatically. We identified task-specific quality dimensions, quantified them through user ratings, and developed models that predict perceived quality for new generated texts.
TextQ-German is a dataset suite for studying the Quality of Experience (QoE) of machine-generated German text. It covers two Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/nphamdinh/textq-german.
