datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
picotron_bench
Wrapup results:
compute mfu for each results
change status of jobs
Push to hub
add scripts reproductible
add topology
bandwidth etc
nan-nli
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
Natural Language Inference
Text Classification
Languages
en
Dataset Structure
Data Instances
Data Fields
premise:
hypothesis:
label:
Data Splits
Evaluation: 258 samples
Dataset Creation
Curation Rationale
Extracting samples corresponding to different linguistics constructions of… See the full description on the dataset page: https://huggingface.co/datasets/joey234/nan-nli.nangang_sports_centerlocal-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.R3C-Universal-Nanofabrication
R3C — Reservoir-Rank and Reaction-Repair Compiler
Reservoir-rank engineering and finite-bandwidth reaction repair toward programmable nanofabrication
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiVersion: 1.0.0 — 17 September 2026Repository: PureOne/R3C-Universal-NanofabricationResource type: theoretical research report + reproducible software + entirely synthetic datasetsScientific status: conditional finite-model theory; no physical… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/R3C-Universal-Nanofabrication.GUEinsult-datasetcovid_fake_newsConstraint@AAAI2021 - COVID19 Fake News Detection in English
@misc{patwa2020fighting,
title={Fighting an Infodemic: COVID-19 Fake News Dataset},
author={Parth Patwa and Shivam Sharma and Srinivas PYKL and Vineeth Guptha and Gitanjali Kumari and Md Shad Akhtar and Asif Ekbal and Amitava Das and Tanmoy Chakraborty},
year={2020},
eprint={2011.03327},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
NT_dataablation_tokensvlwnc-if-vf-universal-class-nanofabricator-v1
Vaelorium Luminex / The Weave NooCathedral InfiLattice / Veyrglass Fabricator "VLWNC-IF-VF" - Universal Class
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiScientific release: v1.0.0 · Hub packaging: hf.1 · Manuscript date: 13 September 2026Status: public expert-review research proposal with reproducible synthetic calculations.
Light-addressed physical compilation for heterogeneous fabrication: a proposed multi-cartridge “light printer in a box” combining… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/vlwnc-if-vf-universal-class-nanofabricator-v1.GPTMicro-Nanowire-Sintering
GPTMicro — Nanowire Sintering & Symbolic Regression Dataset
Curated data for data-driven discovery of governing equations in nanowire
sintering. It pairs raw molecular-dynamics (MD) trajectories with the ML-ready
train/validation/test splits used to learn closed-form models for the sintering
dynamics (change in flattening ddelta and rotation dtheta) and for two
effective material properties (effective diffusion coefficient D_eff and
effective relaxation/viscosity coefficient… See the full description on the dataset page: https://huggingface.co/datasets/Kiarash99/GPTMicro-Nanowire-Sintering.feni-nanoparticles
Machine learning-based prediction of FeNi nanoparticle magnetization
Public data for "Machine learning-based prediction of FeNi nanoparticle magnetization", F. Williamson et al., Journal of Materials Research and Technology (2024). https://doi.org/10.1016/j.jmrt.2024.10.142.
ML Scripts
ML scripts are available on GitHub.
Data
Nanoparticles were simulated using LAMMPS.
A single LAMMPS input script from this extended repository was modiffied to obtain various NP… See the full description on the dataset page: https://huggingface.co/datasets/Ailurion/feni-nanoparticles.nanobody_type
Nanobody Type Classification Dataset
Dataset Overview
This dataset helps classify different types of single-domain antibodies (sDAbs). Nanobodies are a special type of single-domain antibody, mainly from camelid heavy-chain antibodies. Correctly identifying different types of sDAbs is important for understanding their structural properties, binding ability, and potential applications.
The dataset contains sDAb sequences from different sources, with the goal of classifying… See the full description on the dataset page: https://huggingface.co/datasets/ZYMScott/nanobody_type.atcoder_abc_contests
Notification
Atcoder is selling this data now. If you are interested in accessing it please contact them.
Dataset Summary
This dataset aims to facilitate the creation of sophisticated, multi-turn dialogue datasets focused on coding
that could be used for training reasoning Large Language Models (LLMs), particularly for Supervised Fine-Tuning (SFT) and Knowledge Distillation techniques.
It also serves as a robust foundation for problem-solving in Large Language… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/atcoder_abc_contests.Nemotron-3-Nano-RL-Training-Blend-prompt-only
Nemotron-3-Nano-RL-Training-Blend-prompt-only
Prompt-only extraction from nvidia/Nemotron-3-Nano-RL-Training-Blend.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-3-Nano-RL-Training-Blend-prompt-only.finetune_datananochat-german-eval-data
nanochat German: Evaluation Data
This repository hosts the translated evaluation data used for assessing a German nanochat model.
Background information: The original nanochat implementation by Andrej Karpathy uses the "Mosaic Eval Gauntlet" (version v0.3.0) benchmark. More information about this benchmark can be found in Mosaic's blog post and this paper.
To evaluate our German nanochat model, we translated several datasets to German using Gemini 2.5 Pro. While this translation… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/nanochat-german-eval-data.Nanobody_Sequence_DatasetRepresentative sequence dataset extracted after clustering of Integrated NANOBODY® Database for Immunoinformatics (INDI) by MMseqs2 program.
For more information, please visit https://github.com/DynaX-C/EvoNB.
IPC_and_BNS_transformationft_dataNanoAntimicrobialDB
NanoAntimicrobialDB
The NanoAntimicrobialDB dataset contains comprehensive data on the antimicrobial properties of silver (Ag) and gold (Au) nanoparticles. This dataset is structured into four distinct files, each representing the Minimum Inhibitory Concentration (MIC) and Minimum Bactericidal Concentration (MBC) values for both Ag and Au nanoparticles against various bacterial strains. The dataset was developed to support research in nanomaterial synthesis, bioactivity assessment… See the full description on the dataset page: https://huggingface.co/datasets/QueHayRich/NanoAntimicrobialDB.nan_twfinguard-finance-injection-dataset
FinGuard: Finance-Specific Prompt Injection Detection Dataset
Dataset Summary
FinGuard is the first open dataset for detecting prompt injection attacks
against agentic financial AI systems. It combines 6 public datasets with
synthetically generated finance-specific attack examples across 4 enterprise
agent types.
Dataset Structure
Split
Rows
SAFE
ATTACK
Train
10,699
5,375 (50.2%)
5,324 (49.8%)
Test
3,047
2,006 (65.8%)
1,041 (34.2%)… See the full description on the dataset page: https://huggingface.co/datasets/nandhak12/finguard-finance-injection-dataset.nano-syn-catalytic-nanoparticle-activation-failure-v0.1Goal
Predict catalytic activation failure.
This is the failure mode wherenanoparticles sinter or restructureand activity collapses.
The key signal
Conversion and structure metricsstop moving together.
Inputs
reactor temperature
space velocity
conversion and selectivity
in situ XRD crystallite size
BET surface area
sintering indicator
Required outputs
activation_coherence_score
decoupling_flag
decoupling_type
activity_loss_horizon_hr
sintering_risk
stabilization_intervention_set
Decoupling… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/nano-syn-catalytic-nanoparticle-activation-failure-v0.1.standardized-food-nutrition-dataset
Daily Food & Nutrition Dataset
A comprehensive dataset containing nutritional information for 587 common food items with detailed macronutrient breakdowns.
Dataset Overview
This dataset provides detailed nutritional composition for a wide variety of foods, including calories, protein, carbohydrates, fat, fiber, sugars, sodium, and cholesterol content. Each food item is categorized and associated with meal types for easy filtering and analysis.
Dataset Statistics:… See the full description on the dataset page: https://huggingface.co/datasets/Nandhini3737/standardized-food-nutrition-dataset.nanbeige-ai-errors
Technical Challenge: Blind Spots of Frontier Models
This dataset was created as part of a technical challenge to identify and document the "blind spots" of a recent, moderately-sized base model. The goal was to browse models released in the last 6 months (between 0.6B and 6B parameters), select one, and systematically probe its failures to understand its limitations.
Dataset: Nanbeige4.1-3B AI Errors
This dataset contains 10 examples where the Nanbeige/Nanbeige4.1-3B… See the full description on the dataset page: https://huggingface.co/datasets/abdulrahman245/nanbeige-ai-errors.turkish-social-media-offensive-dataset
Overwiev
It is a 4-class Turkish bullying data set obtained from Twitter.
Cinsiyetçilik
Irkçılık
Kızdırma
Nötr
Sum
601
490
910
1387
3388
Authors
Seyma SARIGIL: seymasargil@gmail.com
Elif SARIGIL KARA: elifsarigil@gmail.com
Murat KOKLU: mkoklu@selcuk.edu.tr
Alaaddin Erdinç DAL: aerdincdal@icloud.com
harvey-nanobody-polyreactivity
Harvey Nanobody Polyreactivity Dataset (Novo Nordisk Preprocessing)
Dataset Summary
This dataset contains 141,021 nanobody (VHH) sequences with binary polyreactivity labels, preprocessed according to the methodology described in Sakhnini et al. 2025 (Novo Nordisk & University of Cambridge). The dataset was originally published by Harvey et al. 2022 and contains synthetic nanobodies assessed by PSR (Poly-Specificity Reagent) assay via FACS sorting and deep sequencing.
This… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/harvey-nanobody-polyreactivity.nano-syn-nanoparticle-surface-ligand-decoupling-v0.1Goal
Detect ligand decoupling.
This is the failure mode where:
Ligand metricsstop predictingfunctional stability and targeting.
A batch can look acceptableby ligand density alonewhile zeta potential, size, and bindingdrift into failure.
Inputs
ligand density
zeta potential
hydrodynamic diameter
aggregation rate
in vitro binding efficiency
stability window
Required outputs
ligand_coherence_score
decoupling_flag
decoupling_type
targeting_failure_probability
aggregation_risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/nano-syn-nanoparticle-surface-ligand-decoupling-v0.1.
