datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nrvbench-review
NR Video Editing Benchmark
This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions.
The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.amd-nrGIT-SCRAPED
🚀 SKT-NRS / GIT-SCRAPED
This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets.
Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures.
📂 Repository Structure
All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.sphere_cohere_embed-english-v3.0moe-5g-nrxreddit_question_best_answersQuestion & question body together with the best answers to that question from Reddit.
The score for the question / answer is the upvote count (i.e. positive-negative upvotes).
Only questions / answers that have these properties were extracted:
min_score = 3
min_title_len = 20
min_body_len = 100
FBHM
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
Accepted at EMNLP 2026 Main 🎉
Authors: Paramananda Bhaskar*, Naquee Rizwan*, Daksh Jogchand, Saurabh Kumar Pandey, Animesh Mukherjee(*) denotes equal contribution
Left: suite of 5,000 FBHM memes spread across 25 functionalities. Each tile presents the functionality number, its description and the corresponding number of memes in that functionality. Right: examples of constructing ten memes… See the full description on the dataset page: https://huggingface.co/datasets/nrizwan/FBHM.nrk_quiz_qa
Dataset Card for NRK-Quiz-QA
Dataset Details
Dataset Description
NRK-Quiz-QA is a multiple-choice question answering (QA) dataset designed for zero-shot evaluation of language models' Norwegian-specific and world knowledge. It comprises 4.9k examples from over 500 quizzes on Norwegian language and culture, spanning both written standards of Norwegian: Bokmål and Nynorsk (the minority variant). These quizzes are sourced from NRK, the national public broadcaster… See the full description on the dataset page: https://huggingface.co/datasets/ltg/nrk_quiz_qa.nr-bundles-public
nr-bundles-public
A curated, multi-modal dataset of blockchain validator infrastructure under attack and benign workloads. 1227 bundles across 40 chains and network stacks (Aptos, Bitcoin, bnb-smart-chain, Cardano, casper, celestia, conflux, Cosmos, Dogecoin, dusk, Ethereum, ethereum-consensus-layer, Filecoin, http2, http3, ic, icon, IOTA, ipfs, kaspa, libp2p, Litecoin, Monero, namada, Near, nimiq, optimism, polkadot, polkadot-substrate, polygon-pos, qtum, quic, Solana, Solana… See the full description on the dataset page: https://huggingface.co/datasets/NullRabbit/nr-bundles-public.microrpusvn-ocr-documents-eval
vn-ocr-documents-eval v0.3
107 single-page Vietnamese documents for evaluating PDF / image → DOCX
OCR pipelines. Six configs covering the full register matrix
(formal + business + conversational + literary) plus real PD scans
and synthetic receipts.
Config
n
Source
License
real
9
chinhphu.vn + hanoi.gov.vn signed scans
Public Domain (Luật SHTT VN, Điều 15)
formal
24
UDHR-vie articles + scan artifacts
CC0 (rendered) — UDHR text is PD
news_business
24
wiki_vi article… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-ocr-documents-eval.wikipedia_0715_clean_cohere_embed-english-v3.0nr_ahr_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_ahr_tox21.nr_aromatase_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_aromatase_tox21.SoccerNet_Field_SegmentationProcessed data from the Soccernet 2023 dataset. Processing notebook is included in this repo.
To see an example:
def show_item(item):
fig, axs = plt.subplots(nrows = 1, ncols = 4, figsize = (20, 4))
axs[0].imshow(item['image'])
axs[0].set_title("Image")
axs[0].axis('off')
axs[1].imshow(overlay_mask(item['image'], item['outlines']))
axs[1].set_title("Outlines")
axs[1].axis('off')
axs[2].imshow(show_segments(item['segments']))
axs[2].set_title("Segments")… See the full description on the dataset page: https://huggingface.co/datasets/nreHieW/SoccerNet_Field_Segmentation.nras-cypa-macrocyclic-glues-GA-II
NRAS–Cyclophilin A Macrocyclic Glue Designs (GA-II)
Why this target matters. NRAS-mutant melanoma has no approved targeted therapy and poor outcomes once immunotherapy fails; RAS(ON) tri-complex glues are among the very few mechanisms that engage NRAS at all.
180 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the NRAS·Cyclophilin A protein–protein interface, with macrocyclic ring closure imposed during generation.
Each molecule was… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/nras-cypa-macrocyclic-glues-GA-II.elephantNREL_Sky_Imagery
NREL SRRL Minute-Resolution Sky Imagery Dataset
Homepage
https://huggingface.co/datasets/knl2366/NREL_Sky_Imagery
Paper
Hammond & Korgel (2026), Journal of Data-centric Machine Learning Research
Contact
Joshua E. Hammond (jeh5975@utexas.edu)
Summary
A continuously growing dataset of minute-resolution sky images from the EKO ASI-16 all-sky imager at NREL's Solar Radiation Research Laboratory (SRRL) in Golden, Colorado (39.742°N, 105.180°W, 1829… See the full description on the dataset page: https://huggingface.co/datasets/knl2366/NREL_Sky_Imagery.nri-fin-reasoning
nri-fin-reasoning
A Japanese instruction dataset with reasoning traces from openai/gpt-oss-120b, specialized for the financial domain.
Overview
A large-scale dataset of 632,636 samples (~6.35 billion tokens), featuring multi-turn conversations (up to 3 turns) with explicit reasoning traces. Designed for supervised fine-tuning to improve LLM reasoning in the financial domain.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nri-ai/nri-fin-reasoning.en-sentiment-nrc
GRADIEND English Sentiment (NRC Adjective) Data
Masked tweet contexts where the masked word is a sentiment adjective:
top 10 adjectives per valence attested as spaCy ADJ in
cardiffnlp/tweet_eval
(sentiment), with polarity taken from the NRC Emotion Lexicon
(Mohammad & Turney, 2013) for target selection.
Frozen training artifact for
gradiend.examples.train_sentiment.
Not a discrete emotion taxonomy (joy/anger/…). Binary polarity cloze
over adjectives.
Companion neutrals:… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/en-sentiment-nrc.viirs-sst-daily-nrt
VIIRS/SNPP Daily Sea Surface Temperature — 4 km, Near-Real-Time
Daily global sea surface temperature from the VIIRS instrument aboard Suomi-NPP,
Level-3 Standard Mapped Image at 4 km, as distributed by the NASA Ocean Biology
Processing Group. This is the near-real-time (NRT) feed — produced within hours
of acquisition, with preliminary calibration.
If you are training a model or computing a trend, use the science-quality feed
instead: PranavKonijeti/viirs-sst-daily-nonNRT.
See… See the full description on the dataset page: https://huggingface.co/datasets/PranavKonijeti/viirs-sst-daily-nrt.nr_er_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_er_tox21.nRadioWaveDatasetnr_ar_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_ar_tox21.en-sentiment-nrc-neutral
GRADIEND English Sentiment (NRC) Neutral Data
Filtered tweet_eval texts with no NRC polarity (positive / negative)
lexicon words — not only the top-20 mask targets used by
aieng-lab/en-sentiment-nrc.
For neutral evaluation.
Usage
from datasets import load_dataset
neutral = load_dataset("aieng-lab/en-sentiment-nrc-neutral", split="train")
texts = neutral["text"]
One split: train.
Dataset Details
Description
Neutral evaluation text… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/en-sentiment-nrc-neutral.all-nr-plasmid-training-public
All-NR Clean Plasmid 50Mbp 1:4 Top50-Switchable Dataset actL4000
This public all-NR plasmid-vs-host training dataset is freshly sampled from the clean NR profile skani_minaf80_ani99_afsym97p5_affull99_afpart95. It is not derived from a curr1 dataset. Plasmid sampling uses a 50Mbp host-genus budget and the strict training_clean config targets four clean host negatives per sampleable plasmid-positive segment after minimap2 filtering.
This upload was produced with config group… See the full description on the dataset page: https://huggingface.co/datasets/neuralbioinfo/all-nr-plasmid-training-public.B24-ota-v2nanorpusanylearning-data
AnyLearning datasets
This repository contains reproducible sample datasets used to develop and test
AnyLearning OSS.
Dataset licenses are recorded in LICENSES.md. The repository's
scripts and original documentation are Apache-2.0, but that license does not
override the terms of any dataset. Check the dataset license before use.
Licence-cleared
Task
Dataset
Licence
Image classification
ZhangLabData: Chest X-Ray
CC BY 4.0
Object detection
Safety Helmet… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/anylearning-data.dolphin
