datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nrvbench-review
NR Video Editing Benchmark
This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions.
The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.sphere_cohere_embed-english-v3.0reddit_question_best_answersQuestion & question body together with the best answers to that question from Reddit.
The score for the question / answer is the upvote count (i.e. positive-negative upvotes).
Only questions / answers that have these properties were extracted:
min_score = 3
min_title_len = 20
min_body_len = 100
FBHM
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
Accepted at EMNLP 2026 Main 🎉
Authors: Paramananda Bhaskar*, Naquee Rizwan*, Daksh Jogchand, Saurabh Kumar Pandey, Animesh Mukherjee(*) denotes equal contribution
Left: suite of 5,000 FBHM memes spread across 25 functionalities. Each tile presents the functionality number, its description and the corresponding number of memes in that functionality. Right: examples of constructing ten memes… See the full description on the dataset page: https://huggingface.co/datasets/nrizwan/FBHM.nrk_quiz_qa
Dataset Card for NRK-Quiz-QA
Dataset Details
Dataset Description
NRK-Quiz-QA is a multiple-choice question answering (QA) dataset designed for zero-shot evaluation of language models' Norwegian-specific and world knowledge. It comprises 4.9k examples from over 500 quizzes on Norwegian language and culture, spanning both written standards of Norwegian: Bokmål and Nynorsk (the minority variant). These quizzes are sourced from NRK, the national public broadcaster… See the full description on the dataset page: https://huggingface.co/datasets/ltg/nrk_quiz_qa.microrpusvn-ocr-documents-eval
vn-ocr-documents-eval v0.3
107 single-page Vietnamese documents for evaluating PDF / image → DOCX
OCR pipelines. Six configs covering the full register matrix
(formal + business + conversational + literary) plus real PD scans
and synthetic receipts.
Config
n
Source
License
real
9
chinhphu.vn + hanoi.gov.vn signed scans
Public Domain (Luật SHTT VN, Điều 15)
formal
24
UDHR-vie articles + scan artifacts
CC0 (rendered) — UDHR text is PD
news_business
24
wiki_vi article… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-ocr-documents-eval.wikipedia_0715_clean_cohere_embed-english-v3.0nr_ahr_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_ahr_tox21.nr_aromatase_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_aromatase_tox21.nras-cypa-macrocyclic-glues-GA-II
NRAS–Cyclophilin A Macrocyclic Glue Designs (GA-II)
Why this target matters. NRAS-mutant melanoma has no approved targeted therapy and poor outcomes once immunotherapy fails; RAS(ON) tri-complex glues are among the very few mechanisms that engage NRAS at all.
180 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the NRAS·Cyclophilin A protein–protein interface, with macrocyclic ring closure imposed during generation.
Each molecule was… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/nras-cypa-macrocyclic-glues-GA-II.elephantnri-fin-reasoning
nri-fin-reasoning
A Japanese instruction dataset with reasoning traces from openai/gpt-oss-120b, specialized for the financial domain.
Overview
A large-scale dataset of 632,636 samples (~6.35 billion tokens), featuring multi-turn conversations (up to 3 turns) with explicit reasoning traces. Designed for supervised fine-tuning to improve LLM reasoning in the financial domain.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nri-ai/nri-fin-reasoning.en-sentiment-nrc
GRADIEND English Sentiment (NRC Adjective) Data
Masked tweet contexts where the masked word is a sentiment adjective:
top 10 adjectives per valence attested as spaCy ADJ in
cardiffnlp/tweet_eval
(sentiment), with polarity taken from the NRC Emotion Lexicon
(Mohammad & Turney, 2013) for target selection.
Frozen training artifact for
gradiend.examples.train_sentiment.
Not a discrete emotion taxonomy (joy/anger/…). Binary polarity cloze
over adjectives.
Companion neutrals:… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/en-sentiment-nrc.nr_er_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_er_tox21.nRadioWaveDatasetnr_ar_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_ar_tox21.en-sentiment-nrc-neutral
GRADIEND English Sentiment (NRC) Neutral Data
Filtered tweet_eval texts with no NRC polarity (positive / negative)
lexicon words — not only the top-20 mask targets used by
aieng-lab/en-sentiment-nrc.
For neutral evaluation.
Usage
from datasets import load_dataset
neutral = load_dataset("aieng-lab/en-sentiment-nrc-neutral", split="train")
texts = neutral["text"]
One split: train.
Dataset Details
Description
Neutral evaluation text… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/en-sentiment-nrc-neutral.all-nr-plasmid-training-public
All-NR Clean Plasmid 50Mbp 1:4 Top50-Switchable Dataset actL4000
This public all-NR plasmid-vs-host training dataset is freshly sampled from the clean NR profile skani_minaf80_ani99_afsym97p5_affull99_afpart95. It is not derived from a curr1 dataset. Plasmid sampling uses a 50Mbp host-genus budget and the strict training_clean config targets four clean host negatives per sampleable plasmid-positive segment after minimap2 filtering.
This upload was produced with config group… See the full description on the dataset page: https://huggingface.co/datasets/neuralbioinfo/all-nr-plasmid-training-public.B24-ota-v2nanorpusdolphincocostackexchange_cohere_embed-english-v3.0picorpusnrftw-boss-arena
NRFTW Boss Arena — 同步视频 + 20Hz 遥测
📌 勘误(2026-08-31)
初次发布时本数据集包含 13 段视频,其中 6 段实际是纯黑画面已被移除。
原因:OBS 的游戏捕获在 2026-08-29 通宵批次没有挂上钩子,录下的是全黑帧。
当时用「文件体积在增长」判断录制正常 —— 这个判据是错的:NVENC 走固定码率,
纯黑同样会被填满到目标码率(30 分钟纯黑照样 2.1 GB)。
这 6 个 session 的日志完整可用,已在 metadata.jsonl 里标记为 quality: LOGS_ONLY。
现有 7 段视频均已逐个抽帧验证非黑。
No Rest for the Wicked 中 13 场 BOSS 战(其中 7 场带录像),每一帧画面都与同一时刻的模拟状态和玩家输入对齐。
全部战斗发生在同一个固定竞技场,由一个自动化 BOT 完成,因此 BOSS 行为之外的变量被刻意压到最小。
Telemetry from 13 boss fights (7 with… See the full description on the dataset page: https://huggingface.co/datasets/teawhite/nrftw-boss-arena.corpusyapeichang-hotpotqa-filtered-MC_hf_qwen3_8b-NC_5-ET_0-MNT_512-NRS_3-T_1.0-S_0no-asr-eval-data-verbatim-sample
NRK Norwegian Speech Dataset (Sample)
Dataset Description
Note: This is a sample dataset containing a subset of chunks for demonstration and preview purposes.
The full dataset is available privately.
This dataset contains Norwegian speech data from NRK TV sports broadcasts, processed for automatic speech recognition (ASR) evaluation and research.
Dataset Statistics
Total chunks: 102
Episodes: 34
Total duration: 0.22 hours
Chunk types:… See the full description on the dataset page: https://huggingface.co/datasets/NRK-KIHUB/no-asr-eval-data-verbatim-sample.reddit_question_best_answers_langs
