datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nrvbench-review
NR Video Editing Benchmark
This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions.
The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.nr_ahr_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_ahr_tox21.nr_aromatase_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_aromatase_tox21.nras-cypa-macrocyclic-glues-GA-II
NRAS–Cyclophilin A Macrocyclic Glue Designs (GA-II)
Why this target matters. NRAS-mutant melanoma has no approved targeted therapy and poor outcomes once immunotherapy fails; RAS(ON) tri-complex glues are among the very few mechanisms that engage NRAS at all.
180 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the NRAS·Cyclophilin A protein–protein interface, with macrocyclic ring closure imposed during generation.
Each molecule was… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/nras-cypa-macrocyclic-glues-GA-II.nr_er_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_er_tox21.nr_ar_tox21
Dataset Details
Dataset Description
Tox21 is a data challenge which contains qualitative toxicity measurements
for 7,831 compounds on 12 different targets, such as nuclear receptors and stress
response pathways.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
data source
assay name
Citation
BibTeX:
@article{Huang2017,
doi = {10.3389/fenvs.2017.00003},
url = {https://doi.org/10.3389/fenvs.2017.00003},
year = {2017}… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/nr_ar_tox21.all-nr-plasmid-training-public
All-NR Clean Plasmid 50Mbp 1:4 Top50-Switchable Dataset actL4000
This public all-NR plasmid-vs-host training dataset is freshly sampled from the clean NR profile skani_minaf80_ani99_afsym97p5_affull99_afpart95. It is not derived from a curr1 dataset. Plasmid sampling uses a 50Mbp host-genus budget and the strict training_clean config targets four clean host negatives per sampleable plasmid-positive segment after minimap2 filtering.
This upload was produced with config group… See the full description on the dataset page: https://huggingface.co/datasets/neuralbioinfo/all-nr-plasmid-training-public.anylearning-data
AnyLearning datasets
This repository contains reproducible sample datasets used to develop and test
AnyLearning OSS.
Dataset licenses are recorded in LICENSES.md. The repository's
scripts and original documentation are Apache-2.0, but that license does not
override the terms of any dataset. Check the dataset license before use.
Licence-cleared
Task
Dataset
Licence
Image classification
ZhangLabData: Chest X-Ray
CC BY 4.0
Object detection
Safety Helmet… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/anylearning-data.test_gopro_020926_nr2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1,
"total_frames": 465,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hamish-grant/test_gopro_020926_nr2.nrftw-boss-arena
NRFTW Boss Arena — 同步视频 + 20Hz 遥测
📌 勘误(2026-08-31)
初次发布时本数据集包含 13 段视频,其中 6 段实际是纯黑画面已被移除。
原因:OBS 的游戏捕获在 2026-08-29 通宵批次没有挂上钩子,录下的是全黑帧。
当时用「文件体积在增长」判断录制正常 —— 这个判据是错的:NVENC 走固定码率,
纯黑同样会被填满到目标码率(30 分钟纯黑照样 2.1 GB)。
这 6 个 session 的日志完整可用,已在 metadata.jsonl 里标记为 quality: LOGS_ONLY。
现有 7 段视频均已逐个抽帧验证非黑。
No Rest for the Wicked 中 13 场 BOSS 战(其中 7 场带录像),每一帧画面都与同一时刻的模拟状态和玩家输入对齐。
全部战斗发生在同一个固定竞技场,由一个自动化 BOT 完成,因此 BOSS 行为之外的变量被刻意压到最小。
Telemetry from 13 boss fights (7 with… See the full description on the dataset page: https://huggingface.co/datasets/teawhite/nrftw-boss-arena.yapeichang-hotpotqa-filtered-MC_hf_qwen3_8b-NC_5-ET_0-MNT_512-NRS_3-T_1.0-S_0TMF921-intent-to-config-25k
TMF921-Grounded Intent-to-Network-Configuration Dataset (25K)
The most comprehensive open dataset for training LLMs to translate natural language network intents into spec-compliant 5G/6G configurations.
25,000 samples (22,500 train / 2,500 test) of natural language intents paired with structured network configurations across 6 target specification layers and 8 lifecycle operations, all grounded in real telecom standards.
What Makes This Dataset Unique… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/TMF921-intent-to-config-25k.local_politicsumls-nrp-dataset
Next Relation Prediction on the UMLS KG
The dataset trains and evaluates a model that generates the subsequent relation to pursue (if available) based on a given MedQA question and the relations explored thus far, otherwise indicating the end (END).
telecom-intent-config-sft-10k
Telecom Intent→Config SFT Dataset (10K)
The first open SFT dataset for training LLMs to translate natural language network intents into structured 5G/6G configurations.
This dataset addresses the #1 gap identified in the telecom LLM research landscape: there is no public training dataset for intent-to-policy translation. All existing telecom datasets (TeleQnA, ORANBench-13K, 6G-Bench) are MCQ evaluation benchmarks — not instruction-following format. This dataset fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/telecom-intent-config-sft-10k.NR-short-32k-16-thoughts-4k-8-thoughts-Qwen3-1.7Bfashion-datasetTMF921-intent-to-config-research-sota
TMF921 Intent-to-Config Research SOTA Splits
This dataset is a research-oriented derivative of nraptisss/TMF921-intent-to-config-augmented. It provides reproducible training and OOD evaluation splits for supervised fine-tuning models that translate natural-language telecom/network-slicing intents into structured JSON configuration objects.
This dataset is intended for research. It is not a production-certified telecom configuration generator and should not be used to deploy network… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/TMF921-intent-to-config-research-sota.NR-short-32k-16-thoughts-1k-8-thoughtsSubsample of JackHsieh/NR-short-32k-16-thoughts-4k-128-thoughts, with a smaller test set.
NR-short-32k-16-thoughts-4k-128-thoughtsNR-short-32k-16-thoughts-4k-8-thoughtsSubsample of JackHsieh/NR-short-32k-16-thoughts-4k-128-thoughts, with a smaller test set.
strawberry_50_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/nrburns/strawberry_50_3.NeversNet5G_Vehicular_5G_NR_Dataset
NeversNet5G
NeversNet5G is a city-scale 5G NR vehicular network dataset generated through
SUMO and Simu5G/OMNeT++ co-simulation over the real urban road network of
Nevers, France. The release contains processed per-UE event-level CSV files,
not the raw OMNeT++ .vec/.sca outputs.
Release Contents
data/: 716 per-UE metric CSV files, grouped by scenario part.
configs/: OMNeT++/Simu5G configuration files used for the scenario. The
released omnetpp.ini is trimmed to… See the full description on the dataset page: https://huggingface.co/datasets/Askedrin/NeversNet5G_Vehicular_5G_NR_Dataset.eval_act_merged_Nrc0_Nh2g_50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "arxr5_bimanual",
"total_episodes": 1,
"total_frames": 511,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yutian929/eval_act_merged_Nrc0_Nh2g_50.test_nredThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 30,
"total_frames": 4907,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shuohsuan/test_nred.eval_act_merged_Nrc0_Nh2g_50_failedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "arxr5_bimanual",
"total_episodes": 1,
"total_frames": 478,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yutian929/eval_act_merged_Nrc0_Nh2g_50_failed.Tox21_NRERNR-short-1k-1-thought-1k-1-thoughtSubsample of JackHsieh/NR-short-32k-16-thoughts-4k-128-thoughts, with a much smaller test set.
NR-short-128-128nrps_modules_asdb4.0The dataset was extracted from antismash-db 4.0 postgresql dump, the corresponding description can be found here:
https://antismash-db.secondarymetabolites.org/
If you want to use it, please refer to the original licensing terms and properly cite the authors.
Each line in .csv file corresponds to NRPS module, both with the monomer produced by it.
Each module has a specific sequence of domains. Module possible structure is described in the literature.
To clean the data, I referred to wiki… See the full description on the dataset page: https://huggingface.co/datasets/latticetower/nrps_modules_asdb4.0.
