datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MoleculeNet_HIV
MoleculeNet HIV
HIV dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict ability of molecules to inhibit HIV replication.
Characteristic
Description
Tasks
1
Task type
classification
Total samples
41127
Recommended split
scaffold
Recommended metric
AUROCWarning: in newer RDKit vesions, 7 molecules from the original dataset are not read correctly due to disallowed
hypervalent states of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_HIV.wan22-physics-videos
Wan2.2 Physics Video Generation Dataset
A dataset of AI-generated physics simulation videos with intermediate denoising latent tensors, created using the Wan2.2 T2V-A14B text-to-video model.
Dataset Summary
480 videos (.mp4) generated from 60 unique physics prompts (8 random seeds each)
Intermediate latent tensors (.pt) saved at every 2 denoising steps (10 snapshots per video)
Physics categories: ball bouncing, pendulum motion, objects sliding on inclined planes… See the full description on the dataset page: https://huggingface.co/datasets/hivamoh/wan22-physics-videos.srt-hivemind
SRT hivemind artifacts
Every number in Where the Hivemind Comes From has a file here. This repository is
the evidence, not a model: sampled continuations, hidden states, fitted-map scores,
execution outcomes and run logs.
Paper: paper_hivemind.md.
Code: github.com/space-bacon/SRT.
Reproduction principle. If a claim is in the paper and you cannot rebuild it from
the files here plus the scripts in the GitHub repo, that is a bug and we want to hear
about it. Four claims were… See the full description on the dataset page: https://huggingface.co/datasets/RiverRider/srt-hivemind.HIVAU-70k_XD-Violence
다운받아서 압축해제 후 사용하세요
현재위치에 다운
hf download backseollgi/HIVAU-70k_XD-Violence --repo-type dataset --local-dir .
cat xd-violence_clips.tar.part_* | tar -xvf -
code-layerC-200klanguage:
en
license: cc-by-4.0
size_categories:
100K<n<1M
pretty_name: Code LayerC — training-ready (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:cc-by-4.0
domain:code
competitive-programming
reasoning
synthetic
sft
codegen
python
algorithms
data-structures
debugging
nvidia
opencodereasoning
deepseek-r1
codeforces
atcoder
codechef
leetcode
layer-c
layer-training-ready
Code LayerC (training-ready) — 200k
Built… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerC-200k.pii-bench
PII-Bench (ru)
Span-level benchmark for evaluating personal-data (PII) detection in Russian text.
Annotations use explicit character offsets (start, end, type) rather than
IO/BIO/BILOU token tags. This makes the benchmark agnostic to tokenization and
lets you evaluate a complete pipeline — ML model, regular expressions,
post-processing, or a hybrid such as Presidio —
instead of only the model in isolation.
Released alongside GLiNER Guard, a unified safety + PII encoder family:… See the full description on the dataset page: https://huggingface.co/datasets/hivetrace/pii-bench.hiv
Dataset Details
Dataset Description
The HIV dataset was introduced by the Drug Therapeutics Program (DTP)
AIDS Antiviral Screen, which tested the ability to inhibit HIV replication for
over 40,000 compounds.
Curated by:
License: CC BY 4.0
Dataset Sources
data source
corresponding publication
data source
Citation
BibTeX:
@article{Wu2018,
doi = {10.1039/c7sc02664a},
url = {https://doi.org/10.1039/c7sc02664a},
year = {2018},
publisher = {Royal… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/hiv.pick_toys_21_jan_1cam_bg_replace_06This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3_real",
"total_episodes": 261,
"total_frames": 23992,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:261"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/pick_toys_21_jan_1cam_bg_replace_06.general-layerC-200klanguage:
en
license: apache-2.0
size_categories:
100K<n<1M
pretty_name: General LayerC — training-ready (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:apache-2.0
domain:general
reasoning
synthetic
sft
general-knowledge
science
creative-writing
glaive
deepseek-r1-distill
distillation
layer-c
layer-training-ready
General LayerC (training-ready) — 200k
Built from glaiveai/reasoning-v1-20m (spread-sampled across… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerC-200k.pick_toy_light_night_1cam_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3_real",
"total_episodes": 163,
"total_frames": 15607,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:163"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/pick_toy_light_night_1cam_2.pick_toys_21_jan_1cam_brightcontrast_smallThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3_real",
"total_episodes": 130,
"total_frames": 12176,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:130"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/pick_toys_21_jan_1cam_brightcontrast_small.hive-partitioned-data-demomedical-layerC-200kHivemath-layerC-200Kpick_toy_1cam_idThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3_real",
"total_episodes": 74,
"total_frames": 7298,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:74"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/pick_toy_1cam_id.pick_toy_qualiaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3_real",
"total_episodes": 163,
"total_frames": 15607,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:163"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/pick_toy_qualia.retail_restock_11_feb_no_verbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3",
"total_episodes": 124,
"total_frames": 13344,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:124"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/retail_restock_11_feb_no_verb.code-layerB-200klanguage:
en
license: cc-by-4.0
size_categories:
100K<n<1M
pretty_name: Code LayerB — teacher-supervision (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:cc-by-4.0
domain:code
competitive-programming
reasoning
synthetic
sft
codegen
python
algorithms
data-structures
debugging
nvidia
opencodereasoning
deepseek-r1
codeforces
atcoder
codechef
leetcode
layer-b
layer-teacher-supervision
teacher-supervision
Code LayerB… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerB-200k.pick_toy_full_1cam_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3_real",
"total_episodes": 263,
"total_frames": 30183,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:263"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/pick_toy_full_1cam_2.code-layerA-200klanguage:
en
license: cc-by-4.0
size_categories:
100K<n<1M
pretty_name: Code LayerA — canonical (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:cc-by-4.0
domain:code
competitive-programming
reasoning
synthetic
sft
codegen
python
algorithms
data-structures
debugging
nvidia
opencodereasoning
deepseek-r1
codeforces
atcoder
codechef
leetcode
layer-a
layer-canonical
Code LayerA (canonical) — 200k
Built from… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerA-200k.HiveA Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
Kai Li*, Jintao Cheng*, Chang Zeng, Zijun Yan, Helin Wang, Zixiong Su, Bo Zheng, Xiaolin Hu
Tsinghua University, Shanda AI, Johns Hopkins University
*Equal contribution
Completed during Kai Li's internship at Shanda AI.
📜 Arxiv 2026 | 💻 Code | 🎶 Demo
Usage
from datasets import load_dataset
# Load full dataset
dataset = load_dataset("ShandaAI/Hive")
# Load… See the full description on the dataset page: https://huggingface.co/datasets/AlayaLab/Hive.prompt-2-prompt-injection-v2-dataset-ruПереведённый с помощью Gemini 2.5 flash и Gemini 2.0 flash вариант датасета r1char9/prompt-2-prompt-injection-v2-dataset
pick_toy_light_id_1cam_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3_real",
"total_episodes": 67,
"total_frames": 6122,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:67"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/pick_toy_light_id_1cam_2.africa-who-treatment-success-rate-hiv-positive-tb-cases
Africa — WHO GHO: Treatment success rate: HIV-positive TB cases | Africa (World Health Organization)
Size category: n<1K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-treatment-success-rate-hiv-positive-tb-cases.LOGIC-701
LOGIC-701 Benchmark
This is a synthetic and filtered dataset for benchmarking large language models (LLMs). It consists of 701 medium and hard logic puzzles with solutions on 10 distinct topics.
A feature of the dataset is that it tests exclusively logical/reasoning abilities, offering only 5 answer options. There are no or very few tasks in the dataset that require external knowledge about events, people, facts, etc.
Languages
This benchmark is also part of an… See the full description on the dataset page: https://huggingface.co/datasets/hivaze/LOGIC-701.Brain-HIVE_Visual_Embeddingshiv-multimodalHIVAU-70k_UCF-crime
다운받아서 압축해제 후 사용하세요
현재위치에 다운
hf download backseollgi/HIVAU-70k_UCF-crime --repo-type dataset --local-dir .
cat ucf-crime.tar.part_* | tar -xvf -
HiveA Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
Kai Li*, Jintao Cheng*, Chang Zeng, Zijun Yan, Helin Wang, Zixiong Su, Bo Zheng, Xiaolin Hu
Tsinghua University, Shanda AI, Johns Hopkins University
*Equal contribution
📜 Arxiv 2026 | 💻 Code | 🎶 Demo
Usage
from datasets import load_dataset
# Load full dataset
dataset = load_dataset("ShandaAI/Hive")
# Load specific split
train_data = load_dataset("ShandaAI/Hive"… See the full description on the dataset page: https://huggingface.co/datasets/JusperLee/Hive.
