datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MoleculeNet_HIV
MoleculeNet HIV
HIV dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict ability of molecules to inhibit HIV replication.
Characteristic
Description
Tasks
1
Task type
classification
Total samples
41127
Recommended split
scaffold
Recommended metric
AUROCWarning: in newer RDKit vesions, 7 molecules from the original dataset are not read correctly due to disallowed
hypervalent states of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_HIV.pii-bench
PII-Bench (ru)
Span-level benchmark for evaluating personal-data (PII) detection in Russian text.
Annotations use explicit character offsets (start, end, type) rather than
IO/BIO/BILOU token tags. This makes the benchmark agnostic to tokenization and
lets you evaluate a complete pipeline — ML model, regular expressions,
post-processing, or a hybrid such as Presidio —
instead of only the model in isolation.
Released alongside GLiNER Guard, a unified safety + PII encoder family:… See the full description on the dataset page: https://huggingface.co/datasets/hivetrace/pii-bench.hiv
Dataset Details
Dataset Description
The HIV dataset was introduced by the Drug Therapeutics Program (DTP)
AIDS Antiviral Screen, which tested the ability to inhibit HIV replication for
over 40,000 compounds.
Curated by:
License: CC BY 4.0
Dataset Sources
data source
corresponding publication
data source
Citation
BibTeX:
@article{Wu2018,
doi = {10.1039/c7sc02664a},
url = {https://doi.org/10.1039/c7sc02664a},
year = {2018},
publisher = {Royal… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/hiv.general-layerC-200klanguage:
en
license: apache-2.0
size_categories:
100K<n<1M
pretty_name: General LayerC — training-ready (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:apache-2.0
domain:general
reasoning
synthetic
sft
general-knowledge
science
creative-writing
glaive
deepseek-r1-distill
distillation
layer-c
layer-training-ready
General LayerC (training-ready) — 200k
Built from glaiveai/reasoning-v1-20m (spread-sampled across… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerC-200k.medical-layerC-200kmath-layerC-200Kcode-layerB-200klanguage:
en
license: cc-by-4.0
size_categories:
100K<n<1M
pretty_name: Code LayerB — teacher-supervision (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:cc-by-4.0
domain:code
competitive-programming
reasoning
synthetic
sft
codegen
python
algorithms
data-structures
debugging
nvidia
opencodereasoning
deepseek-r1
codeforces
atcoder
codechef
leetcode
layer-b
layer-teacher-supervision
teacher-supervision
Code LayerB… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerB-200k.code-layerA-200klanguage:
en
license: cc-by-4.0
size_categories:
100K<n<1M
pretty_name: Code LayerA — canonical (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:cc-by-4.0
domain:code
competitive-programming
reasoning
synthetic
sft
codegen
python
algorithms
data-structures
debugging
nvidia
opencodereasoning
deepseek-r1
codeforces
atcoder
codechef
leetcode
layer-a
layer-canonical
Code LayerA (canonical) — 200k
Built from… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerA-200k.HiveA Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
Kai Li*, Jintao Cheng*, Chang Zeng, Zijun Yan, Helin Wang, Zixiong Su, Bo Zheng, Xiaolin Hu
Tsinghua University, Shanda AI, Johns Hopkins University
*Equal contribution
Completed during Kai Li's internship at Shanda AI.
📜 Arxiv 2026 | 💻 Code | 🎶 Demo
Usage
from datasets import load_dataset
# Load full dataset
dataset = load_dataset("ShandaAI/Hive")
# Load… See the full description on the dataset page: https://huggingface.co/datasets/AlayaLab/Hive.prompt-2-prompt-injection-v2-dataset-ruПереведённый с помощью Gemini 2.5 flash и Gemini 2.0 flash вариант датасета r1char9/prompt-2-prompt-injection-v2-dataset
africa-who-treatment-success-rate-hiv-positive-tb-cases
Africa — WHO GHO: Treatment success rate: HIV-positive TB cases | Africa (World Health Organization)
Size category: n<1K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-treatment-success-rate-hiv-positive-tb-cases.LOGIC-701
LOGIC-701 Benchmark
This is a synthetic and filtered dataset for benchmarking large language models (LLMs). It consists of 701 medium and hard logic puzzles with solutions on 10 distinct topics.
A feature of the dataset is that it tests exclusively logical/reasoning abilities, offering only 5 answer options. There are no or very few tasks in the dataset that require external knowledge about events, people, facts, etc.
Languages
This benchmark is also part of an… See the full description on the dataset page: https://huggingface.co/datasets/hivaze/LOGIC-701.Brain-HIVE_Visual_Embeddingshiv-multimodalHiveA Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
Kai Li*, Jintao Cheng*, Chang Zeng, Zijun Yan, Helin Wang, Zixiong Su, Bo Zheng, Xiaolin Hu
Tsinghua University, Shanda AI, Johns Hopkins University
*Equal contribution
📜 Arxiv 2026 | 💻 Code | 🎶 Demo
Usage
from datasets import load_dataset
# Load full dataset
dataset = load_dataset("ShandaAI/Hive")
# Load specific split
train_data = load_dataset("ShandaAI/Hive"… See the full description on the dataset page: https://huggingface.co/datasets/JusperLee/Hive.asia-who-treatment-success-rate-hiv-positive-tb-cases
Treatment success rate: HIV-positive TB cases | Asia (WHO GHO)
🌏 551 observations · 45 Asia countries · 1999–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 551 observations of Treatment success rate: HIV-positive TB cases data across 45 Asia countries, spanning 1999–2023, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License: cc-by-4.0
Topic:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-treatment-success-rate-hiv-positive-tb-cases.strongrejectPlusPlus
🌍 strongREJECT++ Dataset
Welcome to the strongREJECT++ dataset!
This dataset is a collection of translations from the original strongREJECT dataset.
"It's cutting-edge benchmark for evaluating jailbreaks in Large Language Models (LLMs)"
Available Languages
You can find translations provided by native speakers in the following languages:
🇺🇸 English
🇷🇺 Russian
🇺🇦 Ukrainian
🇧🇾 Belarusian
🇺🇿 Uzbek
Hive-ALLA Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
Kai Li*, Jintao Cheng*, Chang Zeng, Zijun Yan, Helin Wang, Zixiong Su, Bo Zheng, Xiaolin Hu
Tsinghua University, Shanda AI, Johns Hopkins University
*Equal contribution
📜 Arxiv 2026 | 🎶 Demo | 🤗 Metadata | 🤗 Hive-ALL Audio | 🤗 Space
💥 News
[2026-05-21] Hive-ALL is now also available on ModelScope for users in China who prefer faster downloads via the… See the full description on the dataset page: https://huggingface.co/datasets/JusperLee/Hive-ALL.asia-who-tested-tb-patients-hiv-positive
Tested TB patients HIV-positive (%) | Asia (WHO GHO)
🌏 735 observations · 44 Asia countries · 2003–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 735 observations of Tested TB patients HIV-positive (%) data across 44 Asia countries, spanning 2003–2024, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License: cc-by-4.0
Topic: Tested TB patients… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-tested-tb-patients-hiv-positive.general-layerA-200klanguage:
en
license: apache-2.0
size_categories:
100K<n<1M
pretty_name: General LayerA — canonical (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:apache-2.0
domain:general
reasoning
synthetic
sft
general-knowledge
science
creative-writing
glaive
deepseek-r1-distill
distillation
layer-a
layer-canonical
General LayerA (canonical) — 200k
Built from glaiveai/reasoning-v1-20m (spread-sampled across all 709 shards).… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerA-200k.hiv
Dataset Card for "hiv"
More Information needed
hivex-leaderboard-datageneral-layerB-200klanguage:
en
license: apache-2.0
size_categories:
100K<n<1M
pretty_name: General LayerB — teacher-supervision (200,000 samples)
tags:
task_categories:text-generation
task_categories:text2text-generation
language:en
license:apache-2.0
domain:general
reasoning
synthetic
sft
general-knowledge
science
creative-writing
glaive
deepseek-r1-distill
distillation
layer-b
layer-teacher-supervision
teacher-supervision
General LayerB (teacher-supervision) — 200k
Built from… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/general-layerB-200k.cs217-rlhf-dataset
CS217 Fixed HH-RLHF Dataset
Fixed subset of Anthropic's HH-RLHF dataset for reproducible RLHF experiments
Created for Stanford CS217: Hardware Accelerators for Machine Learning - Final Project
🔗 GitHub Repository: CS217-Final-Project
Dataset Description
This is a fixed subset of the Anthropic/hh-rlhf dataset, created to ensure reproducible experiments across all runs. The dataset contains human preference pairs for training reward models and RLHF (Reinforcement Learning… See the full description on the dataset page: https://huggingface.co/datasets/hivamoh/cs217-rlhf-dataset.code-layerB-final
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerB-final.HIVBench
HIVBench
HIVBench is a clinical reasoning benchmark designed to evaluate Large Language Models on the management of advanced HIV disease. It consists of 269 expert-level multiple-choice questions rigorously synthesized from official clinical protocols to address a critical gap in medical AI evaluation.
Background & Motivation
HIV remains one of the most significant global health challenges. Despite advancements, the burden is disproportionately concentrated in… See the full description on the dataset page: https://huggingface.co/datasets/Larxel/HIVBench.code-layerC-final
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hivemind-research/code-layerC-final.asia-who-hiv-positive-tb-patients-on-art
HIV-positive TB patients on ART (antiretroviral therapy) (%) | Asia (WHO GHO)
🌏 609 observations · 42 Asia countries · 2003–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 609 observations of HIV-positive TB patients on ART (antiretroviral therapy) (%) data across 42 Asia countries, spanning 2003–2024, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-hiv-positive-tb-patients-on-art.asia-who-tb-patients-with-known-hiv-status
TB patients with known HIV status (%) | Asia (WHO GHO)
🌏 834 observations · 47 Asia countries · 2003–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 834 observations of TB patients with known HIV status (%) data across 47 Asia countries, spanning 2003–2024, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License: cc-by-4.0
Topic: TB patients with… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-tb-patients-with-known-hiv-status.MoleculeNet_HIV
Mirrored by Aurigene AI
Discovery stage: Hit generation
Compounds screened for inhibition of HIV replication. The classic large-scale virtual screening benchmark.
Rows: 41,127 (hiv.csv 41,127)
Pairs with Aurigene-AI/MoLFormer-XL-both-10pct from our model catalogue.
Upstream: scikit-fingerprints/MoleculeNet_HIV - all credit to the original authors and to the researchers who produced the underlying data; the dataset card and licence below are theirs.
Explore the rest of the… See the full description on the dataset page: https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_HIV.
