datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.MoleculeNet_Lipophilicity
MoleculeNet Lipophilicity
Lipophilicity dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict octanol/water distribution coefficient (logD) at pH 7.4. Targets are already log transformed, and are a unitless ratio.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
4200
Recommended split
scaffold
Recommended metric
RMSE
References
[1]
Wu, Zhenqin, et al.… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_Lipophilicity.Lipoprotein-a-Binder-Designs
Lipoprotein(a) Binder Designs — apo(a) KIV-8 / 8TCE
Why this target matters. Lp(a) is the most common genetic cardiovascular risk factor in the world, affecting roughly a fifth of the population, and the only major one with no approved therapy.
100 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the lysine-binding site of kringle IV type 8 (KIV-8) of apolipoprotein(a), using the 1.07 Å co-crystal 8TCE — Lp(a) KIV-8 in complex with… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/Lipoprotein-a-Binder-Designs.MoleculeNet_Lipophilicity
Mirrored by Aurigene AI
Discovery stage: Lead optimization
Octanol/water distribution coefficient (logD at pH 7.4). Regression, a core ADMET endpoint.
Rows: 4,200 (lipophilicity.csv 4,200)
Pairs with Aurigene-AI/MoLFormer-XL-both-10pct from our model catalogue.
Upstream: scikit-fingerprints/MoleculeNet_Lipophilicity - all credit to the original authors and to the researchers who produced the underlying data; the dataset card and licence below are theirs.
Explore the rest of the… See the full description on the dataset page: https://huggingface.co/datasets/Aurigene-AI/MoleculeNet_Lipophilicity.lipo
Dataset Card for lipo
Dataset Summary
lipo is a dataset included in MoleculeNet. It measures the experimental results of octanol/water distribution coefficient(logD at pH 7.4)
Dataset Structure
Data Fields
Each split contains
smiles: the SMILES representation of a molecule
selfies: the SELFIES representation of a molecule
target: octanol/water distribution coefficient(logD at pH 7.4)
Data Splits
The dataset is split into an 80/10/10… See the full description on the dataset page: https://huggingface.co/datasets/zpn/lipo.MassSpecGym_LipidWildVid-LIP
WildVid-LIP: In-The-Wild Temporal Anchors for Visual Speech Recognition
WildVid-LIP is a large-scale, open-source dataset mapping over 100,000 curated temporal segments from unconstrained, real-world YouTube videos. It provides precise timestamp anchors optimized for training Visual Speech Recognition (VSR / Lip-Reading), audio-visual synchronization, and multimodal self-supervised models.
Instead of distributing heavy, monolithic video files—which introduces platform friction… See the full description on the dataset page: https://huggingface.co/datasets/Rizul2159/WildVid-LIP.MoleculeNet_Lipophilicity
MoleculeNet Lipophilicity
Lipophilicity dataset, part of MoleculeNet [1] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict octanol/water distribution coefficient (logD) at pH 7.4. Targets are already log transformed, and are a unitless ratio.
Characteristic
Description
Tasks
1
Task type
regression
Total samples
4200
Recommended split
scaffold
Recommended metric
RMSE
References
[1]
Wu, Zhenqin… See the full description on the dataset page: https://huggingface.co/datasets/mdrahmanshakil/MoleculeNet_Lipophilicity.LiProSlipohrevaldatasetThe dataset can be used to evaluate a intent based classifier with the folowing intents
['training_request',
'performance_review',
'access_request',
'relocation_request',
'safety_incident_report',
'time_off_report',
'benefits_enrollment',
'harassment_report',
'goal_setting',
'it_issue_report']
There are 5 instances for each intent
Lipophilicity-Predictionlipinski_prediction
Lipinski Prediction Dataset
This dataset is part of the Deep Principle Bench collection.
Files
lipinski_prediction.csv: Main dataset file
Usage
import pandas as pd
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("yhqu/lipinski_prediction")
# Or load directly as pandas DataFrame
df = pd.read_csv("hf://datasets/yhqu/lipinski_prediction/lipinski_prediction.csv")
Citation
Please cite this work if you use this dataset… See the full description on the dataset page: https://huggingface.co/datasets/yhqu/lipinski_prediction.Lipo
