datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LipFD
[NeruIPS 2024] Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-syncing DeepFakes
Abstract. In recent years, DeepFake technology has achieved unprecedented success in high-quality video synthesis, but these methods also pose potential and severe security threats to humanity. DeepFake can be bifurcated into entertainment applications like face swapping and illicit uses such as lip-syncing fraud. However, lip-forgery videos, which neither… See the full description on the dataset page: https://huggingface.co/datasets/huahua123313/LipFD.lipreadingchinese-lips-speech-slide-probe
Chinese-LiPS Speech + Slide Probe
A self-contained probe set for testing whether visual slide context helps
simultaneous speech translation — with the input as audio, not transcripts.
Why audio matters: feeding a transcript to a text LLM deletes the acoustic
ambiguity (homophones, polysemy) that slide context is meant to resolve; the
transcript already commits to one reading. Any honest test of "does vision help
streaming ST" must consume speech.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.lip_detectionlipylipika-eval
Lipika eval — Indic font recognition benchmark
The frozen validation set behind loopdesk-ai/lipika
(Indic font recognizer): 6,876 synthetic text crops covering 553 freely-licensed font
families across 13 scripts (Devanagari, Bengali, Gujarati, Gurmukhi, Kannada, Malayalam,
Meetei Mayek, Odia, Ol Chiki, Perso-Arabic, Tamil, Telugu, Latin).
This is the set reported as "synthetic val" in the model card (Lipika v2.4 scores 0.849
family top-1 / 0.977 top-5 / 0.991 script). Use it to… See the full description on the dataset page: https://huggingface.co/datasets/loopdesk-ai/lipika-eval.vton-temp-uploadsStudentHabitsvsAcademicPerformance
Student Habits vs Academic Performance
Dataset & Analysis by Lipaz Levy
This project explores how students’ daily habits — including study hours, sleep, social media use, lifestyle factors, mental health, exercise, diet quality, and more — influence their academic performance.
We used a dataset of 1001 students, each with lifestyle habits, well-being indicators, and exam scores.The goal was to answer:
How do students’ daily habits affect their exam performance?
Data… See the full description on the dataset page: https://huggingface.co/datasets/lipaz/StudentHabitsvsAcademicPerformance.diamondsProject Walkthrough Video Click here to watch the video
0. Project Overview & Initial Workflow
0.1 Project Goal
The main objective of this project was to analyze the diamonds dataset and understand:
What drives the price of a diamond,
How physical attributes (carat, x/y/z dimensions) behave,
How categorical qualities (cut, color, clarity) influence pricing,
And which features are the most important for predictive modeling.
This includes EDA, feature engineering… See the full description on the dataset page: https://huggingface.co/datasets/lipaz/diamonds.lipophilicity-multimodalLiposome-RBC_Scoping_Review_Field_Overview
Liposome-RBC Interactions Field Overview Dataset
Description
This dataset contains structured information extracted from 487 title-abstracts that examine the interactions between liposomes and red blood cells (RBCs). The data was gathered through a systematic PRISMA-compatible scoping review process enhanced with AI assistance (Claude 3.7 Sonnet) and structured with JSON schema for reproducible analysis.
Dataset Content
The dataset captures key information from… See the full description on the dataset page: https://huggingface.co/datasets/UtopiansRareTruth/Liposome-RBC_Scoping_Review_Field_Overview.VeGame-AtariLHRS_Benchlipid_droplets
Dataset Card for "lipid_droplets"
More Information needed
lipid-dropletsLIPimagesLiProSlipslipreading-wordslevihello levi
yoyo
yes yes
lipid-droplets-v3lipid-droplets-v4imnet1k_lipstick_lip_rougelipid-droplets-v2lipsFinalmiracl-lips-datasetlipstick-image-dataset
Dataset Card for keerthikoganti/lipstick-image-dataset
Dataset Details
Dataset Description
This dataset consists of images labeled as lipstick (1) or no_lipstick (0). It was created as part of a classroom exercise in supervised learning and data augmentation, with the goal of practicing binary image classification and experimenting with dataset curation, preprocessing, and augmentation.
Curated by: Fall 2025 24-679 course at Carnegie Mellon
Shared by :… See the full description on the dataset page: https://huggingface.co/datasets/keerthikoganti/lipstick-image-dataset.lipid-droplets-v5Odia-lipi-ocr-data
Dataset Card for Odia Lipi OCR Dataset
Dataset Description
This dataset is part of the Odia Lipi Project, an open and collaborative initiative led in collaboration with AHRC IIT Bhubaneswar, aimed at building high-quality OCR resources for the Odia language. It contains scanned page images paired with human-validated Odia text, intended to support OCR model training, post-OCR correction, NLP research, and the digital preservation of Odia literature and documents.
The… See the full description on the dataset page: https://huggingface.co/datasets/OdiaGenAIOCR/Odia-lipi-ocr-data.
