datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASVspoof2021_DF
ASVspoof 2021 DF
Benchmark-ready packaging of the DeepFake (DF) evaluation partition from ASVspoof 2021 for speech anti-spoofing and synthetic / deepfake voice detection.
Overview
This dataset contains the DF evaluation subset of the ASVspoof 2021 challenge. The task is binary classification: bonafide (genuine human speech) vs. spoof (synthetic, converted, or otherwise manipulated speech). The original dataset is available at… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarainala152/ASVspoof2021_DF.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for Indian languages. All annotations are human-corrected with time-aligned, speaker-attributed transcriptions. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/indic-diarbench.big_patent
Dataset Card for Big Patent
Dataset Summary
BIGPATENT, consisting of 1.3 million records of U.S. patent documents along with human written abstractive summaries.
Each US patent application is filed under a Cooperative Patent Classification (CPC) code.
There are nine such classification categories:
a: Human Necessities
b: Performing Operations; Transporting
c: Chemistry; Metallurgy
d: Textiles; Paper
e: Fixed Constructions
f: Mechanical Engineering; Lightning; Heating;… See the full description on the dataset page: https://huggingface.co/datasets/manoj8890/big_patent.ca-sco-properties
CA SCO Unclaimed Property
California State Controller's Office unclaimed property records.
Bucketed into alphabetical splits by owner last name first letter
so every split stays under HF's 5 GB filter-index limit.
Updated weekly via GitHub Actions.
Manosaba_Benchmark_Reorderedfootball-players
Dataset Labels
['football', 'player']
Number of Images
{'valid': 87, 'train': 119}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("manot/football-players", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/konstantin-sargsyan-wucpb/football-players-2l81z/dataset/1
Citation
@misc{… See the full description on the dataset page: https://huggingface.co/datasets/manot/football-players.pothole-segmentation
Dataset Labels
['potholes', 'object', 'pothole', 'potholes']
Number of Images
{'valid': 157, 'test': 80, 'train': 582}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("manot/pothole-segmentation", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/abdulmohsen-fahad-f7pdw/road-damage-xvt2d/dataset/3
Citation… See the full description on the dataset page: https://huggingface.co/datasets/manot/pothole-segmentation.MedHallu
Dataset Card for MedHallu
MedHallu is a comprehensive benchmark dataset designed to evaluate the ability of large language models to detect hallucinations in medical question-answering tasks.
Dataset Details
Dataset Description
MedHallu is intended to assess the reliability of large language models in a critical domain—medical question-answering—by measuring their capacity to detect hallucinated outputs. The dataset includes two distinct splits:… See the full description on the dataset page: https://huggingface.co/datasets/Manoharareddy/MedHallu.gigahands-vitra-mano
GigaHands → VITRA Stage-1, official-MANO annotations
VITRA Stage-1 hand annotations for GigaHands, with all joint positions taken from GigaHands'
official MANO fit instead of mixing in triangulated keypoints. Annotations only — no videos
(get those from GigaHands; the mapping is described in §5).
episodes
13,247 (train 11,904 / test 1,343)
frames
3,395,733
camera
brics-odroid-001_cam0 (static rig; one constant extrinsic per scene)
source
GigaHands params/ +… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/gigahands-vitra-mano.Manosaba_Benchmarkclaude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/manojdahal191gom/claude-opus-4.6-4.7-reasoning-8.7k.github-pullrequestsnoticias_ptbrpothole-segmentation2
Dataset Labels
['pothole']
Number of Images
{'valid': 133, 'test': 66, 'train': 466}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("manot/pothole-segmentation2", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/gurgen-hovsepyan-mbrnv/pothole-detection-gilij/dataset/2
Citation
@misc{… See the full description on the dataset page: https://huggingface.co/datasets/manot/pothole-segmentation2.manus-mano-poses
Manus MANO Poses
This dataset contains a right-hand Manus glove recording converted into the 21-landmark hand-pose convention used by the orca_teleop pipeline and its retargeters.
Contents
Split: train
Frames: 2278
Duration: 37.95 s
Sampling rate: 60.00 Hz
Handedness: right
Each row contains frame, elapsed timestamps, handedness, and keypoints, a (21, 3) float32 array of wrist-relative 3D landmarks in meters.
Landmark Order
The keypoints array follows the… See the full description on the dataset page: https://huggingface.co/datasets/fracapuano/manus-mano-poses.control_image2diamond-price-predictor-logs2
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/manojdec25/diamond-price-predictor-logs2.philosloppy_encyclopediaanveshana
Dataset Card for Anveshana
Dataset Details
Dataset Description
we embarked on a comprehensive benchmarking study to explore and evaluate current state-of-the-art models for Cross-Lingual Information Retrieval (CLIR) from English to Sanskrit. Our primary objective is to assess the effectiveness of these models in accurately retrieving Sanskrit documents based on English queries. To achieve this, we meticulously assembled a robust dataset, focusing on the… See the full description on the dataset page: https://huggingface.co/datasets/manojbalaji1/anveshana.ego4d-hand-mano
Dataset summary
Per-frame 3D hand annotations and text captions for 1,302,538 egocentric video clips
drawn from 3,320 Ego4D videos. Every clip is 121 frames at 30 fps (4.03 s) at a
540-pixel short side.
This is the annotation release accompanying Controllable Egocentric Video Generation via
Occlusion-Aware Sparse 3D Hand Joints (ECCV 2026). Each clip carries, for both hands and every
frame: 21 3D joints, MANO pose parameters, root rotation and translation, 2D wrist position
and… See the full description on the dataset page: https://huggingface.co/datasets/bochen123/ego4d-hand-mano.replicated_emotions
Dataset Summary
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper.
This dataset is a processed form of "dair-ai/emotion" dataset. [https://huggingface.co/datasets/dair-ai/emotion]
In this one, I have replicated/duplicated the samples for minority classes so that all the emotion classes have [approximate] equal sample count.
dataset_info:
features:
name:… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarvohra/replicated_emotions.ManoloPueblo__LLM_MERGE_CC2-details
Dataset Card for Evaluation run of ManoloPueblo/LLM_MERGE_CC2
Dataset automatically created during the evaluation run of model ManoloPueblo/LLM_MERGE_CC2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ManoloPueblo__LLM_MERGE_CC2-details.plantvillageManoloPueblo__LLM_MERGE_CC3-details
Dataset Card for Evaluation run of ManoloPueblo/LLM_MERGE_CC3
Dataset automatically created during the evaluation run of model ManoloPueblo/LLM_MERGE_CC3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ManoloPueblo__LLM_MERGE_CC3-details.ManoloPueblo__ContentCuisine_1-7B-slerp-details
Dataset Card for Evaluation run of ManoloPueblo/ContentCuisine_1-7B-slerp
Dataset automatically created during the evaluation run of model ManoloPueblo/ContentCuisine_1-7B-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ManoloPueblo__ContentCuisine_1-7B-slerp-details.amplified_emotions
Dataset Summary
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper.
This dataset is a processed form of "dair-ai/emotion" dataset. [https://huggingface.co/datasets/dair-ai/emotion]
In this one, I have amplified the samples for minority classes so that all the emotion classes have [approximate] equal sample count.
There is another dataset with duplicate… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarvohra/amplified_emotions.Sanskrit-to-English-Vocabulary-v1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: Manoj, Nandish, Mayank, Abhiram
Language(s) (NLP): Sanskrit, English
License: MIT License
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/Manoj2702/Sanskrit-to-English-Vocabulary-v1.Manosaba_Scriptsblender_duplicates
Dataset Card for Dataset Name
Contains reduced description of issues reported at https://projects.blender.org/blender/blender/issues and points to duplicate issues in order to categorize similarity.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Each report has been shortened by removing frequently repeated texts such as System Information, Blender Version… See the full description on the dataset page: https://huggingface.co/datasets/mano-wii/blender_duplicates.salesdata
