datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PointCloudCorruptionCoSyn-point
CoSyn-point
CoSyn-point is a collection of diverse computer-generated images that are annotated with queries and answer points.
It can be used to train models to return points in the image in response to a user query.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
The code used to generate this data is open source.
Synthetic question-answer data is also available in a seperate repo.
Quick links:
📃 CoSyn… See the full description on the dataset page: https://huggingface.co/datasets/allenai/CoSyn-point.pixmo-points
PixMo-Points
PixMo-Points is a dataset of images paired with referring expressions and points marking the locations the
referring expression refers to in the image. It was collected using human annotators and contains a diverse
range of points and expressions, with many high-frequency (10+) expressions.
PixMo-Points is a part of the PixMo dataset collection and was used to
provide the pointing capabilities of the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with Videos… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-points.pointerbench
Pointerbench
Pointerbench is a small GUI grounding benchmark suite for computer-use models.
Each example has one screenshot, one instruction, target geometry in absolute
pixels, and a binary evaluation rule.
Links:
GitHub: https://github.com/warmwindOS/pointerbench
Blog post: https://about.warmwind.com/pointer-bench/
Add your model to the official benchmark leaderboard: https://warmwind.com/contact
The suite has three subsets:
Subset
Examples
What it tests… See the full description on the dataset page: https://huggingface.co/datasets/WarmwindOS/pointerbench.puregen-poisonspointer-retrievalrelease-910k-new
SOMA UMR Release v260717 T2M 910k (realigned)
Realigned from release_v260710_t2m_910k per
code/soma-motion-tokenizer/doc/handoff_hml_split_realign.md.
split_seed: 20260717
HumanML3D: official tags.split
bones-seed: Kimodo test pinned; val+extra-test carved from Kimodo-train
others: per-subset 80:5:15 by source group
Old release_v260710_t2m_910k is preserved untouched.
common_crawl_pointer_indicesllm-graph-poisoning-data
Generation-Time Poisoning of LLM-Generated Social Networks
This dataset contains synthetic personas, LLM-generated social graphs, cached
text embeddings, and evaluation metrics for clean generation and three
generation-time attack families. All names and profiles are synthetic and do
not represent real people.
Dataset variants
Variant
Nodes
Generator
Graph seeds per condition
Attack rates
p50
50
Qwen3-Max
10
10%, 20%, 30%, 40%, 50%
p200
200… See the full description on the dataset page: https://huggingface.co/datasets/Kevynf/llm-graph-poisoning-data.pixmo-point-count-concat_0-20pixmo-point-count-gen-undmsmarco-llm-reranking-pointwisePDL-Bench
PDL-Bench
PDL-Bench is Poindexter Labs' multi-domain evaluation benchmark: six mini-benchmarks of
specialist, auditable tasks written by domain experts to beat frontier models. Each task ships
with what its grading needs — a final answer, a rubric, a worked solution, or an
executable test suite — so results can be reproduced locally.
This is an open benchmark (HLE-style): the answer key is published alongside the prompts.
See Contamination for the trade-off we accept.
This… See the full description on the dataset page: https://huggingface.co/datasets/Poindexter-Labs/PDL-Bench.image-pointing-1M-sft-swiftmodelnet40_normal_resampled-compressedPointBenchro_sft_pixmo_points
Dataset Description
PixmoPoints is a dataset of images paired with referring expressions and points marking the locations the referring expression refers to in the image.
Here we provide the Romanian translation of the PixmoPoints dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_points.ESdB-Embeddings-for-Sequential-data-Benchmark
ESdB: Embeddings for Sequential Data Benchmark
ESdB provides reproducible splits, evaluation shifts, and downstream targets
for benchmarking representations of sequential data.
This repository contains benchmark annotations only. It does not redistribute
the original events or input features. Original datasets must be obtained from
their respective sources and can be reproduced with the preprocessing code in
the ESdB repository.
Structure
Each dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/On-Point-Rnd/ESdB-Embeddings-for-Sequential-data-Benchmark.point-in-time-us-equity-fundamentals-sample
Tradevo Data — honest point-in-time US equity fundamentals
Fundamentals with filed-date stamps, so a backtest only sees what was public — and restatements are flagged, not silently applied.
A deliberately small public proof pack of point-in-time US equity fundamentals, built from SEC EDGAR.
Every value is stamped with the date it first became public (first_filed), so a join that
filters by first_filed <= as_of only sees what was knowable on that date — and later
revisions are… See the full description on the dataset page: https://huggingface.co/datasets/Tradevodata/point-in-time-us-equity-fundamentals-sample.s3dis-compressedglobal-geo-poison-v1SAM_PointPrompt_Dataset
Abstract
The remarkable capabilities of the Segment Anything Model (SAM) for tackling image segmentation tasks in an intuitive and interactive manner has sparked interest in the design of effective visual prompts. Such interest has led to the creation of automated point prompt selection strategies, typically motivated from a feature extraction perspective. However, there is still very little understanding of how appropriate these automated visual prompting strategies are… See the full description on the dataset page: https://huggingface.co/datasets/gOLIVES/SAM_PointPrompt_Dataset.pixmo-pointsYe
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/POISONX/Ye.soma-umr-release_v260714_a2m_52kGPT_Pointillism_Style_Images
GPT Pointillism Style Images
Dataset description
This is a synthetic GPT-generated pointillism style image dataset. It contains 70 image-caption pairs with colorful dotted texture, painterly lighting, and pointillism-inspired compositions.
The images focus on colorful stippled brushwork, luminous dotted texture, painterly lighting, cozy subjects, gardens, landscapes, interiors, and decorative still-life scenes.
Contents
images/
metadata.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Pointillism_Style_Images.us-warn-act-layoffs-point-in-time-snapshots
US WARN Act layoff notices — point-in-time (as-of) snapshot archive
25 daily vintages, 2026-08-30 → 2026-09-24.
1,067,077 total rows, 42 MB compressed. One new vintage every day, forever.
This is the same US WARN Act layoff dataset as
the daily mirror — except you can load it as it stood on a past
date, instead of only as it stands today.
from datasets import load_dataset
# the table exactly as it was published on 5 September 2026
past =… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-warn-act-layoffs-point-in-time-snapshots.Point-MAE-Zero
Procedural 3D Synthetic Shapes Dataset
Overview
This dataset contains 152,508 procedurally synthesized 3D shapes in order to help people better reproduce results for Semantic-Free Procedural 3D Shapes Are Surprisingly Good Teachers. The shapes are created using a procedural 3D program that combines primitive shapes (e.g., cubes, spheres, and cylinders) and applies various transformations and augmentations to enhance geometric diversity.
Our dataset is collected based on… See the full description on the dataset page: https://huggingface.co/datasets/uva-cv-lab/Point-MAE-Zero.ICSE-2027-public
ICSE 2027 Heap Membership-Inference Dataset — Public Competition Release
This is the public competition release of the ICSE 2027 membership-inference
benchmark for Go, Java, Python, Ruby, and Rust. It contains train and
validation only.
The release contains 54,040 source-code files.
Validation label masking
Each language has 1,000 validation rows in their original order. The first 500
rows retain their membership labels; membership is null for rows 500–999.
In… See the full description on the dataset page: https://huggingface.co/datasets/Poisoned-Chalice/ICSE-2027-public.pixmo-points-filtered-below10_imgContained
