datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glassballai
GlassBallAI: A Dataset of LLM Market Predictions (Made with Google Gemini Models)
The dataset contains thousands of stock market predictions generated across multiple Google Gemini models.This dataset contains live-captured inference states that cannot be reproduced due to model updates and information leakage.
Gemini 2.5 Pro Example Evaluations:
Blue lines represent 10-day predictions, while the red line represents the actual trend
Discussion: Hugging Face Forum Website:… See the full description on the dataset page: https://huggingface.co/datasets/louidev/glassballai.pine-of-glass-sessions
Coding agent session traces for thomasmustier/pine-of-glass-sessions
This dataset contains redacted coding agent session traces collected while working on tmustier/pine-of-glass. The traces were exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pine-of-glass-sessions.corruption-glass_blur
Corruption Dataset: Glass_Blur
Dataset Description
This dataset contains corrupted versions of ImageNet-1K images using glass_blur corruption. It is part of the ImageNet-C benchmark for evaluating model robustness to common image corruptions.
Dataset Structure
Train: 1,281,167 corrupted images
Validation: 50,000 corrupted images
Classes: 1000 ImageNet-1K classes
Format: Arrow (Hugging Face Datasets)
Corruption Type: Glass_Blur
Applies glass blur… See the full description on the dataset page: https://huggingface.co/datasets/MarMaster/corruption-glass_blur.rayban-meta-glasses
Ray-Ban Meta Glasses
Brand: Ray-Ban Meta
Item: link
metacam-glasses
metacam-datasets
Currently holding glasses images
SynGallery-abl4-tex-light-glass-frame
SynGallery-abl4-tex-light-glass-frame: + frame variety
Rung 4 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting, glass and frame molding variant + color/roughness/metallic, while freezing camera pose (the only frozen factor). Same schema, source images and index↔painting… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl4-tex-light-glass-frame.spin-glass-benchmarksovos-wake-word-bench-picovoice-view-glass
OVOS wake_word bench — picovoice-view-glass
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-view-glass.robocasa_20260430T030150Z_full_run_setup_wine_glasses ---
pretty_name: RoboCasa Trajectories Single
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
---
# RoboCasa Trajectories Single
This dataset contains one row per RoboCasa trajectory / episode.
## Structure
Each row is one trajectory / episode.
Episode-level JSON is stored inline:
adapted_trajectory
original_trajectory
execution_metadata
Step-level data is stored in aligned sequence columns:… See the full description on the dataset page: https://huggingface.co/datasets/DorianAtSchool/robocasa_20260430T030150Z_full_run_setup_wine_glasses.SynGallery-abl3-tex-light-glass
SynGallery-abl3-tex-light-glass: + glass
Rung 3 of the SynGallery instance-level artwork-recognition ablation ladder. 4,898 MET paintings × 5 camera viewpoints = 24,490 synthetic RGB images at 512×512, paired with their source photos and museum metadata.
In this rung, the scene varies textures, lighting and a glass sheet present with probability 0.25, while freezing frame variant/color, camera pose. Same schema, source images and index↔painting mapping as every other rung — they… See the full description on the dataset page: https://huggingface.co/datasets/patryk-bartkowiak/SynGallery-abl3-tex-light-glass.synthetic-glass-with-liquid-filled
🥃 Glass Half Full — Synthetic Glass with Liquid Filled
8,000 synthetic images of drinking glasses with varying liquid fill levels,
rendered with Blender Cycles (physically-based path tracer) at 256×256
resolution. Every image ships with perfect YOLO-format bounding-box labels
for two classes — glass and liquid — computed directly from 3D geometry
(no human annotation).
Built for the Existential Glass Analyzer,
a browser-based model that answers the timeless question: is your… See the full description on the dataset page: https://huggingface.co/datasets/Aspirin4/synthetic-glass-with-liquid-filled.Synthetic-Glass-Transparent-Packaging-Dataset-Sample
Transparent Packaging & Glass Benchmark
Watch our benchmark breakdown: Why Object Detection Fails on Glass.
Synthetic Transparent Glass & Packaging Dataset
A photorealistic synthetic computer vision dataset for transparent glass and
packaging object detection and instance segmentation. The dataset is
designed for models dealing with challenging transparent and reflective
materials, including transparency, reflections, refractions, specular
highlights, and harsh… See the full description on the dataset page: https://huggingface.co/datasets/Ji0134ch/Synthetic-Glass-Transparent-Packaging-Dataset-Sample.shlyokavitsa-pairs
Shlyokavitsa → Cyrillic restoration pairs
210,236 (Latin, Cyrillic) phrase pairs for restoring shlyokavitsa (Bulgarian typed on a
Latin keyboard) back into Cyrillic. Built from Bulgarian Wikipedia, so it can be shared
under the same licence as its source.
{"latin": "sreshta se na dalbochina okolo", "cyrillic": "среща се на дълбочина около", "n_words": 5, "page_id": 1041}
Filed under translation because that is the closest category the Hub offers, but the task is
script… See the full description on the dataset page: https://huggingface.co/datasets/glassbox/shlyokavitsa-pairs.glasseye-dataset
GlassEye 3-Way Unified Building Defect Dataset (640px)
The GlassEye 3-Way Unified Building Defect Dataset is a curated, multi-domain benchmark corpus designed for training computer vision models to detect structural building defects (cracks, spalling, efflorescence) across close-up façade photography and high-altitude aerial drone surveys.
Primary Repository: GlassEye GitHub
Model Checkpoint: sanjeevafk/glasseye-yolo
Image Count: 2,899 total training images + 289 held-out… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/glasseye-dataset.glassformingface-glasses-inference-v1
Face & Glasses Inference Dataset v1
Dataset Summary
This dataset is generated through a distributed, single-pass inference pipeline designed for face detection and glasses classification. It includes images along with metadata and CLIP embeddings, making it ideal for tasks such as face detection, glasses classification, zero-shot inference, and multimodal research.
Supported Tasks
Face Detection & Glasses Classification: Evaluate models on… See the full description on the dataset page: https://huggingface.co/datasets/jhabikash2829/face-glasses-inference-v1.glassdoorspy-glass-dataXiang_Watermelon_Hair_Cut_glasses_Stand_In_Videos_Captioned
Use Lora From: https://huggingface.co/svjack/Flux_Xiang_lora
Reference Image:
glassdoor_reviewsPrince_Xiang_Z_Image_Turbo_glasses_Images
Use lora from https://huggingface.co/svjack/Prince_Xiang_Z_Image_Turbo_Lora
Xiang Glasses Grid Image
celeba-hq-glassesglass_alloy_compositionThis is an alloy composition datasetpart4-put_in_take_out_glassesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "hand",
"total_episodes": 2688,
"total_frames": 815105,
"total_tasks": 8,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2688"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/gsethia08/part4-put_in_take_out_glasses.microfibres_Glass_filter
Dataset Card for microfibres_Glass_filter
A dataset of annotated images derived from wastewater sludge samples collected using fibreglass filters ("Glass dataset"), supporting microfibre detection and segmentation through deep learning. Each image is manually annotated to identify microfibres, their location, and area, and is designed for the development and benchmarking of computer vision models, especially for environmental monitoring applications.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/femartip/microfibres_Glass_filter.glassdoor_reviews_gpt4_0
Dataset Card for "glassdoor_reviews_gpt4_0"
More Information needed
wan_putting_on_glassesThis dataset contains videos generated using Wan 2.1 T2V 14B.
glass-ceramic_lithium_thiophosphate_electrolytes_
Cite this dataset Guo, H., and Artrith, N. _glass-ceramic lithium thiophosphate electrolytes _. ColabFit, 2024. https://doi.org/10.60732/0a15fe72
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_cpznjcu51bvg_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/glass-ceramic_lithium_thiophosphate_electrolytes_.glass_uncap_restest1glass_uncap_restest2
