datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers_circleci_workflow_runscifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.Glot500
Glot500 Corpus
A dataset of natural language data collected by putting together more than 150
existing mono-lingual and multilingual datasets together and crawling known multilingual websites.
The focus of this dataset is on 500 extremely low-resource languages.
(More Languages still to be uploaded here)
This dataset is used to train the Glot500 model.
Homepage: homepage
Repository: github
Paper: acl, arxiv
This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.cifar100
Dataset Card for CIFAR-100
Dataset Summary
The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).
Supported Tasks and Leaderboards
image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar100.lambada
Dataset Card for LAMBADA
Dataset Summary
The LAMBADA evaluates the capabilities of computational models
for text understanding by means of a word prediction task.
LAMBADA is a collection of narrative passages sharing the characteristic
that human subjects are able to guess their last word if
they are exposed to the whole passage, but not if they
only see the last sentence preceding the target word.
To succeed on LAMBADA, computational models cannot
simply rely on local… See the full description on the dataset page: https://huggingface.co/datasets/cimec/lambada.SciOL-CI
Scientific Openly-Licensed Publications
This repository contains companion material for the following publication:
Tim Tarsi, Heike Adel, Jan Hendrik Metzen, Dan Zhang, Matteo Finco, Annemarie Friedrich. SciOL and MuLMS-Img: Introducing A Large-Scale Multimodal Scientific Dataset and Models for Image-Text Tasks in the Scientific Domain. WACV 2024.
Please cite this paper if using the dataset, and direct any questions regarding the dataset
to Tim Tarsi
Summary
Scientific… See the full description on the dataset page: https://huggingface.co/datasets/Timbrt/SciOL-CI.cihfh_ci_scan_dataset_bCircuitSense
CircuitSense
This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks.
Dataset Structure
The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.civil_comments
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.TextPecker-1.5M
TextPecker-1.5M: A Dataset for Training and evaluating TextPecker
This repository contains the TextPecker-1.5M dataset, a new benchmark proposed in the paper "TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering".
Code and Project Page
The official implementation and project details for the TextPecker and TextPecker-1.5M dataset can be found on the GitHub repository:
https://github.com/CIawevy/TextPecker
Sample Usage
You… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/TextPecker-1.5M.CIC-IoT-2023CI-VID
📄 CI-VID: A Coherent Interleaved Text-Video Dataset
CI-VID is a large-scale dataset designed to advance coherent multi-clip video generation. Unlike traditional text-to-video (T2V) datasets with isolated clip-caption pairs, CI-VID supports text-and-video-to-video (TV2V) generation by providing over 340,000 interleaved sequences of video clips and rich captions. It enables models to learn both intra-clip content and inter-clip transitions, fostering story-driven generation with… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CI-VID.transformers_flash_attn_cicifar-10-pythondmi-aarhus-predictions
DMI Aarhus Predictions
Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0.
Primary files
File
Purpose
Produced by
predictions_latest.parquet
Current future + verified prediction store
dmi-collector
frontend_snapshot.json
Primary integration contract for the Vercel frontend
dmi-collector
Compatibility files
File
Status
Notes
predictions.parquet
Legacy
Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.transformers_daily_ciCircuitWeaveGeoBenchMeta
GeoBench: A Benchmark for Geometric Image Editing
This repository contains the GeoBench benchmark dataset, introduced in the paper Training-Free Diffusion for Geometric Image Editing.
Project Page & Code: https://github.com/CIawevy/FreeFine
GeoBench is designed to evaluate the capability of diffusion models in geometric image editing tasks. It supports various scenarios including object repositioning, reorientation, reshaping, fine-grained partial editing, structure completion… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/GeoBenchMeta.legacy-satvision-sr-pretrain-small
Satvision Pretraining Dataset - Small
Developed by: NASA GSFC CISTO Data Science Group
Model type: Pre-trained visual transformer model
License: Apache license 2.0
This dataset repository houses the pretraining data for the Satvision pretrained transformers.
This dataset was constructed using webdatasets to
limit the number of inodes used in HPC systems with limited shared storage. Each file has 100000
tiles, with pairs of image input and annotation. The data has been further… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/legacy-satvision-sr-pretrain-small.civit-ai-modelscissp-llmbench
CISSP-LLMBench
cifar100cityscapesversion https://git-lfs.github.com/spec/v1
oid sha256:4bcf87ecfbbb8e07a01b21415a970c8b53a5283bf6872b657040d3f45c9241f7
size 31
transformers_daily_ciIPL-Cityscapes-Illuminants
IPL-CityscapesIlluminants-dataset
Illuminant modified dataset version of the famous autonomous driving semantic segmentation Cityscapes dataset.
Dataset generation
For each image, we generate a flat (constant) light spectrum and compute the pixel reflectances that obtain the RGM pixel values. Once we have the pixel reflectances, we generate different light spectrums of different dominant wavelengths (colors) and saturations and compute the new modified images. We apply… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/IPL-Cityscapes-Illuminants.transformers_pr_cicityscape-adverse
Cityscape‑Adverse
A benchmark for evaluating semantic segmentation robustness under realistic adverse conditions.
Overview
Cityscape‑Adverse extends the original Cityscapes dataset by applying eight realistic environmental modifications—rainy, foggy, spring, autumn, snowy, sunny, night, and dawn—using diffusion‑based image editing. All transformations preserve the original 2048×1024 semantic labels, enabling direct evaluation of model robustness in… See the full description on the dataset page: https://huggingface.co/datasets/naufalso/cityscape-adverse.dmi-aarhus-weather-data
DMI Aarhus Weather Data
Training data and model artifact dataset for the Aarhus weather pipeline. Maintained by Ciroc0.
Primary files
File
Purpose
Produced by
training_matrix.parquet
Current source of truth for training rows and causal observation context
dmi-collector
model_registry.json
Active bucket registry per target
dmi-ml-trainer
model_meta.json
Training timestamp, sample count and training window
dmi-ml-trainer
temperature_models.pkl… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-weather-data.CICIoT2023Small
CICIoT2023
This dataset provides a processed derivative of the CICIoT2023 traffic collection. The repository organizes truncated PCAP files and flow-based CSV extractions aligned to the original CICResearch folder hierarchy.
Processing Workflow
The processing pipeline follows four stages:
Source acquisition from the CICResearch CICIoT2023 portal.
Flow extraction from full PCAP files using TriFlowMeter.
PCAP size reduction by truncating packet payloads to 128 bytes with… See the full description on the dataset page: https://huggingface.co/datasets/somnath0100/CICIoT2023Small.
