datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASPED
ASPED: An Audio Dataset for Detecting Pedestrians
This repo contains the data for the ASPED dataset, presented at ICASSP 2024.
Paper Link, Project Homepage
Pavan Seshadri, Chaeyeon Han, Bon-Woo Koo, Noah Posner, Suhbrajit Guhathakurta, Alexander Lerch
Usage
This dataset contains audio and video recordings of pedestrian activity collected at various locations in and around Georgia Tech.
Labels of pedestrian counts per each second of audio/video are provided as well… See the full description on the dataset page: https://huggingface.co/datasets/urbanaudiosensing/ASPED.ais-em-artifactsais-em-code
(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL
Note: This is a reference implementation accompanying our blog post. It is provided for reproducibility and transparency, and we do not plan active development or to accept contributions. If you encounter issues you can't resolve, please email us at satvik.golechha@dsit.gov.uk or sid.black@dsit.gov.uk.
Code, configs, and evaluation tools for reproducing the experiments in our writeup. We reproduce… See the full description on the dataset page: https://huggingface.co/datasets/asparius/ais-em-code.Environmental_geology-aspect
Aspect
Overview
This repository contains the Aspect dataset, part of the Environmental Geology suite maintained by NORA Research Lab. The data is processed into Cloud-Optimized GeoTIFF (COG) format for efficient geospatial analysis.
Dataset Details
Variable: Aspect
Units: degrees (0-360, 0=N)
Format: Cloud-Optimized GeoTIFF (COG)
Compression: ZSTD
Coordinate System: EPSG:4326 (WGS84)
Spatial Resolution: ~30 metres (1 arc-second)… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Environmental_geology-aspect.wiki_aspWikiAsp is a multi-domain, aspect-based summarization dataset in the encyclopedic
domain. In this task, models are asked to summarize cited reference documents of a
Wikipedia article into aspect-based summaries. Each of the 20 domains include 10
domain-specific pre-defined aspects.wikipedia
Dataset Card for Wikimedia Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/)
with one subset per language, each containing a single train split.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/AspectRS/wikipedia.Persian-Food-Sentiment
Persian Food Sentiment Dataset
This data is orinally from https://hooshvare.github.io/docs/datasets/sa.
BibTeX Citation
If you use this dataset, please cite following paper:
@article{ParsBERT,
title={ParsBERT: Transformer-based Model for Persian Language Understanding},
author={Mehrdad Farahani, Mohammad Gharachorloo, Marzieh Farahani, Mohammad Manthouri},
journal={ArXiv},
year={2020},
volume={abs/2005.12515}
}
baarat-hindi-pretrain-dataASPRM-BON-Evaluation-Dataset-MathThis repository contains the datasets for the paper AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence.
aspectemoAspectEmo dataset: Multi-Domain Corpus of Consumer Reviews for Aspect-Based
Sentiment Analysispixmopoints-imagesASPEDvb
ASPED: Audio-Based Pedestrian Detection Dataset Card
Dataset Summary
The Audio Sensing for PEdestrian Detection (ASPED) v.b dataset is a comprehensive, 1,321-hour roadside collection of audio and video recordings designed for the task of pedestrian detection in the presence of vehicular noise. As urban sound emerges as a cost-effective and privacy-preserving alternative to vision-based or GPS-based monitoring, this dataset addresses the key challenge of detecting… See the full description on the dataset page: https://huggingface.co/datasets/urbanaudiosensing/ASPEDvb.ais-em-main-backup
(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL
Note: This is a reference implementation accompanying our blog post. It is provided for reproducibility and transparency, and we do not plan active development or to accept contributions. If you encounter issues you can't resolve, please email us at satvik.golechha@dsit.gov.uk or sid.black@dsit.gov.uk.
Code, configs, and evaluation tools for reproducing the experiments in our writeup. We reproduce… See the full description on the dataset page: https://huggingface.co/datasets/asparius/ais-em-main-backup.details_Aspik101__WizardVicuna-Uncensored-3B-instruct-PL-lora_unload
Dataset Card for Evaluation run of Aspik101/WizardVicuna-Uncensored-3B-instruct-PL-lora_unload
Dataset Summary
Dataset automatically created during the evaluation run of model Aspik101/WizardVicuna-Uncensored-3B-instruct-PL-lora_unload on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Aspik101__WizardVicuna-Uncensored-3B-instruct-PL-lora_unload.aspect-based-sentiment-analysis-uzbekaspenaspi
ASPI — Ambiguous State Prompt Injection
ASPI is a benchmark that measures LLM-agent vulnerability to prompt injection during a clarification state. It extends AgentDojo (v1.2.2) with an 8-condition design that varies state (execution vs clarification), channel (tool-output vs first-user vs follow-up-user), and wrapper (raw attacker text vs ImportantInstructionsAttack-wrapped) so that the state effect is paired-comparable against the channel effect and the wrapper effect.
When a… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/aspi.visual-genomeZImage-Turbo-200k-multires-aspectbucketed
ZImage-Turbo WebDataset
Generated images from the ZImage-Turbo model using DiffusionDB prompts.
Generation Details
Hardware: 8x NVIDIA RTX 3090 GPUs
Generation Time: ~2 days
Estimated Cost: ~$70 (cloud compute)
Dataset Statistics
Total Samples: 211,081
Total Shards: 216
Samples per Shard: ~1000
Shard Naming Convention
Tarballs are named: {base_resolution}-{aspect_ratio}-{shard_num:04d}-of-{total_shards:04d}.tar
For example:… See the full description on the dataset page: https://huggingface.co/datasets/RareConcepts/ZImage-Turbo-200k-multires-aspectbucketed.muchocine_aspectssynthetic-glass-with-liquid-filled
🥃 Glass Half Full — Synthetic Glass with Liquid Filled
8,000 synthetic images of drinking glasses with varying liquid fill levels,
rendered with Blender Cycles (physically-based path tracer) at 256×256
resolution. Every image ships with perfect YOLO-format bounding-box labels
for two classes — glass and liquid — computed directly from 3D geometry
(no human annotation).
Built for the Existential Glass Analyzer,
a browser-based model that answers the timeless question: is your… See the full description on the dataset page: https://huggingface.co/datasets/Aspirin4/synthetic-glass-with-liquid-filled.MD17-aspirin
Dataset Card for aspirin
Dataset Summary
The aspirin dataset is a molecular dynamics (MD) dataset. The total energy and force labels for each dataset were computed using the PBE+vdW-TS electronic structure method. All geometries are in Angstrom, energies and forces are given in kcal/mol and kcal/mol/A respectively.
Supported Tasks and Leaderboards
aspirin should be used for organic molecular property prediction, a regression task on 1 property. The score used… See the full description on the dataset page: https://huggingface.co/datasets/graphs-datasets/MD17-aspirin.asphalt-3d-laser-texture
Asphalt pavement 3D laser surface height maps with pendulum skid resistance (PTV)
Derivative of the Zenodo record 20606458, "Raw 3D laser scan point clouds and pendulum test value (PTV) and Traction
Watcher One (TWO) friction measurements from 45 asphalt pavement test sections" by Matus Kovac, Matej Brna and Peter Pisca
(University of Zilina, Faculty of Civil Engineering), https://doi.org/10.5281/zenodo.20606458, licensed CC-BY-4.0. This
derivative is redistributed under the… See the full description on the dataset page: https://huggingface.co/datasets/RaymondAllen/asphalt-3d-laser-texture.ethrc_piper_screw_driverdetails_Aspik101__trurl-2-13b-pl-instruct_unload
Dataset Card for Evaluation run of Aspik101/trurl-2-13b-pl-instruct_unload
Dataset Summary
Dataset automatically created during the evaluation run of model Aspik101/trurl-2-13b-pl-instruct_unload on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Aspik101__trurl-2-13b-pl-instruct_unload.details_Aspik101__vicuna-13b-v1.5-PL-lora_unload
Dataset Card for Evaluation run of Aspik101/vicuna-13b-v1.5-PL-lora_unload
Dataset Summary
Dataset automatically created during the evaluation run of model Aspik101/vicuna-13b-v1.5-PL-lora_unload on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Aspik101__vicuna-13b-v1.5-PL-lora_unload.saferdecoding-fine-tuning
Dataset Card for SaferDecoding Fine Tuning Dataset
This dataset aims to fine-tune models in an attempt to defend against jailbreak attacks. It is an extension of SafeDecoding
Dataset Details
Dataset Description
The dataset generation process was adapted from SafeDecoding.
This dataset includes 252 original human-generated adversarial seed prompts, covering 18 harmful categories.
This dataset includes responses generated by Llama2, Vicuna, Dolphin, Falcon… See the full description on the dataset page: https://huggingface.co/datasets/aspear/saferdecoding-fine-tuning.details_Aspik101__StableBeluga-13B-instruct-PL-lora_unload
Dataset Card for Evaluation run of Aspik101/StableBeluga-13B-instruct-PL-lora_unload
Dataset Summary
Dataset automatically created during the evaluation run of model Aspik101/StableBeluga-13B-instruct-PL-lora_unload on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Aspik101__StableBeluga-13B-instruct-PL-lora_unload.TurHistQuAD
Turkish Historic Question Dataset
This data is orinally from https://github.com/okanvk/Turkish-Reading-Comprehension-Question-Answering-Dataset
BibTeX Citation
If you use this dataset, please cite following paper:
@INPROCEEDINGS{9559013,
author={Soygazi, Fatih and Çiftçi, Okan and Kök, Uğurcan and Cengiz, Soner},
booktitle={2021 6th International Conference on Computer Science and Engineering (UBMK)},
title={THQuAD: Turkish Historic Question Answering Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/asparius/TurHistQuAD.details_Aspik101__Llama-2-7b-hf-instruct-pl-lora_unload
Dataset Card for Evaluation run of Aspik101/Llama-2-7b-hf-instruct-pl-lora_unload
Dataset Summary
Dataset automatically created during the evaluation run of model Aspik101/Llama-2-7b-hf-instruct-pl-lora_unload on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Aspik101__Llama-2-7b-hf-instruct-pl-lora_unload.
