datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RamanBench
RamanBench Dataset Mirror
⚠️ RESEARCH MIRROR ONLY — All datasets are provided for research/educational purposes. Original copyrights remain with original authors. See Sources & Licenses below.
A unified mirror of 87 Raman spectroscopy datasets from the RamanBench benchmark. Wide-format Parquet files for fast, reliable access.
Quick Start
from raman_bench import RamanBenchmark
# Fast mirror access (default)
bench = RamanBenchmark(… See the full description on the dataset page: https://huggingface.co/datasets/HTW-KI-Werkstatt/RamanBench.ramanv-image-captions-realhyper-bot-dataramanv-tts-all-raw
ramanv-tts-all-raw
Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages.
ramanv-image-vlm-instructionpolymarket_bot_dataramanv-image-kaggle-realramanv-image-captions-11rampnet-datasetRampNet is a two-stage pipeline that addresses the scarcity of curb ramp detection datasets by using government location data to automatically generate over 210,000 annotated Google Street View panoramas. This new dataset is then used to train a state-of-the-art curb ramp detection model that significantly outperforms previous efforts. In this repo, we provide our generated curb ramp dataset that we use to train the model.
Each parquet row contains a panoramic image and… See the full description on the dataset page: https://huggingface.co/datasets/projectsidewalk/rampnet-dataset.ramanv-image-editing-pairsramanv-document-ocr-2kalshi_bot_dataramanv-image-real-physics-atmosphereramanv-image-real-style-editorialramanv-image-real-3d-rendersdroid_1.0.1npb_data_appEpstein-Files
The Epstein Files Dataset
Dataset Description
Dataset Summary
This dataset comprises a curated collection of publicly available documents and related materials concerning Jeffrey Epstein. It includes unsealed court filings, FBI reports, DOJ publications, and other official investigative records. These files have been aggregated from reputable public sources, such as the U.S. Department of Justice Epstein Library, House Oversight Committee releases, and… See the full description on the dataset page: https://huggingface.co/datasets/ramvorg/Epstein-Files.ramanv-image-real-restorationramanv-image-captions-11ramanv-image-real-faces-diversityramanv-image-real-commerceportallib-tasks
PorTAL 14-Task Multiple-Choice Suite
This dataset is the normalized task suite used by the training examples in
portallib. It contains 129,212 training examples and
19,548 validation examples. Each row has the following fields:
task: stable task name
prompt: language-model context with no trailing whitespace; it is empty only when a source
sentence places its blank first
choices: candidate continuations with exactly one leading space and no trailing whitespace
gold_idx:… See the full description on the dataset page: https://huggingface.co/datasets/RampPublic/portallib-tasks.processed_bert_dataset
Dataset Card for "processed_bert_dataset"
More Information needed
chess-av-mates
ChessBench Action-Values + Mate-in-N
This dataset is derived from the action-value data released with "Grandmaster-Level Chess Without Search" (DeepMind). It provides:
A reorganization of the original action-value format into a per-position structure: each FEN maps to a list of all legal moves with per-move win probabilities (as provided in the upstream release).
A mate-in-N augmentation: for moves with win probability >= 99.99% or <= 0.01%, Stockfish mate search adds a mate depth… See the full description on the dataset page: https://huggingface.co/datasets/Ramora0/chess-av-mates.ramanv-image-real-pose-depthramanv-stt-domainsramanv-image-vqa-benchmarksmulti-label-class-github-issues-text-classification
Dataset Card for "multi-label-class-github-issues-text-classification"
More Information needed
ramanv-stt-augmented
