hsb
Datasets
All datasets matching “hsb”HS-Bench-results
HS-Bench Results
Benchmark results for HumanStudy-Bench (HS-Bench): evaluation outputs from running AI agents through reconstructed human-subject experiments.
Dataset description
This dataset contains:
12 studies (study_001–study_012): replicated experiments from published human-subject research (cognition, strategic interaction, social psychology).
Multiple model × agent-design runs per study: e.g. different LLMs (Mistral, GPT, Claude, Gemini, etc.) and presets… See the full description on the dataset page: https://huggingface.co/datasets/fuyyckwhy/HS-Bench-results.multilingual_toxicity_dataset
Multilingual Toxicity Detection Dataset
[2025] We extend our binary toxicity classification dataset to more languages! Now also covered: Italian, French, Hebrew, Hindglish, Japanese, Tatar. The data is prepared for TextDetox 2025 shared task.
[2024] For the shared task TextDetox 2024, we provide a compilation of binary toxicity classification datasets for each language.
Namely, for each language, we provide 5k subparts of the datasets -- 2.5k toxic and 2.5k non-toxic samples.… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/multilingual_toxicity_dataset.ocr_datasets
Combined OCR Dataset for Text Recognition
Dataset Description
This is a large-scale dataset (~11M training, ~0.9M validation images) for Optical Character Recognition (OCR), aggregated from several common benchmarks and sources (see Sources below). It includes scene text, handwritten text, and synthetic images with corresponding text labels.
Training Code: https://github.com/ducto489/lib_ocr
Dataset Structure
./data/
├── train/
│ ├── images/*.jpg… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/ocr_datasets.hsb_audio_corpusThis is a collection of speech recordings in Upper Sorbian. Several speakers have contributed their voice to this dataset.
Audio files are stored in subfolders of the sig folder. The corresponding written text can be found at the same path in the trl folder.
Subfolders are constructed as follows:
sig/ID_of_resource/ID_of_speaker/recording_session/files.wav
resp.
trl/ID_of_resource/ID_of_speaker/recording_session/files.trl
Matching speaker IDs inside different resources indicate the same… See the full description on the dataset page: https://huggingface.co/datasets/zalozbadev/hsb_audio_corpus.AI-vs-Real
🖼️ AI-vs-Real Dataset
A balanced dataset for AI-generated vs Real image classification.This dataset is designed to help researchers, developers, and practitioners build and evaluate models that can distinguish between synthetic (AI-generated) and authentic (human-captured) images.
📊 Dataset Overview
Classes:
0 → AI-generated images
1 → Real (human-captured) images
Balance:The dataset is properly balanced across both classes.This ensures that… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/AI-vs-Real.toxic-vs-clean-dataset
Dataset Name: Toxic vs Clean Text Dataset
Описание
Данный датасет предназначен для обучения бинарного классификатора текстовой безопасности (выявление токсичности, вредоносного контента и скрытых промпт-инъекций).
Структура данных
Датасет разделен на три сплита (train, validation, test) в пропорции 80/10/10 с сохранением стратификации классов.
Каждый сплит содержит колонки:
text (string): Текст запроса или команды.
label (int): Метка класса (0 —… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/toxic-vs-clean-dataset.
