datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
TSP_EXECUTION_RUNSkernelbench-v3-runs
KernelBench-v3 — Agent Runs
2071 agent evaluations from the v3 sweep (2026-02): 10 frontier models × {RTX 3090, H100, B200} × 43–58 problems per GPU. Each row is one (model, gpu, problem) triple with correctness, speedup, baseline timing, token usage, cost, and a pointer to the agent's winning solution.py.
Companion datasets:
Infatoshi/kernelbench-v3-problems — 60 problem definitions
Infatoshi/kernelbench-hard-runs — newer KernelBench-Hard sweep (12 models × 7 problems on Blackwell… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-v3-runs.cantabile-runs
cantabile-runs
Work queue and checkpoint store for the Cantabile dynamics study. The directory tree
is the plan — there is no plan file and no database.
main/<song>/<method>/.gitkeep queued, unclaimed
main/<song>/<method>/<seed>/CLAIM-<worker> a worker holds it (mtime = heartbeat)
main/<song>/<method>/<seed>/*.pt done: 5M / 6M / 7M / 8M checkpoints
main/<song>/<method>/<seed>/FAILED crashed, needs a human
A worker lists main/, takes… See the full description on the dataset page: https://huggingface.co/datasets/well-balanced/cantabile-runs.atlas-25-sequential-tool-runtime-upgrade
ATLAS report 25: the sequential tool runtime on verl V1
1. Question and links
Read this first. Every stage of the bring-up ran to its evidence; the report is complete for the correctness acceptance of issue 59 and for its performance stack (a second pass: the call parser fixed after an independent judgement, a boundary rollout at a 1024-token cap, one stacked performance ladder whose first tier, a48k, is now the campaign's default) and for its first research use:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-25-sequential-tool-runtime-upgrade.mimose_runsrussian-names
Russian Names with Popularity Scores
Description
This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.
Usage
The dataset can be loaded using the Hugging Face datasets library.
from datasets import load_dataset
dataset = load_dataset("rustemgareev/russian-names", split='train')… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.RUListening
RUListening: Building Perceptually-Aware Music-QA Benchmarks
Multimodal LLMs, particularly Large Audio Language Models (LALMs), have shown progress in music understanding tasks due to text-only LLM initialization. However, we find that seven of the top ten Music Question Answering (Music-QA) models are text-only models, suggesting these benchmarks rely on reasoning rather than audio perception. To address this limitation, we present RUListening: Robust Understanding through… See the full description on the dataset page: https://huggingface.co/datasets/yongyizang/RUListening.ru_sentiment_dataset
Dataset with sentiment of Russian text
Contains aggregated dataset of Russian texts from 6 datasets.
Labels meaning
0: NEUTRAL
1: POSITIVE
2: NEGATIVE
Datasets
Sentiment Analysis in Russian
Sentiments (positive, negative or neutral) of news in russian language from Kaggle competition.
Russian Language Toxic Comments
Small dataset with labeled comments from 2ch.hk and pikabu.ru.
Dataset of car reviews for machine learning (sentiment analysis)
Glazkova A.… See the full description on the dataset page: https://huggingface.co/datasets/MonoHime/ru_sentiment_dataset.nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/penikmatrumput/nasa-cmapss-rul.PriceFMPaper link: https://arxiv.org/pdf/2508.04875
Model link: https://huggingface.co/RunyaoYu/PriceFM
Github link: https://github.com/runyao-yu/PriceFM
medicines_from_zakupki_gov_ruДанные для исследования существования focal points (https://www.jstor.org/stable/3132148) в гос. закупках лекарств в России.
pixar_movies
Pixar Movies Dataset
A comprehensive dataset of Pixar movies, including details on their release dates, directors, cast, box office performance, and ratings. This dataset is gathered from official sources, including Pixar, Rotten Tomatoes, and IMDb. For more information, visit Pixar.
How the Data is Compiled
All information in this dataset has been collected from public sources, including official information from Pixar, Rotten Tomatoes, and IMDb. Cells are each… See the full description on the dataset page: https://huggingface.co/datasets/RummageLabs/pixar_movies.cars_from_drom.ru_archive_2007-2025More information on the parsing process can be found here: https://github.com/zavzyatiy/drom_archive_parser.
This dataset is also published on Kaggle: https://www.kaggle.com/datasets/assaabramovich/resaled-cars-from-drom-ruarchive-2018-2023/.
Main dataset with all data: drom_archive_2007-2025_full.csv
Dataset with (almost) all configurations from Drom for cars in data: additional_data/drom-24-07-2025-all_main_cars_configurations.csv
Dataset with identification of regions for all cities in… See the full description on the dataset page: https://huggingface.co/datasets/zavzyatiy/cars_from_drom.ru_archive_2007-2025.evo-runrwt-rutabert
RWT-RuTaBERT
Dataset based on Russian Web Tables (RWT), which is a corpus of Russian language tables from Wikipedia.
Only relational tables were chosen from RWT with headers matching selected 170 DBpedia semantic types.
Dataset contains 1 441 349 columns, and has fixed train / test split.
Split
Columns
Tables
Avg. columns per table
Test
115 448
55 080
2.096
Train
1 325 901
633 426
2.093
Train statistics
Most frequent column sizes… See the full description on the dataset page: https://huggingface.co/datasets/aidalab/rwt-rutabert.trace-x-runtime-dataSrilanka-vegetable-prices
Sri Lanka Daily Price Dataset Pipeline
Automates extraction of selected items from the CBSL Daily Price Report PDF and appends them to a long-format CSV. It also enriches each date with rainfall for Nuwara Eliya and Polonnaruwa using Open-Meteo.
Output Schema
Columns in data/price_dataset.csv:
date (YYYY-MM-DD)
item
unit
retail_pettah
retail_dambulla
retail_narahenpita
wholesale_pettah
wholesale_dambulla
rainfall_nuwara_eliya_mm
rainfall_polonnaruwa_mm
source_pdf… See the full description on the dataset page: https://huggingface.co/datasets/Rusiru-erandaka/Srilanka-vegetable-prices.russian-business-registries
Russian Business Registries — Aggregated Statistics
Aggregated, ready-to-analyse slices of Russian state registers. Every figure
comes from an official open-data source; nothing here is modelled, imputed or
estimated. Individual companies are not published — only aggregates, with one
deliberate exception described below.
Собрано из открытых данных российских госреестров. Все цифры — из официальных
источников, без моделирования и досчётов. Публикуются агрегаты, не сведения
об… See the full description on the dataset page: https://huggingface.co/datasets/Roman-Kpro/russian-business-registries.Russian_bank_reviews
Dataset Card for bank reviews dataset
Dataset Summary
The dataset is collected from the banki.ru website.
It contains customer reviews of various banks. In total, the dataset contains 12399 reviews.
The dataset is suitable for sentiment classification.
The dataset contains this fields - bank name, username, review title, review text, review time, number of views,
number of comments, review rating set by the user, as well as ratings for special categories… See the full description on the dataset page: https://huggingface.co/datasets/Romjiik/Russian_bank_reviews.nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/mdrafsanisdani/nasa-cmapss-rul.SpatialEpiBench
SpatialEpiBench
Dataset Summary
SpatialEpiBench is a benchmark collection of 11 spatiotemporal epidemic forecasting datasets. The benchmark covers multiple public-health surveillance modalities, including influenza-like illness surveillance rates, confirmed cases, test positivity, inpatient and outpatient hospitalizations, hospital admissions, doctor visits, and deaths. The datasets span the United States, Canada, and Australia, with daily or weekly temporal resolution… See the full description on the dataset page: https://huggingface.co/datasets/ruiqil/SpatialEpiBench.whisper-browser-benchmarks
whisper-browser-benchmarks
Measurements from a Whisper transcription pipeline running entirely inside a
browser tab: which audio and video containers the browser will actually decode,
how accurate the smallest usable Whisper size is on clean synthetic speech, how
long transcription takes relative to the length of the clip, what the first
load pulls over the wire, and what happens to clips longer than the model's
30-second window.
Everything here was measured, not quoted from a… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/whisper-browser-benchmarks.llm-metric-mrewardbenchRUEmoCorp
RUEmoCorp
The largest publicly available, human-annotated, inter-annotator-agreement-validated emotion dataset for Roman Urdu.
Dataset Overview
RUEmoCorp (Roman Urdu Emotion Corpus) is a large-scale, manually curated, expert-annotated dataset of Roman Urdu social media and conversational texts labeled across 7 emotion categories: joy, anger, sadness, fear, disgust, surprise, and none. It is the training corpus behind roman-urdu-emotion-xlmr-v2 — the… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/RUEmoCorp.repro-chain-of-thought-gradient-descent-runs-sol
Chain-of-Thought Gradient Descent reproduction runs
Immutable outputs for the independent scaled reproduction of ICML 2026 paper
#443, OpenReview uZ8JZ1Lw9a.
gpu-l4-seed443/: successful NVIDIA L4 checkpoint, result JSON, and cost CSV.
figures/: interactive logbook figures and raw CSVs.
poster/: Posterly source, zero-warning gate report, PDF/PNG, and
self-contained poster_embed.html.
reproduction-bundle/: complete clean download-and-rerun bundle.
Successful Job:… See the full description on the dataset page: https://huggingface.co/datasets/JG1310/repro-chain-of-thought-gradient-descent-runs-sol.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/kemhug11/nasa-cmapss-rul.nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/saediscrazy/nasa-cmapss-rul.russian_oil_gas_news_telegram_dataset
Description in English:
Dataset collected from 30 Russian-language Telegram news channels on the topic of Oil ang Gas Industry,
collected and marked up automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url - link to the publication. sourceLink - link to Telegram. subSourceLink - link to the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/russian_oil_gas_news_telegram_dataset.
