CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K274 likes24k downloads3y agoHugging Face02ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1k downloads1y agoHugging Face03asingh15 /amazon-c2-varied-rubrics Amazon C2 varied-rubric distillation This release exposes six balanced C2 SFT configurations: latent-state and non-diverse candidate panels at K=1, K=2, and K=4 rubrics per retained reviewer. Each rubric-writer target is paired with one full-rubric listwise judge target over the same variant's frozen 40-candidate panel. The K arms within a variant share one reviewer cohort and are exact nested prefixes. Config Train rows Validation Test Train reviewers… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-c2-varied-rubrics.tabulartext-generation100K<n<1M0 likes141 downloads4d agoHugging Face04eugeneshilow /rubench RuBench RuBench is a repository-level agentic coding benchmark whose task specifications are natively authored in Russian. Each task is a real bug from a live open-source repository (aiohttp, aiogram, Laravel, NestJS, or Fastify), specified in Russian in the style of an actual customer request rather than translated from an English issue. Solutions are judged by the upstream maintainers' regression tests, which are withheld from this release. Paper: arXiv:2607.06411 Hub and… See the full description on the dataset page: https://huggingface.co/datasets/eugeneshilow/rubench.tabulartext-generationn<1K2 likes87 downloads2mo agoHugging Face05ammarnasr /the-stack-ruby-clean Dataset 1: TheStack - Ruby - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Ruby, a popular statically typed language. Target Language: Ruby Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Ruby as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-ruby-clean.tabulartext-generation100K<n<1M3 likes82 downloads3y agoHugging Face06philosopher-from-god /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K2 likes73 downloads1y agoHugging Face07jebkoralav /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/jebkoralav/ru-big-russian-dataset.tabulartext-generation1M<n<10M0 likes62 downloads10mo agoHugging Face08Emrahisik /rubric-dataset Rubric Dataset — Startup Investability & Digital Marketing Training and evaluation data for a single behaviour: fill in a rubric, do not score. The model never produces the overall number. For each criterion it returns one rating and the verbatim quotes that justify it — or it says the case is silent on that criterion. The weighted total is computed deterministically outside the model. This dataset trains the first half of that split; the arithmetic half is not a model problem.… See the full description on the dataset page: https://huggingface.co/datasets/Emrahisik/rubric-dataset.tabulartext-generation1K<n<10K0 likes50 downloads2mo agoHugging Face09asingh15 /glm52-aligned-rubric-traces GLM-5.2 Aligned Rubric-Writing Traces 23777 teacher traces from GLM-5.2 on the aligned rubric-writing task, collected to distill / warmstart a smaller rubric-writer. For each (user, book) example the teacher is shown a persona-conditioned prompt (a user's past book reviews) and asked to (1) predict what that user would likely write about a new book and (2) produce a <rubric> of numbered criteria for scoring candidate reviews on coverage of that prediction. The full generation —… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/glm52-aligned-rubric-traces.tabulartext-generation10K<n<100K1 likes37 downloads2mo agoHugging Face10mehmetdavut /RubyCraft-3.4-Eval-Logs 🚀 RubyCraft-3.4 Evaluation Logs This dataset contains the comprehensive evaluation logs, including raw and processed outputs, for our research on the adaptation of Small Language Model (SLM) architectures to Ruby 3.4 syntax. It covers more than 26,000 evaluation rows generated across 96 LoRA configurations, 4 base models, and multiple teacher models. ⚡ Quick Performance Summary (The DSP Impact) Our Diagnostic Sanitization Procedure (DSP) revealed massive hidden… See the full description on the dataset page: https://huggingface.co/datasets/mehmetdavut/RubyCraft-3.4-Eval-Logs.tabulartext-generation10K<n<100K0 likes29 downloads5mo agoHugging Face11xxx-cha666-xxx /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K0 likes23 downloads10mo agoHugging Face12NovelHacja /RubricHub_v1_config RubricHub RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of… See the full description on the dataset page: https://huggingface.co/datasets/NovelHacja/RubricHub_v1_config.tabulartext-generation100K<n<1M2 likes15 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.