datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RULER_50
RULER_50 Official-Code Qwen3 Subset
This dataset is a fixed 50-sample-per-group subset of RULER synthetic tasks.
It was generated from the official NVIDIA/RULER GitHub code, not from a
third-party pre-generated mirror.
Official generation source:
Repository: https://github.com/NVIDIA/RULER
Branch: main
Commit: 38da79d79519ef87aa46ae804f838e1eab7f86d7
Generation entrypoint: scripts/data/prepare.py
Benchmark config: scripts/synthetic.yaml
Generation settings:
tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/VenusChenyy/RULER_50.Venus
Venus: A dataset for fine-grained code generation control
🎉 What is Venus? Venus is the dataset used to train Afterburner (WIP). It is an extension of the original Mercury dataset and currently includes 6 languages: Python3, C++, Javascript, Go, Rust, and Java.
🚧 What is the current progress? We are in the process of expanding the dataset to include more programming languages.
🔮 Why Venus stands out? A key contribution of Venus is that it provides runtime and memory… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/Venus.sycophantic-praise
SyPR Benchmark
This dataset contains fixed input artifacts for evaluating sycophantic praise calibration.
Each row is a single (persona, utterance, prompt_condition) evaluation instance.
The evaluated model response is generated dynamically at evaluation time and is not
included in the benchmark artifact.
Reasoning utterances use real benchmark questions. GSM8K target questions and
final answers are pulled from openai/gsm8k. MMLU-Pro Chemistry and Economics
target questions… See the full description on the dataset page: https://huggingface.co/datasets/vennemeyerd/sycophantic-praise.london_venues_synthetic
London Venues Synthetic Dataset 🇬🇧
Project Overview
This dataset contains 10,000 synthetic rows of fictional venues in London, designed to train and test a Semantic Search & Recommendation System.
The goal of this project was to solve the "problem" in recommendation engines. Real-world user reviews are often messy, sparse, or lack specific "intent" or "vibe" contexts (e.g., explicitly mentioning "good for studying" or "cosy cafe"). By generating synthetic data, we… See the full description on the dataset page: https://huggingface.co/datasets/uleeberber/london_venues_synthetic.fol-data
FOL Reasoning Dataset
A preprocessed and vocabulary-augmented dataset derived from the ProofWriter (Kaggle) OWA splits, built for training a Natural Language → First-Order Logic translation model.
The source dataset contains natural-language premises and questions in English along with structured proof metadata. Our preprocessing adds two things that the original does not provide:
FOL translations — each natural-language statement is converted to First-Order Logic via a rule-based… See the full description on the dataset page: https://huggingface.co/datasets/Venkatdatta/fol-data.vendor-verify
VENDOR_VERIFY
A preference dataset for VENDOR_VERIFY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits
80/10/10 train /… See the full description on the dataset page: https://huggingface.co/datasets/316usman/vendor-verify.ventset
Ventset: Raw & Real Conversations with an Empathic AI
Ventset is a dataset of human-AI dialogues featuring an AI designed to respond with empathy, humor, or tough love. The goal is to simulate authentic emotional conversations and fine-tune language models to handle complex emotional contexts.
⚠️ This dataset is still under development — contributions and feedback are welcome!
⚠️ Some messages may be misinterpreted. The creator is not a psychologist. Misuse or misinterpretation… See the full description on the dataset page: https://huggingface.co/datasets/archIBARBUgrr/ventset.pubmed24To recreate, checkout scripts in download.py, extract.py, ungzip.py
This dataset contains Pubmed Open Access articles up to 18-6-2024.
Language_Identification_v1
Dataset Card for Language Identification Dataset
Dataset Summary
A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications.
Languages and Distribution
Language Distribution:
Urdu 1000
Hindi 1000
Odia 1000
Tamil 1000
Kannada 1000
Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.Code-170k-venda
Dataset Description
Code-170k-venda is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Venda, making coding education accessible to Venda speakers.
🌟 Key Features
176,999 high-quality conversations about programming and coding
Pure Venda language - democratizing coding education
Multi-turn dialogues covering various programming concepts
Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-venda.constitucion-venezuela-1000
Constitucion de Venezuela - Dataset de 1000 Instrucciones
Descripcion General
Este dataset contiene 1000 pares de instruccion-respuesta cuidadosamente curados sobre la Constitucion de la Republica Bolivariana de Venezuela de 1999. Ha sido diseñado especificamente para el entrenamiento y evaluacion de modelos de lenguaje en tareas de comprension y generacion de texto sobre contenido constitucional venezolano.
El dataset abarca los aspectos mas importantes de la… See the full description on the dataset page: https://huggingface.co/datasets/niuvaroza/constitucion-venezuela-1000.venue-manager-v2-agent-traces
Venue Manager v2 Agent Traces
Synthetic 100-case trace capture for Floodlight Venue Manager v2.
Source cases: product/5-idea-venue-manager/2-sport-agnostic-venue-agent/eval/cases/booking_100_message_cases.jsonl
Dataset target: build-small-hackathon/venue-manager-v2-agent-traces
Model: nvidia/Nemotron-Cascade-2-30B-A3B
Runtime: Modal HTTP / vLLM / safetensors / bf16
Privacy: synthetic booking messages only
Proof boundary: trace capture only; not judge readiness, public release… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/venue-manager-v2-agent-traces.Hindi-Marathi-Synonyms
Multilingual Synonyms Dataset (बहुभाषी पर्यायवाची शब्द संग्रह)
Overview
This dataset contains a comprehensive collection of words and their synonyms across multiple Indian languages including Hindi and Marathi. It is designed to assist NLP research, language learning, and applications focused on Indian language processing and cross-lingual applications.
The dataset provides word-synonym pairs that can be used for tasks like:
Semantic analysis
Language learning and… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Hindi-Marathi-Synonyms.bulwer-lytton
Bulwer-Lytton sentences
This repository contains all data studied in "Dark & Stormy: Modeling Humor in the Worst Sentences Ever Written", a research paper that investigates how intentionally bad textual humor differs from standard humor datasets.
Bulwer-Lytton.tsv contains all entries highlighted on the Bulwer-Lytton Fiction Contest (BLFC) website's archive between 1996 and 2024. It has 4 columns: type (set to 'human' for the entries from BLFC), year, sentence, category (or genre… See the full description on the dataset page: https://huggingface.co/datasets/venkatasg/bulwer-lytton.hindi-antonyms
Hindi Antonyms Dataset (हिंदी विलोम शब्दकोश)
Overview
This dataset contains a comprehensive collection of Hindi words and their antonyms (विलोम शब्द). It is designed to assist NLP research, language learning, and applications focused on Hindi language processing. The dataset provides word-antonym pairs that can be used for tasks like:
Semantic analysis
Language learning and education
Text enrichment
Linguistic research
Vocabulary expansion
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/hindi-antonyms.veneto-mistral-dataset
Veneto Mistral Dataset
A conversational dataset for training AI models in Venetian language (vèneto).
Description
This dataset was created to preserve and digitalize the Venetian language through artificial intelligence. It contains approximately 11,800 examples of conversations, texts, and translations in Venetian language, extracted from authentic sources and validated for linguistic quality.
The dataset was specifically designed for fine-tuning large language models… See the full description on the dataset page: https://huggingface.co/datasets/Marco-Danz/veneto-mistral-dataset.Sanskrit-verb-forms
Sanskrit Verb Forms Dataset (संस्कृत धातु रूप संग्रह)
Overview
description: |
A comprehensive dataset containing Sanskrit verb conjugations (dhatu roop) with 10,348 unique entries.
Each entry provides the complete information about a Sanskrit verb form, including:
धातु (Dhatu): The root verb
पद (Pada): Voice of the verb (परस्मैपद/आत्मनेपद)
लकार (Lakara): Tense/mood of the verb
पुरुष (Purusha): Person (प्रथम/मध्यम/उत्तम)
वचन (Vachana): Number (एकवचन/द्विवचन/बहुवचन)… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-verb-forms.IRONWORKS-VENOM-preview
IRONWORKS VENOM
Supply Chain Security Training Dataset — Preview v0.1
by IronGate Digital
What this is
A synthetic instruction-tuning dataset focused on software supply chain security.
Built from real threat intelligence sources including security advisories, research
blogs, and vulnerability databases.
This is an early preview. More datasets are in progress.
Coverage
35,000+ labeled training pairs covering:
Dependency confusion and typosquatting… See the full description on the dataset page: https://huggingface.co/datasets/IronGateDigi/IRONWORKS-VENOM-preview.Pure-Telugu-Alpaca
Pure Telugu Alpaca Dataset
This dataset is a cleaned version of Telugu-MultiTask-Instruct-77K with Telugu keys.
It uses enhanced_prompt as instruction and enhanced_completion as output.
Processing Steps
Extracted enhanced_prompt → సూచన (instruction) and enhanced_completion → అవుట్పుట్ (output)
Filtered to keep only entries with no English letters
Removed duplicate entries
Normalized whitespace
Format
Each entry follows the Alpaca format with… See the full description on the dataset page: https://huggingface.co/datasets/VenkataRamanaKurumallajaddangi/Pure-Telugu-Alpaca.constitucion-venezuela-250
Constitución de Venezuela - Dataset de Instrucciones
Descripción del Dataset
Este dataset contiene 250 pares de instrucción-respuesta basados en la Constitución de la República Bolivariana de Venezuela de 1999. Ha sido diseñado específicamente para el entrenamiento de modelos de lenguaje en tareas de comprensión y respuesta sobre contenido constitucional venezolano.
Información del Dataset
Idioma: Español (es)
Licencia: CC-BY-4.0
Tamaño: 250 ejemplos
Formato:… See the full description on the dataset page: https://huggingface.co/datasets/niuvaroza/constitucion-venezuela-250.ginecologia-venezuela
Ginecología Venezuela Dataset
Descripción del Dataset
Este dataset contiene 250 instrucciones especializadas en ginecología y obstetricia, enfocadas específicamente en el contexto de la salud pública venezolana. Está diseñado para entrenar modelos de lenguaje en el dominio médico ginecológico con consideraciones específicas del sistema de salud venezolano.
Contenido
Tamaño: 250 ejemplos de instrucciones
Idioma: Español (Venezuela)
Dominio: Ginecología y… See the full description on the dataset page: https://huggingface.co/datasets/yensonalvi6/ginecologia-venezuela.venba5kதமிழ் வெண்பாக்கள் ~5000, பதவுரை குறிப்புரையுடன்.
