CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01VenusChenyy /RULER_50 RULER_50 Official-Code Qwen3 Subset This dataset is a fixed 50-sample-per-group subset of RULER synthetic tasks. It was generated from the official NVIDIA/RULER GitHub code, not from a third-party pre-generated mirror. Official generation source: Repository: https://github.com/NVIDIA/RULER Branch: main Commit: 38da79d79519ef87aa46ae804f838e1eab7f86d7 Generation entrypoint: scripts/data/prepare.py Benchmark config: scripts/synthetic.yaml Generation settings: tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/VenusChenyy/RULER_50.text-generation1 likes376 downloads1mo agoHugging Face02Elfsong /Venus Venus: A dataset for fine-grained code generation control 🎉 What is Venus? Venus is the dataset used to train Afterburner (WIP). It is an extension of the original Mercury dataset and currently includes 6 languages: Python3, C++, Javascript, Go, Rust, and Java. 🚧 What is the current progress? We are in the process of expanding the dataset to include more programming languages. 🔮 Why Venus stands out? A key contribution of Venus is that it provides runtime and memory… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/Venus.tabulartext-generation1K<n<10K6 likes159 downloads6mo agoHugging Face03vennemeyerd /sycophantic-praise SyPR Benchmark This dataset contains fixed input artifacts for evaluating sycophantic praise calibration. Each row is a single (persona, utterance, prompt_condition) evaluation instance. The evaluated model response is generated dynamically at evaluation time and is not included in the benchmark artifact. Reasoning utterances use real benchmark questions. GSM8K target questions and final answers are pulled from openai/gsm8k. MMLU-Pro Chemistry and Economics target questions… See the full description on the dataset page: https://huggingface.co/datasets/vennemeyerd/sycophantic-praise.tabulartext-generation10K<n<100K0 likes82 downloads4mo agoHugging Face04uleeberber /london_venues_synthetic London Venues Synthetic Dataset 🇬🇧 Project Overview This dataset contains 10,000 synthetic rows of fictional venues in London, designed to train and test a Semantic Search & Recommendation System. The goal of this project was to solve the "problem" in recommendation engines. Real-world user reviews are often messy, sparse, or lack specific "intent" or "vibe" contexts (e.g., explicitly mentioning "good for studying" or "cosy cafe"). By generating synthetic data, we… See the full description on the dataset page: https://huggingface.co/datasets/uleeberber/london_venues_synthetic.imagetext-generationn<1K0 likes69 downloads8mo agoHugging Face05Venkatdatta /fol-data FOL Reasoning Dataset A preprocessed and vocabulary-augmented dataset derived from the ProofWriter (Kaggle) OWA splits, built for training a Natural Language → First-Order Logic translation model. The source dataset contains natural-language premises and questions in English along with structured proof metadata. Our preprocessing adds two things that the original does not provide: FOL translations — each natural-language statement is converted to First-Order Logic via a rule-based… See the full description on the dataset page: https://huggingface.co/datasets/Venkatdatta/fol-data.texttranslation100K<n<1M0 likes49 downloads5mo agoHugging Face06316usman /vendor-verify VENDOR_VERIFY A preference dataset for VENDOR_VERIFY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate. Format Standard preference / DPO schema — each row: column meaning prompt the request (originally prompt) chosen the human-preferred response rejected a worse response to the same prompt source the dataset/URL the row was harvested from Splits 80/10/10 train /… See the full description on the dataset page: https://huggingface.co/datasets/316usman/vendor-verify.texttext-generation1K<n<10K0 likes46 downloads7d agoHugging Face07archIBARBUgrr /ventset Ventset: Raw & Real Conversations with an Empathic AI Ventset is a dataset of human-AI dialogues featuring an AI designed to respond with empathy, humor, or tough love. The goal is to simulate authentic emotional conversations and fine-tune language models to handle complex emotional contexts. ⚠️ This dataset is still under development — contributions and feedback are welcome! ⚠️ Some messages may be misinterpreted. The creator is not a psychologist. Misuse or misinterpretation… See the full description on the dataset page: https://huggingface.co/datasets/archIBARBUgrr/ventset.texttext-generationn<1K0 likes39 downloads1y agoHugging Face08venera-ai /pubmed24To recreate, checkout scripts in download.py, extract.py, ungzip.py This dataset contains Pubmed Open Access articles up to 18-6-2024. texttext-generation1M<n<10M0 likes30 downloads2y agoHugging Face09Process-Venue /Language_Identification_v1 Dataset Card for Language Identification Dataset Dataset Summary A comprehensive dataset for Indian language identification and text classification. The dataset contains text samples across 10 major Indian languages, making it suitable for developing language identification systems and multilingual NLP applications. Languages and Distribution Language Distribution: Urdu 1000 Hindi 1000 Odia 1000 Tamil 1000 Kannada 1000 Bengali… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Language_Identification_v1.texttext-classification1K<n<10K1 likes28 downloads2y agoHugging Face10michsethowusu /Code-170k-venda Dataset Description Code-170k-venda is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Venda, making coding education accessible to Venda speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Venda language - democratizing coding education Multi-turn dialogues covering various programming concepts Diverse topics: algorithms, data… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-venda.texttext-generation100K<n<1M0 likes27 downloads11mo agoHugging Face11niuvaroza /constitucion-venezuela-1000 Constitucion de Venezuela - Dataset de 1000 Instrucciones Descripcion General Este dataset contiene 1000 pares de instruccion-respuesta cuidadosamente curados sobre la Constitucion de la Republica Bolivariana de Venezuela de 1999. Ha sido diseñado especificamente para el entrenamiento y evaluacion de modelos de lenguaje en tareas de comprension y generacion de texto sobre contenido constitucional venezolano. El dataset abarca los aspectos mas importantes de la… See the full description on the dataset page: https://huggingface.co/datasets/niuvaroza/constitucion-venezuela-1000.textquestion-answeringn<1K0 likes26 downloads1y agoHugging Face12build-small-hackathon /venue-manager-v2-agent-traces Venue Manager v2 Agent Traces Synthetic 100-case trace capture for Floodlight Venue Manager v2. Source cases: product/5-idea-venue-manager/2-sport-agnostic-venue-agent/eval/cases/booking_100_message_cases.jsonl Dataset target: build-small-hackathon/venue-manager-v2-agent-traces Model: nvidia/Nemotron-Cascade-2-30B-A3B Runtime: Modal HTTP / vLLM / safetensors / bf16 Privacy: synthetic booking messages only Proof boundary: trace capture only; not judge readiness, public release… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/venue-manager-v2-agent-traces.text-generation0 likes24 downloads3mo agoHugging Face13Process-Venue /Hindi-Marathi-Synonyms Multilingual Synonyms Dataset (बहुभाषी पर्यायवाची शब्द संग्रह) Overview This dataset contains a comprehensive collection of words and their synonyms across multiple Indian languages including Hindi and Marathi. It is designed to assist NLP research, language learning, and applications focused on Indian language processing and cross-lingual applications. The dataset provides word-synonym pairs that can be used for tasks like: Semantic analysis Language learning and… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Hindi-Marathi-Synonyms.texttext-classification1K<n<10K0 likes23 downloads2y agoHugging Face14venkatasg /bulwer-lytton Bulwer-Lytton sentences This repository contains all data studied in "Dark & Stormy: Modeling Humor in the Worst Sentences Ever Written", a research paper that investigates how intentionally bad textual humor differs from standard humor datasets. Bulwer-Lytton.tsv contains all entries highlighted on the Bulwer-Lytton Fiction Contest (BLFC) website's archive between 1996 and 2024. It has 4 columns: type (set to 'human' for the entries from BLFC), year, sentence, category (or genre… See the full description on the dataset page: https://huggingface.co/datasets/venkatasg/bulwer-lytton.texttext-generation1K<n<10K0 likes21 downloads11mo agoHugging Face15Process-Venue /hindi-antonyms Hindi Antonyms Dataset (हिंदी विलोम शब्दकोश) Overview This dataset contains a comprehensive collection of Hindi words and their antonyms (विलोम शब्द). It is designed to assist NLP research, language learning, and applications focused on Hindi language processing. The dataset provides word-antonym pairs that can be used for tasks like: Semantic analysis Language learning and education Text enrichment Linguistic research Vocabulary expansion Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/hindi-antonyms.texttext-classification1K<n<10K1 likes17 downloads2y agoHugging Face16Marco-Danz /veneto-mistral-dataset Veneto Mistral Dataset A conversational dataset for training AI models in Venetian language (vèneto). Description This dataset was created to preserve and digitalize the Venetian language through artificial intelligence. It contains approximately 11,800 examples of conversations, texts, and translations in Venetian language, extracted from authentic sources and validated for linguistic quality. The dataset was specifically designed for fine-tuning large language models… See the full description on the dataset page: https://huggingface.co/datasets/Marco-Danz/veneto-mistral-dataset.texttext-generation10K<n<100K1 likes17 downloads8mo agoHugging Face17Process-Venue /Sanskrit-verb-forms Sanskrit Verb Forms Dataset (संस्कृत धातु रूप संग्रह) Overview description: | A comprehensive dataset containing Sanskrit verb conjugations (dhatu roop) with 10,348 unique entries. Each entry provides the complete information about a Sanskrit verb form, including: धातु (Dhatu): The root verb पद (Pada): Voice of the verb (परस्मैपद/आत्मनेपद) लकार (Lakara): Tense/mood of the verb पुरुष (Purusha): Person (प्रथम/मध्यम/उत्तम) वचन (Vachana): Number (एकवचन/द्विवचन/बहुवचन)… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Sanskrit-verb-forms.texttext-classification10K<n<100K1 likes16 downloads1y agoHugging Face18IronGateDigi /IRONWORKS-VENOM-preview IRONWORKS VENOM Supply Chain Security Training Dataset — Preview v0.1 by IronGate Digital What this is A synthetic instruction-tuning dataset focused on software supply chain security. Built from real threat intelligence sources including security advisories, research blogs, and vulnerability databases. This is an early preview. More datasets are in progress. Coverage 35,000+ labeled training pairs covering: Dependency confusion and typosquatting… See the full description on the dataset page: https://huggingface.co/datasets/IronGateDigi/IRONWORKS-VENOM-preview.texttext-generation10K<n<100K0 likes14 downloads5mo agoHugging Face19VenkataRamanaKurumallajaddangi /Pure-Telugu-Alpaca Pure Telugu Alpaca Dataset This dataset is a cleaned version of Telugu-MultiTask-Instruct-77K with Telugu keys. It uses enhanced_prompt as instruction and enhanced_completion as output. Processing Steps Extracted enhanced_prompt → సూచన (instruction) and enhanced_completion → అవుట్పుట్ (output) Filtered to keep only entries with no English letters Removed duplicate entries Normalized whitespace Format Each entry follows the Alpaca format with… See the full description on the dataset page: https://huggingface.co/datasets/VenkataRamanaKurumallajaddangi/Pure-Telugu-Alpaca.texttext-generation10K<n<100K0 likes13 downloads1mo agoHugging Face20niuvaroza /constitucion-venezuela-250 Constitución de Venezuela - Dataset de Instrucciones Descripción del Dataset Este dataset contiene 250 pares de instrucción-respuesta basados en la Constitución de la República Bolivariana de Venezuela de 1999. Ha sido diseñado específicamente para el entrenamiento de modelos de lenguaje en tareas de comprensión y respuesta sobre contenido constitucional venezolano. Información del Dataset Idioma: Español (es) Licencia: CC-BY-4.0 Tamaño: 250 ejemplos Formato:… See the full description on the dataset page: https://huggingface.co/datasets/niuvaroza/constitucion-venezuela-250.textquestion-answeringn<1K0 likes12 downloads1y agoHugging Face21yensonalvi6 /ginecologia-venezuela Ginecología Venezuela Dataset Descripción del Dataset Este dataset contiene 250 instrucciones especializadas en ginecología y obstetricia, enfocadas específicamente en el contexto de la salud pública venezolana. Está diseñado para entrenar modelos de lenguaje en el dominio médico ginecológico con consideraciones específicas del sistema de salud venezolano. Contenido Tamaño: 250 ejemplos de instrucciones Idioma: Español (Venezuela) Dominio: Ginecología y… See the full description on the dataset page: https://huggingface.co/datasets/yensonalvi6/ginecologia-venezuela.texttext-generationn<1K1 likes11 downloads1y agoHugging Face22RajuKandasamy /venba5kதமிழ் வெண்பாக்கள் ~5000, பதவுரை குறிப்புரையுடன். text-generation1K<n<10K1 likes8 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.