CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K274 likes24k downloads3y agoHugging Face02fka /prompts.chat a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts. 📢 Notice This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit: 🌐 Website: prompts.chat 📦 GitHub: github.com/f/awesome-chatgpt-prompts About prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.textquestion-answering1K<n<10K9.8k likes23k downloads18d agoHugging Face03stanford-crfm /air-bench-2024 AIRBench 2024 AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning categories of the regulation-based safety categories in the AIR 2024 safety taxonomy. Dataset Details Dataset Description AIRBench 2024 is a AI safety benchmark that aligns with emerging government regulations and company policies. It consists of diverse, malicious prompts spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.texttext-generation10K<n<100K26 likes7.7k downloads2y agoHugging Face04Columbia-NLP /PUPAThis dataset contains the data presented in the paper PAPILLON: Privacy Preservation from Internet-based and Local Language Model Ensembles. Code: https://github.com/siyan-sylvia-li/PAPILLON texttext-generationn<1K3 likes6.7k downloads1y agoHugging Face05HabibaAbderrahim /Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset Description This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations. It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.imagetranslationn<1K0 likes2.9k downloads1y agoHugging Face06SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes2k downloads7mo agoHugging Face07Trelis /function_calling_extendedgated Trelis Function Calling Dataset UPDATE: As of Dec 5th 2023, there is a v3 of this dataset now available from here. Allows models to be fine-tuned for function-calling. The dataset is human generated and does not make use of Llama 2 or OpenAI! Contains 59 training and 17 test rows Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file, clear_chat Access this dataset by purchasing a license HERE. Alternatively… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_extended.textquestion-answeringn<1K51 likes1.9k downloads3y agoHugging Face08Thytu /ChessInstruct ChessInstruct The ChessInstruct Dataset serves as the foundation for training and fine-tuning Language Models (LLMs) specifically in the realm of chess instruction. Derived from the laion/strategic_game_chess dataset, this meticulously curated dataset encompasses a wide array of annotated instructional chess content. Features of the ChessInstruct Dataset: Rich and Diverse Content: Curated with a broad spectrum of instructional resources including annotated games, strategic analyses… See the full description on the dataset page: https://huggingface.co/datasets/Thytu/ChessInstruct.texttext-generation100K<n<1M22 likes1.7k downloads3y agoHugging Face09siddharthmb /2026.RA.Negotiation-Campaigns Rational-Agent Negotiation Campaigns This public dataset contains the complete selected evidence for the ii_mats/experiments/rational_agents negotiation experiments. It includes raw episode JSON, post-hoc annotations, Markdown and HTML transcripts, committed instances, run manifests, campaign selection and exclusion ledgers, machine-readable analysis tables, figures, and integrity manifests. No contaminated, duplicated, stale, failed, or superseded run is included as selected… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Negotiation-Campaigns.tabulartext-generation10K<n<100K0 likes1.3k downloads1mo agoHugging Face10marin-community /token-counts Marin Token Counts Token counts for all datasets used in Marin pretraining runs. Schema Column Type Description dataset string Dataset identifier marin_tokens int Number of tokens after tokenization category string Content domain (web, code, math, academic, books, etc.) synthetic bool Whether the data is LLM-generated or LLM-translated Categories web — Quality-classified Common Crawl text (Nemotron-CC) code — Source code and… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/token-counts.texttext-generationn<1K1 likes1.3k downloads28d agoHugging Face11LAMDA-NeSy /ChinaTravel ChinaTravel Query Dataset This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). ChinaTravel is an open-ended travel-planning benchmark with compositional constraint validation for language agents. See the paper, Hugging Face paper page, code, and bilingual sandbox database (ModelScope mirror) for the complete benchmark resources. Introduction For a given query, a language agent uses the sandbox tools to collect information and… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel.tabulartext-generation1K<n<10K13 likes1.3k downloads8d agoHugging Face12liupf /ChEBI-20-MM ChEBI-20-MM Dataset Overview The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is designed to provide a comprehensive benchmark for evaluating various models' capabilities in the field of molecular science. This benchmark integrates multi-modal data, including InChI, IUPAC, SELFIES, and images, making it a versatile tool for a wide range of molecular tasks. Dataset Description ChEBI-20-MM is an expansion of the… See the full description on the dataset page: https://huggingface.co/datasets/liupf/ChEBI-20-MM.imagetext-generation10K<n<100K9 likes1.1k downloads2y agoHugging Face13starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes916 downloads2y agoHugging Face14ErenJaegerYeager /ClawBenchPro ClawBenchPro ClawBenchPro is a compact, builder-based workplace-agent benchmark package exported from Nanoclaw. It contains task YAML files, prompts, task-local environment builders, skills, evaluation manifests, provenance metadata, and checksums. Included Splits Dataset Tasks Groups round_01_aligned_mix_800 800 base, hard_aligned, multi_turn_aligned, skills_aligned persona_aligned_mix_200 200 base, hard, multi_turn, skills Directory Layout… See the full description on the dataset page: https://huggingface.co/datasets/ErenJaegerYeager/ClawBenchPro.texttext-generation1K<n<10K0 likes878 downloads4mo agoHugging Face15gplsi /cocoterosCOCOTEROS Dataset V1.1 Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.texttext-generation1K<n<10K0 likes859 downloads11mo agoHugging Face16DeepMostInnovations /saas-sales-conversations saas-sales-conversations Dataset Description This is a synthetic dataset of sales conversations for SaaS (Software as a Service) companies, designed for training sales conversion prediction models. The dataset was created following the methodology presented in "SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization" (Nandakishor M, 2025). The dataset contains realistic dialogues between sales representatives and… See the full description on the dataset page: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations.tabulartext-classification100K<n<1M49 likes795 downloads1y agoHugging Face17Chenyu-Zhou /OR-Space OR-Space A full-lifecycle workspace benchmark for industrial optimization agents. OR-Space evaluates whether language-model agents can work reliably with operations research problems represented as executable, multi-file workspaces. Rather than presenting a self-contained mathematical prompt, each task distributes evidence across business requirements, structured data, source code, execution logs, and solver records. The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.textquestion-answeringn<1K4 likes634 downloads2mo agoHugging Face18MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes624 downloads1y agoHugging Face19syrgkanislab /CausalReasoningBenchmark Automated Causal Reasoning Benchmark Overview The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.tabularquestion-answeringn<1K6 likes514 downloads5mo agoHugging Face20SZLHOLDINGS /thesis-corpus-v18 Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. SZLHOLDINGS/thesis-corpus-v18 The v18 Ouroboros Invariant thesis — LaTeX chapters, the 179 formal blocks (theorem / lemma / definition / axiom environments) as a flat CSV, and the per-version delta ledger that tracks how every formal block evolved v1 → v18. Contents File… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/thesis-corpus-v18.texttext-generationn<1K0 likes480 downloads23d agoHugging Face21AfriSpeech /africa-corpus Africa Corpus Verse-aligned text for 693 African languages, plus several world languages, for building monolingual and parallel corpora. Every language is aligned on a shared verse key, so any single language can be pulled on its own or any two joined into a parallel corpus: Monolingual corpus for any single language African ↔ English (English is the default pair) African ↔ African (e.g. Twi ↔ Yoruba, Hausa ↔ Amharic) African ↔ other language (French, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/africa-corpus.texttranslation10M<n<100M3 likes469 downloads3mo agoHugging Face22Mxode /Chinese-Psychology-Books 免责声明与使用须知 (Disclaimer and Usage Notice) 数据集内容 本数据集包含从互联网上多个来源收集的 中文心理学电子书 的集合。 许可证 本数据集的组织结构、汇编方式以及由维护者添加的任何元数据或注释根据 知识共享署名-非商业性使用 4.0 国际许可协议 (Creative Commons Attribution-NonCommercial 4.0 International License - CC BY-NC 4.0) 提供。这意味着您可以基于非商业目的分享和修改这部分内容,但必须给出适当的署名。 请注意:此 CC BY-NC 4.0 许可证不适用于数据集中包含的原始电子书文件本身。 版权声明 数据集中包含的个别电子书文件极有可能受到版权法保护,其版权归各自的作者、出版商或其他版权所有者所有。 数据集维护者不拥有这些电子书的版权。 这些电子书的来源多样且零散,部分来源可能难以追溯。 使用限制与责任… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Psychology-Books.texttext-generationn<1K9 likes426 downloads1y agoHugging Face23LianHong /zomi-monolingual-corpus Zomi Monolingual Corpus v1.0 The Zomi Monolingual Corpus v1.0 contains 363,401 cleaned, deduplicated, reviewed, and permission-approved Zomi sentences. Zomi is represented with the ISO 639-3 language code ctd (Tedim Chin). Quick start from datasets import load_dataset dataset = load_dataset("LianHong/zomi-monolingual-corpus", split="train") print(dataset.num_rows) # 363401 print(dataset[0]["zomi_text"]) Data fields Field Type Description… See the full description on the dataset page: https://huggingface.co/datasets/LianHong/zomi-monolingual-corpus.texttext-generation100K<n<1M0 likes426 downloads27d agoHugging Face24Berom0227 /tangled-ccs-commits Detecting Multiple Semantic Concerns in Tangled Code Commits using Small Language Models This dataset contains commit data for training and evaluating models on software engineering tasks, specifically focusing on identifying and separating concerns in multi-concern commits. Every tangled (multi-concern) commit in this dataset is composed exclusively of atomic commits from a single repository — resolving a structural weakness in earlier cross-repo tangles (which were trivially… See the full description on the dataset page: https://huggingface.co/datasets/Berom0227/tangled-ccs-commits.texttext-generation1K<n<10K1 likes420 downloads2mo agoHugging Face25jeffboudier /yc-companies-august-2025 Y Combinator Companies Dataset Dataset Description This dataset contains information about 5,404 Y Combinator funded companies that have been publicly launched, sourced from the YC-OSS-API. Dataset Summary Total Companies: 5,404 Time Range: Summer 2005 - Summer 2025 Update Frequency: Snapshot from August 2025 Source: YC-OSS-API Dataset Structure Data Fields id: Unique identifier for each company name: Company name… See the full description on the dataset page: https://huggingface.co/datasets/jeffboudier/yc-companies-august-2025.imagetext-classification1K<n<10K1 likes389 downloads1y agoHugging Face26liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes382 downloads6mo agoHugging Face27p11-p11 /chess_datasets Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/p11-p11/chess_datasets.texttext-generation1M<n<10M0 likes377 downloads2y agoHugging Face28ghananlpcommunity /ghana-corpus This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana Corpus Verse-aligned text for Ghanaian languages, plus several world languages, for building monolingual and parallel corpora. Every language is aligned on a shared verse key, so any single language can be pulled on its own or… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-corpus.texttranslation1M<n<10M0 likes373 downloads3mo agoHugging Face29Ahmed-Selem /Shifaa_Arabic_Medical_Consultations Shifaa Arabic Medical Consultations 🏥📊 Overview 🌍 Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses. 🔍 Why is this dataset important? First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.textquestion-answering10K<n<100K13 likes364 downloads2y agoHugging Face30Abhishekcr448 /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M20 likes328 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.