CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MiliLab /AnesCorpusThe AnesBench Datasets Collection comprises three distinct datasets: AnesBench, an anesthesiology reasoning benchmark; AnesQA, an SFT dataset; and AnesCorpus, a continual pre-training dataset. This repository pertains to AnesCorpus. For AnesBench and AnesQA, please refer to their respective links: https://huggingface.co/datasets/MiliLab/AnesBench and https://huggingface.co/datasets/MiliLab/AnesQA. AnesCorpus AnesCorpus is a large-scale, domain-specific corpus constructed for… See the full description on the dataset page: https://huggingface.co/datasets/MiliLab/AnesCorpus.texttext-generation1M<n<10M4 likes149 downloads1y agoHugging Face02anezatra /dailydialog DailyDialog - ShareGPT Processed Dataset Summary DailyDialog is a high-quality, multi-turn dialogue dataset containing human-written conversations that cover a wide variety of everyday topics.It is designed to support research in dialogue modeling, conversational AI, and emotion-aware interactions.The dataset emphasizes natural, contextually coherent exchanges that resemble real-world human dialogue, making it ideal for training AI systems that need to handle daily… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/dailydialog.texttext-generation10K<n<100K0 likes145 downloads11mo agoHugging Face03Lots-of-LoRAs /task1486_cell_extraction_anem_dataset Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1486_cell_extraction_anem_dataset Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1486_cell_extraction_anem_dataset.texttext-generationn<1K0 likes121 downloads2y agoHugging Face04stanfordaimlab /anesthesia_literacy Anesthesia Literacy Project Adaptive patient education using large-language models Overview This study explores the potential of Large Language Models (LLMs) like OpenAI's Generative Pretrained Transformer (GPT) versions 3.5 and 4 to enhance the readability of preoperative patient instructions, aiming to align them with the American Medical Association's recommendation of a 6th-grade reading level. Acknowledging that nearly 40% of U.S. adults possess basic or below basic… See the full description on the dataset page: https://huggingface.co/datasets/stanfordaimlab/anesthesia_literacy.imagesummarizationn<1K0 likes119 downloads2y agoHugging Face05Lots-of-LoRAs /task1487_organism_substance_extraction_anem_dataset Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1487_organism_substance_extraction_anem_dataset Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1487_organism_substance_extraction_anem_dataset.texttext-generationn<1K0 likes74 downloads2y agoHugging Face06MiliLab /AnesQAThe AnesBench Datasets Collection comprises three distinct datasets: AnesBench, an anesthesiology reasoning benchmark; AnesQA, an SFT dataset; and AnesCorpus, a continual pre-training dataset. This repository pertains to AnesQA. For AnesBench and AnesCorpus, please refer to their respective links: https://huggingface.co/datasets/MiliLab/AnesBench and https://huggingface.co/datasets/MiliLab/AnesCorpus. AnesQA AnesQA is a bilingual question-answering (QA) dataset designed for… See the full description on the dataset page: https://huggingface.co/datasets/MiliLab/AnesQA.texttext-generation10K<n<100K3 likes57 downloads1y agoHugging Face07Anes-03 /aultra-unified-training-data AUltra Unified Training Data This dataset package contains the reconstructed chat-format training data used for the AUltra Unified defensive cybersecurity and code-assistant fine-tune. The dataset was reconstructed from the original preparation scripts, deterministic seeds, local Hugging Face cache, and the same public upstream dataset. The reconstructed split sizes match the documented training run. Transparency Notice This dataset is an experimental, partially… See the full description on the dataset page: https://huggingface.co/datasets/Anes-03/aultra-unified-training-data.texttext-generation10K<n<100K1 likes38 downloads4mo agoHugging Face08anezatra /persona-chat Persona-Chat Dataset Summary Persona-Chat is a high-quality multi-turn dialogue dataset designed to train conversational AI systems with consistent personality and style. Each participant in the dataset is assigned a persona—a short description or set of traits—which guides their responses throughout the conversation. This dataset enables AI models to learn to maintain coherent personas across dialogue turns and produce responses that reflect consistent characteristics… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/persona-chat.texttext-generation10K<n<100K0 likes34 downloads11mo agoHugging Face09anezatra /empathetic-dialogues-sharegpt EmpatheticDialogues - ShareGPT Processed Dataset Summary EmpatheticDialogues is a large-scale, open-domain dialogue dataset designed to help AI systems recognize, understand, and respond to human emotions more naturally. While humans can easily identify and acknowledge others’ feelings during conversation, this remains a major challenge for artificial dialogue agents due to the lack of high-quality empathetic datasets. This dataset introduces a new benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/anezatra/empathetic-dialogues-sharegpt.texttext-generation10K<n<100K0 likes29 downloads10mo agoHugging Face10JvPetas /aneel-legislacao ANEEL Legislação — Corpus NLP Corpus de documentos legislativos e regulatórios publicados pela ANEEL (Agência Nacional de Energia Elétrica), cobrindo os anos 2016, 2021 e 2022. Estatísticas Métrica Valor Documentos 27.060 Caracteres extraídos ~357 milhões Anos cobertos 2016, 2021, 2022 Score qualidade 1.0 97,2% dos documentos Formato JSON estruturado Tipos de documento Tipo Quantidade texto_integral 18.676 voto 6.979… See the full description on the dataset page: https://huggingface.co/datasets/JvPetas/aneel-legislacao.tabulartext-classification10K<n<100K0 likes29 downloads5mo agoHugging Face11igorktech /anekdots_dialogs Anekdots Dialogs Dataset Dataset Summary The Anekdots Dialogs Dataset is a collection of conversational-style dialogs derived from jokes in the original Anekdots Dataset. The dataset consists of dialogues segmented from jokes, allowing for humorous exchanges between multiple participants. It is well-suited for training conversational AI systems, especially those focusing on humor. The dialogues were automatically segmented using the gpt-4o-mini-2024-07-18 model. Due to… See the full description on the dataset page: https://huggingface.co/datasets/igorktech/anekdots_dialogs.tabulartext-generation100K<n<1M4 likes21 downloads2y agoHugging Face12Lots-of-LoRAs /task1485_organ_extraction_anem_dataset Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1485_organ_extraction_anem_dataset Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1485_organ_extraction_anem_dataset.texttext-generationn<1K0 likes15 downloads2y agoHugging Face13aneeshadas02 /smollm3-3b-base-blind-spots SmolLM3-3B-Base Blind Spots Title & Overview A curated set of failure cases for HuggingFaceTB/SmolLM3-3B-Base, showcasing blind spots discovered while probing the 3B-parameter base pre-training checkpoint released in July 2025. Each entry captures a prompt, the expected aligned behaviour, and the model's actual output. The dataset illustrates common failure patterns observed when probing the base model without any instruction tuning, RLHF, or safety fine-tuning applied.… See the full description on the dataset page: https://huggingface.co/datasets/aneeshadas02/smollm3-3b-base-blind-spots.texttext-generationn<1K1 likes10 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.