CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01iamtarun /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K349 likes37k downloads3y agoHugging Face02iamtarun /code_instructions_120k_alpaca Dataset Card for code_instructions_120k_alpaca This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the original source here. texttext-generation100K<n<1M69 likes1.3k downloads3y agoHugging Face03iamtarun /code_contest_python3_alpaca Dataset Card for Code Contest Processed Dataset Summary This dataset contains coding contest questions and their solution written in Python3. This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source. Columns Description id : unique string associated with a problem description : problem description code : one correct code for the problem… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_python3_alpaca.textquestion-answering1K<n<10K8 likes236 downloads3y agoHugging Face04iamjinchen /DD-VQAimagequestion-answering1K<n<10K0 likes176 downloads2y agoHugging Face05i-am-mushfiq /FirstAidQA FirstAidQA: A Synthetic First-Aid and Emergency-Response Question-Answering Dataset Medical safety notice: FirstAidQA is intended for research and educational purposes. It is not a substitute for professional medical advice, emergency services, certified first-aid training, or clinical judgment. Models trained on this dataset may produce incomplete, outdated, or unsafe responses. Dataset Summary FirstAidQA is an English-language synthetic question-answering… See the full description on the dataset page: https://huggingface.co/datasets/i-am-mushfiq/FirstAidQA.textquestion-answering1K<n<10K8 likes136 downloads2mo agoHugging Face06iamtarun /code_contest_processed Dataset Card for Code Contest Processed Dataset Summary This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source. Columns Description id : unique string associated with a problem description : problem description code : one correct code for the problem language : programming language used for code test_samples : contains inputs and their… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_processed.texttext-generation10K<n<100K3 likes127 downloads3y agoHugging Face07iam-tsr /amazon-customer-support Amazon Customer Support Derived from the TWCS corpus (Kaggle: thoughtvector/customer-support-on-twitter), this dataset contains 200 labelled customer-support interactions for evaluation / fine-tuning purposes. Schema Each line is a JSON object with four fields: Field Type Description query string The customer's raw message (input) action string Agent action taken — resolve or escalated_to_human intent string Classified intent — complaint, question… See the full description on the dataset page: https://huggingface.co/datasets/iam-tsr/amazon-customer-support.textquestion-answeringn<1K0 likes61 downloads8d agoHugging Face08iam-tsr /ragmix RAGmix RAGmix is a heterogeneous, multi-domain evaluation dataset for Retrieval-Augmented Generation (RAG) systems. It mixes real-world document styles—policies, meeting minutes, clinical and scientific text, financial disclosures, job postings, and more—so models can be tested outside a single vertical. Each example pairs a full source document with one grounded question and a reference answer. Source PDFs were obtained from Digital Corpora and converted to markdown for this… See the full description on the dataset page: https://huggingface.co/datasets/iam-tsr/ragmix.textquestion-answeringn<1K0 likes35 downloads2mo agoHugging Face09iamshnoo /qa_metacul qa_metacul Summary qa_metacul is an 800-question multiple-choice benchmark used to evaluate metadata-conditioned language models in the Metadata Conditioned LLMs project. The benchmark tests whether a model can answer culturally and geographically grounded factual questions for different parts of the world, and whether metadata-aware models correctly adapt their answers when continent- or country-level context changes. Paper: https://arxiv.org/abs/2601.15236 Project… See the full description on the dataset page: https://huggingface.co/datasets/iamshnoo/qa_metacul.textquestion-answeringn<1K0 likes20 downloads6mo agoHugging Face10iamjry /ai-basic-law-dataset 台灣人工智慧基本法 訓練資料集 Taiwan AI Basic Law (人工智慧基本法) Q&A dataset for LLM finetuning. Files File Description Entries train.jsonl Full training dataset with oversampling ~5000 fulltext.jsonl Clean article fulltext (20 articles) 38 Data Composition Category Unique Repeat Purpose Article Fulltext Q&A ~157 x15 Verbatim article text with topic anchors Alias Recognition ~109 x10 「基本法」「AI基本法」→ 人工智慧基本法 Legislative Reasons ~35 x3 Background… See the full description on the dataset page: https://huggingface.co/datasets/iamjry/ai-basic-law-dataset.textquestion-answering1K<n<10K0 likes20 downloads7mo agoHugging Face11Iamzoo /mental_health_Chatbot Amod/mental_health_counseling_conversations This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue. Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Iamzoo/mental_health_Chatbot.texttext-generation1K<n<10K0 likes20 downloads1mo agoHugging Face12Iambackup /Nemotron-RL-litmus-bench-v0.1 Dataset Description: Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training. This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Iambackup/Nemotron-RL-litmus-bench-v0.1.textreinforcement-learning1K<n<10K0 likes19 downloads3mo agoHugging Face13IAmSkyDra /HCMUT_FAQtextquestion-answering1K<n<10K1 likes17 downloads2y agoHugging Face14iamkoder001 /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K0 likes16 downloads7mo agoHugging Face15IAMRonHIT /RonDistillMed3Mtextquestion-answering1M<n<10M0 likes13 downloads9mo agoHugging Face165digit /I-am-not-happy-with-your-punishment-and-I-don-t-agree-with-ittexttext-classificationn<1K0 likes4 downloads1y agoHugging Face17iamivan11 /russian-spell-correctiontextquestion-answering1K<n<10K0 likes4 downloads10mo agoHugging Face18IAMRonHIT /medmcqaquestion-answering100M<n<1B0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.