CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face02medalpaca /medical_meadow_medical_flashcards Dataset Card for Medical Flashcards Dataset Summary Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge, and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.textquestion-answering10K<n<100K49 likes7.4k downloads3y agoHugging Face03Jainamshahhh /flashfacts FlashFacts: 13,976 rows of earnings extraction, every one from a real SEC filing A training corpus for pulling structured financial facts out of an earnings exhibit: the revenue figure, the period it covers, the units it is reported in, and the year-over-year change. The model must return exact JSON, and it must return null when the filing does not contain the fact. What this dataset proves, and how you check it rows 13,976, and every one is a real… See the full description on the dataset page: https://huggingface.co/datasets/Jainamshahhh/flashfacts.question-answering10K<n<100K0 likes247 downloads2mo agoHugging Face04flwrlabs /medical-meadow-medical-flashcards Dataset Card for medical-meadow-medical-flashcards This dataset originates from the medAlpaca repository. The medical-meadow-medical-flashcards dataset is specifically used for models training of medical question-answering. Dataset Details Dataset Description Each sample is comprised of three columns: instruction, input and output. Language(s): English Dataset Sources The code from the original repository was adopted to post it here. Repository:… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/medical-meadow-medical-flashcards.textquestion-answering10K<n<100K0 likes209 downloads1y agoHugging Face05KrynexLabs /KrynexAI-Dataset-Flash-Instruction 🧠 KrynexAI Dataset English | Русский 📌 Overview KrynexAI Dataset is a high-quality, synthetically expanded collection of 10,000+ instruction-response pairs designed for fine-tuning Large Language Models (LLMs). The dataset covers a wide range of topics including: 💻 Programming (Python, algorithms, data structures) 🤖 AI & Machine Learning (neural networks, transformers, LLMs) 🔭 Science (physics, cosmology, biology, neuroscience) 🧠 Philosophy & Psychology… See the full description on the dataset page: https://huggingface.co/datasets/KrynexLabs/KrynexAI-Dataset-Flash-Instruction.texttext-generation10K<n<100K1 likes149 downloads2d agoHugging Face06SafwanAlbeshti /deepseek-v4-flash-filler-lens-demo DeepSeek-V4-Flash filler-token lens captures — demo subset Per-position logit-lens activations and top-k attention recorded from deepseek-ai/DeepSeek-V4-Flash on a three-product arithmetic task, with and without filler tokens. This is the public demo subset (7 captures) of a larger private collection. It exists so the attention viewer in the accompanying repo runs without special access. Code, full results and write-up: https://github.com/safwanalbeshti/filler-effect-writeup… See the full description on the dataset page: https://huggingface.co/datasets/SafwanAlbeshti/deepseek-v4-flash-filler-lens-demo.question-answeringn<1K0 likes97 downloads19d agoHugging Face07Rallex3 /glm-5.3-flash-function-calling GLM-5.3-Flash Function Calling (synthetic) A synthetic function-calling dataset generated with zai-org/GLM-5.3-Flash via Hugging Face Inference Providers. 513 examples in 8 domains: weather, calendar, finance, travel, e-commerce, devops, smart home, communication. Categories: single-turn tool calls, parallel/multiple calls in one turn, multi-turn trajectories with tool results, and no-tool-needed turns. Format: OpenAI-style — each row has tools (JSON-schema function… See the full description on the dataset page: https://huggingface.co/datasets/Rallex3/glm-5.3-flash-function-calling.textquestion-answeringn<1K0 likes91 downloads15d agoHugging Face08minkyungpark /flashrag_wiki18_aligned FlashRAG Multi-hop QA — wiki-18 aligned gold_doc_ids Multi-hop QA splits from FlashRAG (HotpotQA, MuSiQue, 2WikiMultiHopQA), extended with a per-sample list of paragraph chunk ids in PeterJinGo/wiki-18-corpus that contain the actual supporting evidence. All original FlashRAG fields are preserved verbatim; only one new field is added: metadata.gold_doc_ids: list[str]. Why this dataset exists FlashRAG's multi-hop QA samples ship gold supervision at the Wikipedia article… See the full description on the dataset page: https://huggingface.co/datasets/minkyungpark/flashrag_wiki18_aligned.textquestion-answering100K<n<1M0 likes82 downloads5mo agoHugging Face09mkurman /med-synth-questions-gemma-3-27b-deepseek-v4-flash Med Synth Questions (Gemma-3 + DeepSeek V4 Flash) Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash. Dataset Summary 29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source) 29,148 reasoning turns (99.2% format compliance) Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.tabulartext-generation10K<n<100K1 likes51 downloads2mo agoHugging Face10leandrodevai /medical-meadow-medical-flashcards-splits Dataset Card for Medical Meadow Medical Flashcards - Fixed Splits Dataset Summary This dataset is a reproducible train/validation/test split of flwrlabs/medical-meadow-medical-flashcards, an English medical question-answering dataset from the MedAlpaca project. The source dataset contains 33,955 flashcards in a single training split. This version preserves the original rows and columns while assigning every example to one of three fixed splits using seed 42. No… See the full description on the dataset page: https://huggingface.co/datasets/leandrodevai/medical-meadow-medical-flashcards-splits.textquestion-answering10K<n<100K0 likes34 downloads3mo agoHugging Face11cs-552-2026-Flash-McQueenS-and-TheKing /safety_sft_data Safety SFT data (CS-552, Flash McQueenS and The King) 3,250 English safety multiple-choice items across the 7 SafetyBench categories, used to fine-tune cs-552-2026-Flash-McQueenS-and-TheKing/safety_model (non-thinking SFT). Categories: Unfairness & Bias (BBQ), Ethics & Morality (ETHICS), Physical Health (SafeText), Offensiveness (TweetEval) — plus LLM-generated Mental Health, Illegal Activities, Privacy & Property. Each item: a question with labelled options; the target is a… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-Flash-McQueenS-and-TheKing/safety_sft_data.textquestion-answering1K<n<10K0 likes30 downloads4mo agoHugging Face12CodeWithSomesh /english-hindi-vocab-flashcardsgatedtexttext-classification1K<n<10K2 likes25 downloads1y agoHugging Face13Srinivasmec26 /Educational-Flashcards-for-Global-Learners 1. Educational-Flashcards-for-Global-Learners/README.md Educational Flashcards Dataset Overview A comprehensive collection of 100 educational flashcards covering STEM, humanities, law, arts, and cultural topics. Curated with 70% Indian content, 25% European, and 5% other Asian perspectives to promote diverse knowledge representation. Dataset Structure { "input": "Text description", "output": { "type": "flashcards", "topic": "Subject name"… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Educational-Flashcards-for-Global-Learners.texttext-classificationn<1K1 likes24 downloads1y agoHugging Face141jia /medical_meadow_medical_flashcards Dataset Card for Medical Flashcards Dataset Summary Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge, and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/1jia/medical_meadow_medical_flashcards.textquestion-answering10K<n<100K0 likes16 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.