CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face02medalpaca /medical_meadow_medical_flashcards Dataset Card for Medical Flashcards Dataset Summary Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge, and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.textquestion-answering10K<n<100K49 likes7.3k downloads3y agoHugging Face03ioi-leaderboard /ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limittextn<1K0 likes2.1k downloads2y agoHugging Face04shb777 /gemini-flash-2.0-speech 🎙️ Gemini Flash 2.0 Speech Dataset This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English. 🏅 #1 Trending Audio Dataset in Feb 2025 🏅 Used in training of Kokoro TTS and LLaSA 1B 〽️ Stats Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours) Average duration: 10.83 seconds Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.audiotext-to-speech10K<n<100K60 likes1.4k downloads1y agoHugging Face05openguardrails /tb21-dsv4-flash-0731-dsh Terminal-Bench 2.1 trajectories: DeepSeek-V4-Flash-0731 + dsh sdk-minimal Every trial of this one line, in one place: the 89-task main run, both re-run passes, and the scoring scripts. The trajectories are raw and unedited — each step's reasoning, each tool call, and the verifier's own stdout. This is a re-packaging, not a new measurement. The same files were published before, split across two releases, which made the line look incomplete in both: the first release carried the… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/tb21-dsv4-flash-0731-dsh.texttext-generation10K<n<100K0 likes1.1k downloads3d agoHugging Face06malaiwah /GLM-5.3-Flash-calibration-activations-v1 GLM-5.3-Flash calibration activations v1 (BF16, natural routing) Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.tabularn<1K0 likes891 downloads23d agoHugging Face07TAUR-Lab /Taur_CoT_Analysis_Project___google__gemini-1.5-flash-001text10K<n<100K0 likes877 downloads2y agoHugging Face08Zek-Takai /glm53-flash-harvest GLM-5.3-Flash On-Policy Harvest 86,006 responses / 246,034,910 generated tokens written by zai-org/GLM-5.3-Flash from its reference FP8 weights, across four harvest rounds, 15 registers and both serving modes (22,016 rows carry the model's inline <think>…</think> chain). It is on-policy text: the corpus records what the target model actually generates, which is what a speculative-decoding drafter (EAGLE-3 / DFlash / DSpark family) has to learn to predict. Everything here is MIT.… See the full description on the dataset page: https://huggingface.co/datasets/Zek-Takai/glm53-flash-harvest.tabulartext-generation100K<n<1M3 likes667 downloads20d agoHugging Face09birdsql /bird-critic-1.0-flash-exp BIRD-CRITIC-1.0-Flash BIRD-Critic is the first SQL debugging benchmark designed to answer a critical question: Can large language models (LLMs) fix user issues in real-world database applications? Each task in BIRD-CRITIC has been verified by human experts on the following dimensions: Reproduction of errors on BIRD env to prevent data leakage. Carefully curate test case functions for each task specifically. Soft EX: This metric can evaluate SELECT-ONLY tasks. Soft EX + Parsing:… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-flash-exp.textn<1K8 likes619 downloads6mo agoHugging Face10jzju /wavenet_flashback Dataset Card for "wavenet_flashback" https://cloud.google.com/text-to-speech/docs/reference/rest/v1/text/synthesize#AudioConfig sv-SE-Wavenet-{voice} https://spraakbanken.gu.se/resurser/flashback-dator audioautomatic-speech-recognition10K<n<100K0 likes586 downloads4y agoHugging Face11sollamon /ox-alpha-glm-5.3-flash-distillation-coding-17k-raw Ox Alpha GLM-5.3-Flash Distillation Coding 17K Raw A raw collection of 17,138 synthetic coding samples generated with GLM-5.3-Flash, previously exposed through OpenCode under the stealth-model alias Ox Alpha. The dataset is intended for experimentation with LLM distillation, code-generation models, instruction tuning, supervised fine-tuning, evaluation, and agentic coding systems. text10K<n<100K5 likes488 downloads27d agoHugging Face12skyan1002 /flash-flood-benchmark-data TORRENT — CONUS Flash-Flood Benchmark (L1–L3), agent-friendly Traceable, Observation-constrained, Rapid-Response, Episode–gauge–watershed Network of Testbeds: reproducible flash-flood testbeds for hydrological-response analysis and model intercomparison. This mirror carries the paper-matched v1.0 release (companion paper: TORRENT, Earth System Science Data). Archive of record (v1.0): https://doi.org/10.5281/zenodo.22118051 (concept DOI, always the latest version:… See the full description on the dataset page: https://huggingface.co/datasets/skyan1002/flash-flood-benchmark-data.tabular100K<n<1M0 likes451 downloads27d agoHugging Face13best-distill /glm-5.3-flash-distillation-chat Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI). Split train — successful generations only. field description instruction user prompt from FineToMe response GLM final answer (message.content) reasoning GLM chain-of-thought (reasoning_content), empty if not captured finish stop or length prompt_tokens / completion_tokens / reasoning_tokens usage latency_s request latency source_index original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.tabulartext-generation10K<n<100K5 likes410 downloads8d agoHugging Face14MetonymousAI /Step-3.5-Flash-SFT-No-Tools Step-3.5-Flash-SFT No-Tools Filtered subset of stepfun-ai/Step-3.5-Flash-SFT containing only plain chat rows from the raw JSON shards. Final kept rows: 1493471 No-tool rows before secret filtering: 1495099 Rows removed by accepted secret scan findings: 1628 Primary data files are Parquet shards under data/train-*.parquet. Filter predicate: conversations must be a list, every message must be an object, message roles must be limited to system, user, and assistant, no message may… See the full description on the dataset page: https://huggingface.co/datasets/MetonymousAI/Step-3.5-Flash-SFT-No-Tools.texttext-generation1M<n<10M0 likes349 downloads4mo agoHugging Face15AtomicChat /Qwen3.8-Flash-Next-GGUF-metricstabularn<1K0 likes341 downloads27d agoHugging Face16bjdwh /FlashST-DATAtext2 likes331 downloads2y agoHugging Face17agentic-ptb /dpsk-v4-flash-data dpsk-v4-flash-data Training data built by the AgentPTB arm for cell dpsk-v4-flash — pi / DeepSeek v4-flash @ effort thinking. This is the corpus the arm itself assembled during its 100-hour run: what it downloaded, filtered, rewrote and mixed. It is the input side of the checkpoints published as agentic-ptb/dpsk-v4-flash.h*, and the companion to the run record in agentic-ptb/dpsk-v4-flash-record. field value plot cell dpsk-v4-flash driver pi / DeepSeek v4-flash… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/dpsk-v4-flash-data.text100K<n<1M0 likes291 downloads1mo agoHugging Face1834data /gemini31-flash-lite-traintext10K<n<100K0 likes279 downloads1mo agoHugging Face19programbench /20260729_mini-v2.2.8_gemini-3-5-flashtextn<1K0 likes274 downloads2mo agoHugging Face20malaiwah /glm53-flash-fidelity-root-v1 fidelity--glm53flash.malaiwah.root.bf16 A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from zai-org/GLM-5.3-Flash-BF16. The cut the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-root-v1.tabularn<1K0 likes272 downloads17d agoHugging Face21AlienKevin /SWE-smith-rs-gemini-3-flash-trajectories Trajectories Dataset Top-level fields: messages instance_id resolved model traj_id patch Generated at: 2026-02-27 00:16:39Z Rows: 1449 Shards: 6 Skipped runs (missing/corrupt trajectory): 1 text1K<n<10K1 likes253 downloads7mo agoHugging Face22trjxter /DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x DeepSeek V4 Flash 0731 Teacher Distillation — 40,513 Retained Rows Teacher-distillation corpus generated with deepseek-ai/DeepSeek-V4-Flash-0731. The original manifest contained 45,000 unique seeds. Following generation, QC, retry-based repair, quarantine auditing, and recovery adjudication, 40,513 rows were retained. Composition Bucket Rows Coding 5,601 Agentic 9,982 Cyber blue 13,000 Controlled cyber red 6,999 Tool use 4,931 Total 40,513… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x.tabulartext-generation10K<n<100K4 likes238 downloads1mo agoHugging Face23programbench /20260731_mini-v2.4.2_gemini-3-6-flashtextn<1K0 likes236 downloads2mo agoHugging Face24Reza2kn /Wikipedia-FA-EN-DeepSeek-V4-Flash-0731 Wikipedia Persian to English — DeepSeek V4 Flash 0731 Rolling, machine-generated English translations of Persian Wikipedia articles from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration full_articles_fa_without_en. 129,816 translations are currently published in 26 immutable Parquet shards. The target release contains 129,816 translations; five source rows have empty plain_text and are not translated. Shards are published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.audiotranslation100K<n<1M1 likes220 downloads24d agoHugging Face25TeichAI /DeepSeek-v4-Flash-ChatThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Teich Test This directory contains newline-delimited JSON training examples generated by teich. All assistant responses were generated by deepseek/deepseek-v4-flash. Rows: 6313 Format Each file is newline-delimited JSON where every line is already a training example. Chat-only datasets include messages… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Flash-Chat.text1K<n<10K9 likes199 downloads4mo agoHugging Face26kailasa-ngpt /gemini-3.7-flash-ocr-26-aug-2026 gemini-3.7-flash-ocr-26-aug-2026 Page-image → transcription pairs for finetuning a vision-language model to OCR Devanagari and Tamil printed books. These labels are not human ground truth. They are the output of a teacher model, so its accuracy is the ceiling for anything trained on them. Provenance Teacher model google/gemini-3.7-flash (via OpenRouter, reasoning.effort=low) Page render PyMuPDF at 200 DPI, grayscale JPEG q90 Sampling stratified —… See the full description on the dataset page: https://huggingface.co/datasets/kailasa-ngpt/gemini-3.7-flash-ocr-26-aug-2026.imageimage-to-text10K<n<100K0 likes194 downloads27d agoHugging Face27fireblade2534 /Gemini-2.0-Flash-Aoede-Voiceaudiotext-to-speech1K<n<10K8 likes190 downloads2y agoHugging Face28wjn922-01 /scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928 Scale-SWE DeepSeek V4 Flash 0731 Think Rollouts Successful AweAgent trajectories generated with deepseek-v4-flash-0731 in think mode. Dataset summary Source task instances: 3,393 Rollouts per source instance: 4 Total attempted rollouts: 13,572 Successful exported trajectories: 7,928 Unique instances represented by successful trajectories: 2,250 Scaffold: aweagent Tool-call format: openai_function The export retains assistant reasoning_content, function tool… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928.tabulartext-generation1K<n<10K1 likes189 downloads1mo agoHugging Face29tinkersnot /tb2-k5-n4-flashtext0 likes187 downloads11d agoHugging Face30flwrlabs /medical-meadow-medical-flashcards Dataset Card for medical-meadow-medical-flashcards This dataset originates from the medAlpaca repository. The medical-meadow-medical-flashcards dataset is specifically used for models training of medical question-answering. Dataset Details Dataset Description Each sample is comprised of three columns: instruction, input and output. Language(s): English Dataset Sources The code from the original repository was adopted to post it here. Repository:… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/medical-meadow-medical-flashcards.textquestion-answering10K<n<100K0 likes185 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.