CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenGVLab /VideoChat-Flash-Training-Data 🦜 VideoChat-Flash-Training-Data This repos contains all annotaions and most videos for training VideoChat-Flash. 📕 How to use the LongVid data? For video_dir like longvid_subset/coin_grounding_10k_zip, you need to concat this dir to a zip file as follows: cat ego4dhcap_eventunderstanding_2k_zip/* > ego4dhcap_eventunderstanding_2k.zip ✏️ Citation @article{li2024videochatflash, title={VideoChat-Flash: Hierarchical Compression for Long-Context… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat-Flash-Training-Data.video-text-to-text10K<n<100K16 likes29k downloads1y agoHugging Face02malaiwah /GLM-5.3-Flash-TR3-partsbin-v1 GLM-5.3-Flash TR3 parts bin v1 — K6 + K8 payload stores under one transform seed This dataset is the parts bin for the GLM-5.3-Flash TR3 quantization campaign (2026-08-27/28): the complete per-choice payload stores of the two published uniform quants, plus the preparation artifacts and provenance receipts that produced them. malaiwah/GLM-5.3-Flash-TR3-6bpw (uniform K6) malaiwah/GLM-5.3-Flash-TR3-8bpw (uniform K8) What a parts bin is TR3 (trellis) encoding is… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-TR3-partsbin-v1.0 likes24k downloads23d agoHugging Face03RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face04flashinfer-ai /flashinfer-trace FlashInfer Trace We provide an official dataset called FlashInfer Trace with kernels and workloads in real-world AI system deployment environments. FlashInfer-Bench can use this dataset to measure and compare the performance of kernels. It follows the FlashInfer Trace Schema. It is organized as follows: flashinfer_trace/ # Here ├── definitions/ └── workloads/ flashinfer-trace/ # On Hugging Face ├── solutions/ └── traces/ Example solutions and traces directories, featuring… See the full description on the dataset page: https://huggingface.co/datasets/flashinfer-ai/flashinfer-trace.20 likes7.5k downloads3mo agoHugging Face05medalpaca /medical_meadow_medical_flashcards Dataset Card for Medical Flashcards Dataset Summary Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge, and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.textquestion-answering10K<n<100K49 likes7.2k downloads3y agoHugging Face06stepfun-ai /Step-3.5-Flash-SFT Step-3.5-Flash-SFT Step-3.5-Flash-SFT is a general-domain supervised fine-tuning release for chat models. This repository keeps the full training interface in one place: json/: canonical raw training data tokenizers/: tokenizer snapshots for Step-3.5-Flash and Qwen3, released to preserve chat-template alignment compiled/: tokenizer-specific compiled shards for StepTronOSS training Data Format Each raw shard is a JSON file whose top level is a list of examples.… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/Step-3.5-Flash-SFT.text-generation1M<n<10M347 likes6.3k downloads6mo agoHugging Face07hf-internal-testing /transformers_flash_attn_ci0 likes5.7k downloads22h agoHugging Face08laion /laions_got_talent_enhanced_flash_annotations_and_long_captions18 likes5.6k downloads2y agoHugging Face09brandonmusic /GLM-5.3-Flash-BF16-Teacher-Logits GLM-5.3-Flash BF16 teacher logits This dataset contains full-vocabulary float32 teacher logits from the immutable zai-org/GLM-5.3-Flash-BF16 revision a6c167b62691b2bac901344b65cb651a70f53e43. It keeps the sealed final KLD panel qualification-only and publishes the separate non-final calibration panel under role-specific paths. Qualification-only final windows: 25 Qualification-only final prediction positions: 51175 Vocabulary size: 154880 Teacher receipt:… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/GLM-5.3-Flash-BF16-Teacher-Logits.text-generation4 likes5k downloads26d agoHugging Face10richardchencccc /Flash100K3 likes5k downloads27d agoHugging Face11ByteDance-Seed /Multi-SWE-bench-flash 👋 Overview This repository contains the Multi-SWE-bench dataset, introduced in Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving, to address the lack of multilingual benchmarks for evaluating LLMs in real-world code issue resolution. Unlike existing Python-centric benchmarks (e.g., SWE-bench), this framework spans 7 languages (Java, TypeScript, JavaScript, Go, Rust, C, and C++) with 1,632 high-quality instances, curated from 2,456 candidates by 68 expert annotators… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/Multi-SWE-bench-flash.text-generation3 likes3k downloads9mo agoHugging Face12ydshieh /A10_benchmark_flash_attention0 likes2.6k downloads2y agoHugging Face13laion /laions_got_talent_enhanced_just_flash_annotations0 likes2.6k downloads2y agoHugging Face14malaiwah /GLM-5.3-Flash-fidelity-suite-v1 GLM-5.3-Flash Fidelity Suite v1 Historical distribution-fidelity evidence for GLM-5.3-Flash (released 2026-08-26): BF16-reference and FP8-as-served hidden-state captures, a shared LM head, and receipts from the declared capture/replay path. Compatible candidate captures can be compared on matching published positions without holding the 643 GB reference; this is not a universal native-serving or task-quality score. Protocol: the Qwen3.8-27B fidelity-suite v5 methodology… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-fidelity-suite-v1.0 likes2.4k downloads14d agoHugging Face15ioi-leaderboard /ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limittextn<1K0 likes2k downloads2y agoHugging Face16comoZ /osworld-glm-5.3-flash-trajThese are the trajectory results from our GLM-5.3-Flash evaluation on OSWorld. For detailed evaluation results, configuration, and additional information, please refer to the following GitHub issue: https://github.com/xlang-ai/OSWorld/issues/591 image1K<n<10K0 likes1.7k downloads9d agoHugging Face17victorzhu30 /flashvsr-repro-outputs-v2-part1 FlashVSR 复现实验输出 — part1 FlashVSR 复现及 KV cache 驱逐策略消融实验的逐帧推理输出,以 FFV1 无损编码归档。 本 repo 是全部结果的第 1/2 部分。 内容 reds_bscv_dbl_clean_full_sliding reds_bscv_dbl_clean_full_sliding_kv10 reds_bscv_dbl_clean_full_sliding_kv6 reds_bscv_dbl_full_gate reds_bscv_dbl_full_gate_kv10 reds_bscv_dbl_full_gate_kv6 reds_bscv_dbl_full_gate_lfres reds_bscv_dbl_full_gate_lfres_frame reds_bscv_dbl_full_gate_lfres_frame_kv10 reds_bscv_dbl_full_gate_lfres_frame_kv6… See the full description on the dataset page: https://huggingface.co/datasets/victorzhu30/flashvsr-repro-outputs-v2-part1.videon<1K0 likes1.4k downloads25d agoHugging Face18shb777 /gemini-flash-2.0-speech 🎙️ Gemini Flash 2.0 Speech Dataset This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English. 🏅 #1 Trending Audio Dataset in Feb 2025 🏅 Used in training of Kokoro TTS and LLaSA 1B 〽️ Stats Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours) Average duration: 10.83 seconds Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.audiotext-to-speech10K<n<100K60 likes1.4k downloads1y agoHugging Face19tinkersnot /tb2-k5-a21h-flash0 likes1.1k downloads6d agoHugging Face20tinkersnot /tb2-k5-a21l-flash0 likes1.1k downloads6d agoHugging Face21liuxsh9 /Step-3.5-Flash-SFT-code Step-3.5-Flash-SFT-code Code SFT dataset extracted from stepfun-ai/Step-3.5-Flash-SFT, converted to extended OpenAI SFT format, with multi-dimensional quality labels and thinking-mode classification. 602,595 conversations | 40.30 GB | 80 files Dataset Summary Group Records Size Avg Rounds Description single_turn/slow 445,173 20.75 GB 1.0 Single-turn with chain-of-thought reasoning single_turn/fast 9,746 0.12 GB 1.0 Single-turn without reasoning… See the full description on the dataset page: https://huggingface.co/datasets/liuxsh9/Step-3.5-Flash-SFT-code.text-generation100K<n<1M1 likes1.1k downloads6mo agoHugging Face22openguardrails /tb21-dsv4-flash-0731-dsh Terminal-Bench 2.1 trajectories: DeepSeek-V4-Flash-0731 + dsh sdk-minimal Every trial of this one line, in one place: the 89-task main run, both re-run passes, and the scoring scripts. The trajectories are raw and unedited — each step's reasoning, each tool call, and the verifier's own stdout. This is a re-packaging, not a new measurement. The same files were published before, split across two releases, which made the line look incomplete in both: the first release carried the… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/tb21-dsv4-flash-0731-dsh.texttext-generation10K<n<100K0 likes1.1k downloads3d agoHugging Face23jakeoneijk /FlashSR_weights8 likes1.1k downloads2y agoHugging Face24victorzhu30 /flashvsr-repro-outputs-part3 FlashVSR 复现实验输出 — part3 FlashVSR 复现及 KV cache 驱逐策略消融实验的逐帧推理输出,以 FFV1 无损编码归档。 本 repo 是全部结果的第 3/3 部分。 内容 reds_val30_gate reds_val30_gate_L0-15 reds_val30_gate_L15-30 reds_val30_gate_lam1_fifo reds_val30_gate_lam5_fifo reds_val30_gate_lfres reds_val30_gate_lfres_raw reds_val30_gate_lfres_z reds_val30_gate_sink reds_val30_gate_tjump reds_val30_h2o reds_val30_headwise reds_val30_headwise_kv10 reds_val30_ofr_gate reds_val30_ofr_gate_L0-15 reds_val30_ofr_gate_L15-30… See the full description on the dataset page: https://huggingface.co/datasets/victorzhu30/flashvsr-repro-outputs-part3.video1K<n<10K0 likes1k downloads1mo agoHugging Face25TAUR-Lab /Taur_CoT_Analysis_Project___google__gemini-1.5-flash-001text10K<n<100K0 likes893 downloads2y agoHugging Face26malaiwah /GLM-5.3-Flash-calibration-activations-v1 GLM-5.3-Flash calibration activations v1 (BF16, natural routing) Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.tabularn<1K0 likes891 downloads23d agoHugging Face27flashinfer-ai /mlsys26-contest MLSys 2026 FlashInfer-Bench Challenge Dataset This repository contains the FlashInfer-Bench dataset for the MLSys 2026 Kenrel Generation Challenge. This dataset targets to be used in the FlashInfer-Bench benchmark system. It follows the FlashInfer Trace Schema. To use the dataset in the competition, please refer to our starter kit. Download Use this command to download the dataset: git lfs install git clone https://huggingface.co/datasets/flashinfer-ai/mlsys26-contest… See the full description on the dataset page: https://huggingface.co/datasets/flashinfer-ai/mlsys26-contest.11 likes887 downloads6mo agoHugging Face28skyan1002 /flash-flood-benchmark-data TORRENT — CONUS Flash-Flood Benchmark (L1–L3), agent-friendly Traceable, Observation-constrained, Rapid-Response, Episode–gauge–watershed Network of Testbeds: reproducible flash-flood testbeds for hydrological-response analysis and model intercomparison. This mirror carries the paper-matched v1.0 release (companion paper: TORRENT, Earth System Science Data). Archive of record (v1.0): https://doi.org/10.5281/zenodo.22118051 (concept DOI, always the latest version:… See the full description on the dataset page: https://huggingface.co/datasets/skyan1002/flash-flood-benchmark-data.tabular100K<n<1M0 likes861 downloads26d agoHugging Face29tinkersnot /tb2-k5-cd1-cortex-dsh-flash0 likes738 downloads10d agoHugging Face30Zek-Takai /glm53-flash-harvest GLM-5.3-Flash On-Policy Harvest 86,006 responses / 246,034,910 generated tokens written by zai-org/GLM-5.3-Flash from its reference FP8 weights, across four harvest rounds, 15 registers and both serving modes (22,016 rows carry the model's inline <think>…</think> chain). It is on-policy text: the corpus records what the target model actually generates, which is what a speculative-decoding drafter (EAGLE-3 / DFlash / DSpark family) has to learn to predict. Everything here is MIT.… See the full description on the dataset page: https://huggingface.co/datasets/Zek-Takai/glm53-flash-harvest.tabulartext-generation100K<n<1M3 likes635 downloads19d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.