CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01marin-dna /vertebrate-v1-issue473-fullwindow-cds-random-val marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val CDS full-window vertebrate projection sequences for the issue #473 random validation control. The source is the immutable issue #417 accepted-sequence table. The split uniformly samples 16,384 original-orientation CDS rows without replacement using seed 42. Sampling occurs before reverse-complement augmentation. Selected rows are removed from training; reverse complements are then added only to the remaining training… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-cds-random-val.tabular10M<n<100M0 likes883 downloads1mo agoHugging Face02marin-dna /vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered Review status: draft generated for issue #473 review before upload. Human-anchored 255 bp vertebrate sequences for the ccre_enhancer_centered cohort under the full_window policy. The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.tabular10M<n<100M0 likes506 downloads1mo agoHugging Face03jhanglee /youtube-highlights-full YouTube Highlights 完整媒体与标注 本仓库面向数据集协作交付,提供一个可断点续传的完整 tar 文件。解压后即可得到视频、 官方标签转换结果、字段说明和本地可视化检查页。 数据概况 6 个类别:dog、gymnastics、parkour、skating、skiing、surfing 417 个通过 ffprobe 完整性检查的 MP4 315 个 human_mturk 视频:具有 MTurk 人工软投票分数 102 个 weak_match 视频:只有官方自动匹配弱标签 官方清单中另有 1 个当前不可下载的视频,未进入训练标注 19 个已下载视频存在媒体帧数与官方标注帧号差异,保留在数据集中并单独列入复核清单 Linux 下载与解压 BASE_URL="https://huggingface.co/datasets/jhanglee/youtube-highlights-full/resolve/main" wget -c… See the full description on the dataset page: https://huggingface.co/datasets/jhanglee/youtube-highlights-full.tabularn<1K0 likes335 downloads24d agoHugging Face04MirandaAbhilash /vqav2-full-metadatatabular100K<n<1M0 likes312 downloads6mo agoHugging Face05Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes150 downloads1mo agoHugging Face06bbidpa /flutter-full-examples-v1 Flutter Codegen: Full Examples Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps, there's no step history or diff structure here -- each row is a single, standalone goal -> complete file example. This is the whole-code counterpart to flutter-diff-steps-v1, intended for training/evaluating a baseline that generates the entire file in one shot, to compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.tabulartext-generation10K<n<100K0 likes136 downloads17d agoHugging Face07Gradygu3u /spatial-training-full-release-20260604 Spatial Training Full Data Release Full data staging directory for our current Cambrian-P / SSR-style Spatial VLM reproduction work. The directory contains the lightweight reproduction pack plus raw compressed training archives. Local staging uses hardlinks where possible, but upload payload is the full dataset. Size Logical payload: 1076.281 GiB Files: 254 Max single file: 18.0 GiB VSI-590K raw payload: 216.777 GiB Cambrian-S-3M raw payload: 858.004 GiB… See the full description on the dataset page: https://huggingface.co/datasets/Gradygu3u/spatial-training-full-release-20260604.tabularvisual-question-answering1M<n<10M0 likes106 downloads4mo agoHugging Face08fullstack /stargate_s04e01_100topkdiverse_text2vid imagen<1K0 likes101 downloads2y agoHugging Face09hww123 /MADBench-full MADBench-Full: Multi-Agent Anomaly Detection with Tool Use MADBench-Full contains 5,200 fully labeled execution traces from five LLMs solving synthetic escape-room tasks. The updated release adds paired tool-enabled and no-tool runs over two clue domains, making it possible to study tool selection, argument construction, tool execution failures, error recovery, and downstream error propagation in a controlled multi-agent system. Dataset Overview Property… See the full description on the dataset page: https://huggingface.co/datasets/hww123/MADBench-full.tabularother1K<n<10K0 likes93 downloads8d agoHugging Face10Rabrg /gutenberg_fulltabular10K<n<100K1 likes77 downloads1y agoHugging Face11sxiong /MLR_full_trajectory DeepSeek-R1 Reasoning Trajectories This dataset contains raw reasoning trajectories generated by DeepSeek-R1, used in (ICLR 2026) Enhancing Language Model Reasoning with Structured Multi-Level Modeling. The trajectories capture the full reasoning process produced by the model, including hidden chain-of-thought reasoning and final responses. They are provided to support research on reasoning analysis, trajectory supervision, and multi-step reasoning training.… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_full_trajectory.tabular10K<n<100K1 likes72 downloads3mo agoHugging Face12GroupieSteven /fullstackarena-table-study-v1gated FullStackArena three-site table study (private working dataset) Protocol fullstackarena-three-site-table-study-v1 (status qualified): five exact OpenRouter models × three cells × the signed Arenagram (132), RideApp (120) and Auction (100) study task sets = 45 arms, 5,280 scored attempts, max_steps=30, 10 workers on 10 isolated lanes per batch (5 before protocol 2.10). Cells: table1_no_time_capped, table1_no_time_uncapped (no automatic time; the explicit get_website_time() tool… See the full description on the dataset page: https://huggingface.co/datasets/GroupieSteven/fullstackarena-table-study-v1.tabular1K<n<10K0 likes62 downloads22h agoHugging Face13open-llm-leaderboard /fulim__FineLlama-3.1-8B-detailsgated Dataset Card for Evaluation run of fulim/FineLlama-3.1-8B Dataset automatically created during the evaluation run of model fulim/FineLlama-3.1-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fulim__FineLlama-3.1-8B-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face14skapoor18ancde /repro-fuse-full-spectrum-unlearnable-examples-via-spectral-equalization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes46 downloads2mo agoHugging Face15rntc /biomed-fr-v4-enriched-fulltabular1M<n<10M0 likes45 downloads11mo agoHugging Face16violetxi /tb21-eval-qwen35-4b-base-thinking-full [REDACTED] — Terminal-Bench 2.1 Canonical Terminal-Bench 2.1 evaluation of Qwen/Qwen3.5-4B through the served model ID [REDACTED] with Terminus-2. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero: 92 Exception counts: {"AgentTimeoutError": 89, "Timeout": 1, "VerifierTimeoutError": 2} Agent timeouts / context-length events / output-cap events: 89 / 0 / 0 Mean reward / Pass@1: 0.105618 Pass@5: 0.191011 Total input/output/cache… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-base-thinking-full.tabularreinforcement-learningn<1K0 likes35 downloads2mo agoHugging Face17GIZ /vulnerability_training_data_fulltabulartext-classificationn<1K0 likes30 downloads3y agoHugging Face18iamPi /full_b933bfa0e201tabular100K<n<1M0 likes29 downloads11d agoHugging Face19iamPi /full_71310ea50ad0_0915tabular100K<n<1M0 likes26 downloads10d agoHugging Face20happy8825 /ecva_instruct_ver2_fulltuned /hub_data4/seohyun/saves/ecva_instruct/full/sft · happy8825/valid_ecva_clean results Model: /hub_data4/seohyun/saves/ecva_instruct/full/sft Dataset: happy8825/valid_ecva_clean Generated: 2025-12-19 09:18:31Z Metrics Metric Value Total samples 924 With GT 0 Parsed answers 0 Top-1 accuracy 0 Recall@5 0 MRR 0 The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash. EVQA/ECVA Metrics Metric… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/ecva_instruct_ver2_fulltuned.tabularn<1K0 likes21 downloads9mo agoHugging Face21DreamMachines /eval_actuator_unboxing_pi05_sweep_v2_01_freeze_01_fulltabularn<1K0 likes21 downloads3mo agoHugging Face22helloelwin /Qwen3-32B_16episodes_comparisons_full Qwen3-32B_16episodes_comparisons_full This is a pairwise comparison dataset created from SWE-bench evaluation results. Files Qwen3-32B_16episodes_comparisons_full_comparison_pairs.jsonl: JSONL file containing comparison pairs Metadata { "dataset_name": "Qwen3-32B_16episodes_comparisons_full", "model_name": "Qwen3-32B", "num_episodes": 16, "episodes_used": [ 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13… See the full description on the dataset page: https://huggingface.co/datasets/helloelwin/Qwen3-32B_16episodes_comparisons_full.tabularn<1K0 likes18 downloads1y agoHugging Face23YF0808 /tot-cwq-train-full-plan-mid-outputs61tabular100K<n<1M0 likes17 downloads3mo agoHugging Face24danish-foundation-models /croco-munin-apertus-8b-da-simpo-full-50ktabular10K<n<100K0 likes17 downloads2mo agoHugging Face25Hormold /leafly-full-dump-cannabistabular1K<n<10K3 likes16 downloads1y agoHugging Face26dvyomkesh /nemo-grpo-from083-full-edge-curation Nemotron 0.83 Edge-Prompt Curation This private dataset contains edge-prompt curation rollouts for the DGXChen/Tong CoT dataset. Seed edge prompts: 134 New rollout rows after seed exclusion: 7668 New edge prompts: 1648 Full edge prompts, seed plus rollout: 1782 Full dataset rows: 7830 Edge rate over full dataset: 0.2276 The Hugging Face dataset viewer is configured to load only data/full_edge_prompts_seed_plus_rollout.jsonl. The larger rollout and metadata files remain… See the full description on the dataset page: https://huggingface.co/datasets/dvyomkesh/nemo-grpo-from083-full-edge-curation.tabular1K<n<10K0 likes15 downloads4mo agoHugging Face27Kmisener /ASBCA_Fulltabular1K<n<10K0 likes15 downloads2mo agoHugging Face28haoranli-ml /genvf-tcs-fulltabular1K<n<10K0 likes14 downloads5mo agoHugging Face29Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({ train: Dataset({ features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'], num_rows: 12166802 }) }) image1M<n<10M0 likes12 downloads3y agoHugging Face303N3G /qwen3-4b-hard-math-mix-guided-full-rfttabularn<1K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.