CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01elidek-themis /experimentstabular100K<n<1M0 likes9k downloads21d agoHugging Face02evalstate /transformers-merge-experimentstabularn<1K3 likes1.7k downloads5mo agoHugging Face03LakoreAI /bert-mlm-experiments-en Unified English MLM Pre-training Corpus (80M Rows) This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings. Dataset Details Repository ID: 8Opt/bert-mlm-experiments-en Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.textfill-mask10M<n<100M1 likes1.2k downloads3mo agoHugging Face04kshitijd /platonic-all-experimentstabularn<1K0 likes937 downloads29d agoHugging Face05ASKabalan /jax-fli-experiments jax-fli experiments Data, samples, and reference catalogs for the jax-fli forward-modelling experiments. Each experiment is exposed as one or more HuggingFace dataset configs; load a config with datasets.load_dataset. Experiment 00 — CosmoGrid reference A single CosmoGrid simulation (cosmo_000001) packaged as jax-fli Catalog parquet files, used as reference truth for the lensing / masked-shear experiments. Config Field Shape NSIDE Description… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-experiments.tabular1K<n<10K0 likes586 downloads15d agoHugging Face06namezz /soft-prompt-experiments-archive-20260918 Soft prompt 实验归档 用于查阅和恢复的历史研究记录,涵盖数学与代码任务。共 45 个运行目录,包含教师生成数据、评测输出、原始配置和已有 prompt 检查点。部分目录仅有评测、复核或失败记录,不能将目录数量理解为成功实验数量。 快速查阅 实验总览:模型系列、任务、规模与记录状态。 CSV 索引 / JSON 索引:便于筛选和定位。 archives/:按实验分别压缩的原始文件。 manifests/:各文件 SHA-256 与归档路径。 系列包括 AReaL Boba2、GPT-OSS/Swallow、MiMo、X-Coder/Qwen3、Nemotron、Klear、Mellum2、OLMo3、Polaris 和 Poro2。页面不展开具体方法或实现细节;原始配置仍保留在归档内供恢复。 状态与注意事项 上传完成以 ARCHIVE_COMPLETE.json 为准;文件不存在时表示仍在上传。 每个归档都经过完整下载的 SHA-256 校验。… See the full description on the dataset page: https://huggingface.co/datasets/namezz/soft-prompt-experiments-archive-20260918.tabularn<1K0 likes464 downloads4d agoHugging Face07Changyeli03 /hle-flowbench-experiments-20260829 HLE FlowBench experiment archive Private migration snapshot of the local HLE text-only 100-question research program through 2026-09-01. It preserves the formal and smoke runs, per-question Codex/Claude/Kimi sessions, workflow attempts and metrics, evaluator state, scores, monitoring, experiment controllers, reports, analyses, the paused-run migration package, source Git bundles, and HLE-related host orchestration sessions. The current Chinese experiment status, validity… See the full description on the dataset page: https://huggingface.co/datasets/Changyeli03/hle-flowbench-experiments-20260829.tabularn<1K0 likes455 downloads21d agoHugging Face08publicus-ai /cibench-experiments CIBench Experiments Reproducibility packages for CIBench — the stateless, replayable benchmark engine for the 1M–10M token long-context era. If a benchmark result cannot be replayed from its manifest alone, it did not happen. Every sub-directory in this dataset is a self-contained experiment package: per-run manifests, content-addressed canonical JSON, ResultRecord with full scoring + signed provenance, per-item OpenTelemetry gen_ai_* call metrics, retrieved evidence, a… See the full description on the dataset page: https://huggingface.co/datasets/publicus-ai/cibench-experiments.texttext-retrieval1K<n<10K0 likes431 downloads5mo agoHugging Face09iyatomilab /SciGA-for-experiments-hfimage10K<n<100K0 likes341 downloads1y agoHugging Face10legesher /language-decoded-experiments Language Decoded — Experiment Tracking Central hub for training logs, configurations, evaluation results, and analysis for the Language Decoded project. The project originated as a proposal during Cohere's Tiny Aya Expedition (March 2026 hackathon) and was extended into Phase 3 for the accompanying paper. Submitted paper title (2026-05-26): Language, Decoded: Exploring the Impact of Fine-Tuning a Multilingual Model on Native-Language Code ⚠️ Phase 3 numbers — read… See the full description on the dataset page: https://huggingface.co/datasets/legesher/language-decoded-experiments.tabular10K<n<100K2 likes334 downloads2mo agoHugging Face11TechyCode /bon-pim-hacking-experimentstext10K<n<100K0 likes255 downloads2mo agoHugging Face12tigerking009 /zoya-image-1-experiments ZOYA IMAGE-1 — Reproducible GGUF Experiments Purpose This dataset stores reproducible ZOYA IMAGE-1 image-generation experiments together with the exact generation parameters, model identities, SHA256 fingerprints, and validation reports. The package is designed for controlled comparisons where the tested variable is changed explicitly and all other relevant variables remain fixed. Current baseline Experiment ID: ZOYA_PHASE0_BASELINE_00001… See the full description on the dataset page: https://huggingface.co/datasets/tigerking009/zoya-image-1-experiments.imageimage-to-imagen<1K0 likes250 downloads1mo agoHugging Face13umairinayat /sanad_experimentsimage10K<n<100K0 likes247 downloads3mo agoHugging Face14kamel-usp /jbcs2025_experiments_report JBCS 2025: Experimental Artefacts for AES in Brazilian Portuguese This repository contains all experimental artefacts (logs, configurations, predictions, and evaluation results) described in the paper: Exploring the Usage of LLMs for Automatic Essay Scoring in Brazilian Portuguese EssaysAndré Barbosa, Igor Cataneo Silveira, Denis Deratani MauáTODO 📦 What's in this dataset repo? This dataset is not a training dataset. Instead, it provides comprehensive logs and… See the full description on the dataset page: https://huggingface.co/datasets/kamel-usp/jbcs2025_experiments_report.tabularn<1K0 likes245 downloads1y agoHugging Face15ebcandir /synthetic_experimentstext1M<n<10M0 likes230 downloads1y agoHugging Face16umairinayat /shifaa_experimentsimage10K<n<100K0 likes211 downloads3mo agoHugging Face17inaam1995 /commonvoice17_su_experiments Common Voice 17 -- Single / Long Utterance experiment dataset Built from fixie-ai/common_voice_17_0 (English); the original CV splits are preserved and each is bucketed into single-utterance (1 word) and long-utterance (>= 3 words). Splits: dev_single, dev_long, test_single, test_without_single, train_single, train_long. audio1M<n<10M0 likes171 downloads3mo agoHugging Face18humair-experiments /genshin-voice-english-fXLaudio100K<n<1M0 likes124 downloads4mo agoHugging Face19RL-Forgetting-Experiments-3 /mbpp-code-rl MBPP for code RL (deduplicated against MBPP+) MBPP prepared for RLVR training in verl, with two independent hold-outs so both MBPP+ and MBPP's own canonical test split stay reportable after training on this data. split rows contents train 320 MBPP canonical train + validation + prompt, minus everything in MBPP+ test 378 exactly the problems in evalplus/mbppplus heldout_mbpp_test 276 MBPP's canonical test split (task_id 11-510) that is not in MBPP+… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-rl.texttext-generationn<1K0 likes110 downloads12d agoHugging Face20siddharthmb /2026.mechaptcha.linear-probe-experiments-giant-20260525 siddharthmb/2026.mechaptcha.linear-probe-experiments-giant-20260525 Paired CAPTCHA image experiments for linear probes over a trained Mechaptcha CNN. Each example contains a matched image_a and image_b pair generated from the same seed pool. Intended Use This dataset is designed for linear probe experiments that compare activations from Batch A against Batch B. Use label 1 for image_a and label 0 for image_b. Recommended checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.mechaptcha.linear-probe-experiments-giant-20260525.imageimage-classification1M<n<10M0 likes108 downloads4mo agoHugging Face21Praful932 /abmelt-experiments-exp_20260220_125440textn<1K0 likes104 downloads7mo agoHugging Face22Praful932 /abmelt-experiments-exp_20260220_130124textn<1K0 likes103 downloads7mo agoHugging Face23rmems /grok-1-ternary-quant-experiments Grok-1 SAAQ quantization / route-preservation experiments Dataset author: Raul Montoya Cardenas (rmems) SAAQ stands for Spiking Adaptive Activity Quantization, a term coined by the dataset author. Attribution: Grok Build: Grok 4.5 (high) packaged the original 2026-08-10 dataset. Codex: GPT-5.6-Sol (OpenAI) implemented, executed, validated, and published the canonical issue #85 v4 evidence added on 2026-08-24. Personal research measuring route preservation when packing open… See the full description on the dataset page: https://huggingface.co/datasets/rmems/grok-1-ternary-quant-experiments.tabularothern<1K0 likes96 downloads28d agoHugging Face24mncai /fm-model-experiments-data FM Model Experiments — Synthetic VLM Training Data (KO + EN) Annotation data produced while building a native-resolution Korean+English VLM (GLM-4.6V vision tower transplanted onto a frozen GLM-5.2 743B MoE decoder). Code + technical report: https://github.com/genonai/fm-model-experiments This repo contains ANNOTATIONS ONLY (.jsonl). No images are redistributed. Every row references an image by a relative path (data/...); obtain the images from the original sources listed below… See the full description on the dataset page: https://huggingface.co/datasets/mncai/fm-model-experiments-data.textvisual-question-answering1K<n<10K0 likes84 downloads2mo agoHugging Face25Tonic /trackio-experiments Trackio Experiments Dataset This dataset stores experiment tracking data for ML training runs, particularly focused on SmolLM3 fine-tuning experiments with comprehensive metrics tracking. Dataset Structure The dataset contains the following columns: experiment_id: Unique identifier for each experiment name: Human-readable name for the experiment description: Detailed description of the experiment created_at: Timestamp when the experiment was created status: Current… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/trackio-experiments.textn<1K1 likes76 downloads1y agoHugging Face26Praful932 /abmelt-experiments-exp_20260219_182408textn<1K0 likes58 downloads7mo agoHugging Face27Praful932 /abmelt-experiments-exp_20260218_171120textn<1K0 likes57 downloads7mo agoHugging Face28andstor /peft-unit-test-generation-experiments PEFT Unit Test Generation Experiments Dataset description The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research. Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/andstor/peft-unit-test-generation-experiments.tabularn<1K1 likes54 downloads10mo agoHugging Face29fals3 /peft-unit-test-generation-experiments PEFT Unit Test Generation Experiments Dataset description The PEFT Unit Test Generation Experiments dataset contains metadata and details about a set of trained models used for generating unit tests with parameter-efficient fine-tuning (PEFT) methods. This dataset includes models from multiple namespaces and various sizes, trained with different tuning methods to provide a comprehensive resource for unit test generation research. Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/fals3/peft-unit-test-generation-experiments.tabularn<1K0 likes51 downloads1y agoHugging Face30scholarly360 /contracts-extraction-instruction-llm-experiments Dataset Card for "contracts-extraction-instruction-llm-experiments" More Information needed text1K<n<10K7 likes48 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.