CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01haydn-jones /labbench2-fixed LABBench2 PMID-enriched public mirror This is a public, schema-compatible mirror of EdisonScientific/labbench2, pinned to upstream revision 27d12d72af24e3f70db8a99df63e567366cbdb80. Original columns and source URLs are unchanged. Two columns are added to every configuration: pmids: deduplicated PubMed identifiers resolved for the row's sources. source_pmids: aligned one-to-one with sources; unresolved or non-PubMed sources are null. LABBench2 LABBench2 is a… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/labbench2-fixed.textquestion-answering1K<n<10K0 likes3.1k downloads1mo agoHugging Face02VLM2Vec /ViDoSeek-page-fixedimage1K<n<10K0 likes1.9k downloads11mo agoHugging Face03VLM2Vec /MMLongBench-page-fixedimage1K<n<10K0 likes1.9k downloads11mo agoHugging Face04rajjanardhan00 /Seamless_Dummy_Dataset_Fixed MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes1.5k downloads1y agoHugging Face05aklein4 /proof-pile-2-fixed The original EleutherAI/proof-pile-2 dataset uses a custom python script and .jsonl.zst files, which some versions of the datasets library struggle with. This dataset contains the same data, subsets, and splits as EleutherAI/proof-pile-2, converted into standard parquet format. Each subset and split was also shuffled so that you can directly train on the data without issue. Conversion was performed using the following script: import os importzstandard as zstd import json import pandas as pd… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/proof-pile-2-fixed.texttext-generation10M<n<100M2 likes1.4k downloads8mo agoHugging Face06anthracite-org /nopm_claude_writing_fixedThis is Nopm/Opus_WritingStruct, reuploaded and properly converted to ShareGPT format. text1K<n<10K19 likes909 downloads2y agoHugging Face07galileo-ai /20_Newsgroups_Fixed Dataset Card for 20_Newsgroups_Fixed Dataset Summary This dataset is a version of the 20 Newsgroups dataset fixed with the help of the Galileo ML Data Intelligence Platform. In a matter of minutes, Galileo enabled us to uncover and fix a multitude of errors within the original dataset. In the end, we present this improved dataset as a new standard for natural language experimentation and benchmarking using the Newsgroups dataset. Curation Rationale This… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/20_Newsgroups_Fixed.texttext-classification10K<n<100K3 likes565 downloads4y agoHugging Face08SakethVemula /fixed-tokenizer-morphscore-segmentstabular10M<n<100M0 likes534 downloads6mo agoHugging Face09rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_4license: cc-by-4.0 task_categories: object-detection video-classification tags: biology pretty_name: Seamless_Dummy audion<1K0 likes472 downloads1y agoHugging Face10nz00shuuuu /ltaf-haystack-fixedtabular10K<n<100K0 likes454 downloads5mo agoHugging Face11ReactiveAI /algebraic-stack-fixedtext1M<n<10M0 likes442 downloads9mo agoHugging Face12dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2audio100K<n<1M0 likes435 downloads2mo agoHugging Face13Alignment-Lab-AI /oig-fixedtext1M<n<10M1 likes403 downloads2y agoHugging Face14lapa-llm /hermes3-en-fixed Dataset Card for Hermes 3 Fixed Conversations Dataset Description Dataset Summary hermes3-en-fixed is a [NousResearch/Hermes-3-Dataset]. During preparation we removed all system prompts and normalized the message roles and content to match the common schema we use across our dialog datasets. Languages English (en) Dataset Structure Data Fields conversations: list of messages in a dialog (array of objects) from: normalized sender role — user or assistant… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/hermes3-en-fixed.texttext-generation100K<n<1M0 likes400 downloads11mo agoHugging Face15Humair332 /GLOBE_V2_Fixed A version of the GLOBE dataset that works with load_dataset Important notice Differences between V2 version and the version described in paper: The V2 version provide audio in 44.1kHz sample rate. (Supersampling) The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues. Globe The full paper can be accessed here: arXiv An online demo can be accessed here: Github Abstract This paper introduces GLOBE, a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Humair332/GLOBE_V2_Fixed.audio100K<n<1M0 likes365 downloads10mo agoHugging Face16marcov /super_glue_wsc.fixed_promptsourcetabular1K<n<10K0 likes318 downloads2y agoHugging Face17moca-embed /MMEB-train-fixed MMEB train split used in MoCa Continual Pre-training 🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper Introduction This is a interleaved multimodal pre-training dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from the train split of MMEB by concatenating queries and positive documents. The dataset consists of interleaved multimodal examples. text is a string containing text while images… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/MMEB-train-fixed.text1M<n<10M0 likes313 downloads1y agoHugging Face18SetFit /wsc_fixed Glue WSC Fixed This dataset is a port of the official wsc.fixed dataset on the Hub. Also, the test split is not labeled; the label column values are always -1. tabularn<1K1 likes311 downloads4y agoHugging Face19ysdede /commonvoice_17_tr_fixed Improving CommonVoice 17 Turkish Dataset I recently worked on enhancing the Mozilla CommonVoice 17 Turkish dataset to create a higher quality training set for speech recognition models.Here's an overview of my process and findings. Initial Analysis and Split Organization My first step was analyzing the dataset organization to understand its structure.Through analysis of filename stems as unique keys, I revealed and documented an important aspect of CommonVoice's design… See the full description on the dataset page: https://huggingface.co/datasets/ysdede/commonvoice_17_tr_fixed.audioautomatic-speech-recognition10K<n<100K10 likes311 downloads2y agoHugging Face20batmangiaicuuthegioi /gnl3_fixed_2audio10K<n<100K0 likes298 downloads1y agoHugging Face21rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes297 downloads1y agoHugging Face22avacaondata /sqac_fixedtext10K<n<100K1 likes271 downloads3y agoHugging Face23RedMod /finepdfs_filtered_fixedtext1M<n<10M0 likes255 downloads5mo agoHugging Face24zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes239 downloads1mo agoHugging Face25mrfakename /GLOBE_V2_Fixed A version of the GLOBE dataset that works with load_dataset Important notice Differences between V2 version and the version described in paper: The V2 version provide audio in 44.1kHz sample rate. (Supersampling) The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues. Globe The full paper can be accessed here: arXiv An online demo can be accessed here: Github Abstract This paper introduces GLOBE, a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/GLOBE_V2_Fixed.audio100K<n<1M0 likes233 downloads10mo agoHugging Face26amistele /MTARSI-fixed MTARSI dataset, but actually correctly labeled The Multi-type Aircraft of Remote Sensing Images (MTARSI) original dataset https://zenodo.org/records/3464319 The original MTARSI dataset has a number of aircraft sorted into incorrect categories, and multiple categories completely mislabeled This fixes that issue, by re-defining labels based on the aicraft actually in the dataset, and ensuring all aircraft are actually in their correct categories Removed some images due to poor… See the full description on the dataset page: https://huggingface.co/datasets/amistele/MTARSI-fixed.image1K<n<10K1 likes232 downloads2y agoHugging Face27openmed-community /TheBlueScrubs-v1-fixed openmed-community/TheBlueScrubs-v1-fixed What is this? TheBlueScrubs-v1-fixed is a maintenance fork of the upstream TheBlueScrubs/TheBlueScrubs-v1 train split that resolves a schema bug in the meta column.In the original train files, some rows serialized meta incorrectly (appearing as the literal string "dict"). This fork re-exports the entire train split without meta column, preserving text field and values. Document count: 11,080,331 texts (train) Tokens (upstream… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/TheBlueScrubs-v1-fixed.texttext-generation10M<n<100M13 likes217 downloads1y agoHugging Face28ceselder /lora-text-weight-sonnet5-fixed-a-r1-layer20-3epoch Sonnet source descriptions → fixed-A rank-1 LoRA weights This dataset contains 52,548 aligned examples for raw text-to-LoRA-weight reconstruction with Qwen3-14B. Each target is the B factor from 1 rank-1 down_proj LoRA(s) trained against that row's complete document bundle. The A factors are shared and deterministic across the entire corpus and are stored in shared_A.safetensors. The primary text input is source_description_text, generated from the complete source documents with… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/lora-text-weight-sonnet5-fixed-a-r1-layer20-3epoch.text10K<n<100K0 likes215 downloads2mo agoHugging Face29RAG4Math /targets_fixed_filtered_latextextn<1K0 likes198 downloads4mo agoHugging Face30AgentSuite /DrafterBench-fixed-trajectories AgentSuite/DrafterBench-fixed-trajectories Per-model agent trajectory data for DrafterBench-fixed (public release). Models: 30 Tasks per model: 1,920 One file per model: {model}.jsonl, one JSON object per line. Fields: model_path, user_model_path, benchmark_name, task_name, sampling_params, user_sampling_params, messages, eval_result, meta. sampling_params reflect each benchmark's own implementation; values the benchmark leaves unset are recorded as null (provider default).… See the full description on the dataset page: https://huggingface.co/datasets/AgentSuite/DrafterBench-fixed-trajectories.text10K<n<100K0 likes195 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.