datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
labbench2-fixed
LABBench2 PMID-enriched public mirror
This is a public, schema-compatible mirror of EdisonScientific/labbench2, pinned to upstream revision 27d12d72af24e3f70db8a99df63e567366cbdb80. Original columns and source URLs are unchanged.
Two columns are added to every configuration:
pmids: deduplicated PubMed identifiers resolved for the row's sources.
source_pmids: aligned one-to-one with sources; unresolved or non-PubMed sources are null.
LABBench2
LABBench2 is a… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/labbench2-fixed.ViDoSeek-page-fixedMMLongBench-page-fixedSeamless_Dummy_Dataset_Fixed
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
proof-pile-2-fixed
The original EleutherAI/proof-pile-2 dataset uses a custom python script and .jsonl.zst files, which some versions of the datasets library struggle with.
This dataset contains the same data, subsets, and splits as EleutherAI/proof-pile-2, converted into standard parquet format.
Each subset and split was also shuffled so that you can directly train on the data without issue.
Conversion was performed using the following script:
import os
importzstandard as zstd
import json
import pandas as pd… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/proof-pile-2-fixed.nopm_claude_writing_fixedThis is Nopm/Opus_WritingStruct, reuploaded and properly converted to ShareGPT format.
20_Newsgroups_Fixed
Dataset Card for 20_Newsgroups_Fixed
Dataset Summary
This dataset is a version of the 20 Newsgroups dataset fixed with the help of the Galileo ML Data Intelligence Platform. In a matter of minutes, Galileo enabled us to uncover and fix a multitude of errors within the original dataset. In the end, we present this improved dataset as a new standard for natural language experimentation and benchmarking using the Newsgroups dataset.
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/20_Newsgroups_Fixed.fixed-tokenizer-morphscore-segmentsSeamless_Dummy_Dataset_Fixed_4license: cc-by-4.0
task_categories:
object-detection
video-classification
tags:
biology
pretty_name: Seamless_Dummy
ltaf-haystack-fixedalgebraic-stack-fixedindic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2oig-fixedhermes3-en-fixed
Dataset Card for Hermes 3 Fixed Conversations
Dataset Description
Dataset Summary
hermes3-en-fixed is a [NousResearch/Hermes-3-Dataset]. During preparation we removed all system prompts and normalized the message roles and content to match the common schema we use across our dialog datasets.
Languages
English (en)
Dataset Structure
Data Fields
conversations: list of messages in a dialog (array of objects)
from: normalized sender role — user or assistant… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/hermes3-en-fixed.GLOBE_V2_Fixed
A version of the GLOBE dataset that works with load_dataset
Important notice
Differences between V2 version and the version described in paper:
The V2 version provide audio in 44.1kHz sample rate. (Supersampling)
The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues.
Globe
The full paper can be accessed here: arXiv
An online demo can be accessed here: Github
Abstract
This paper introduces GLOBE, a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Humair332/GLOBE_V2_Fixed.super_glue_wsc.fixed_promptsourceMMEB-train-fixed
MMEB train split used in MoCa Continual Pre-training
🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper
Introduction
This is a interleaved multimodal pre-training dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from the train split of
MMEB by concatenating queries and positive documents.
The dataset consists of interleaved multimodal examples. text is a string containing text while images… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/MMEB-train-fixed.wsc_fixed
Glue WSC Fixed
This dataset is a port of the official wsc.fixed dataset on the Hub.
Also, the test split is not labeled; the label column values are always -1.
commonvoice_17_tr_fixed
Improving CommonVoice 17 Turkish Dataset
I recently worked on enhancing the Mozilla CommonVoice 17 Turkish dataset to create a higher quality training set for speech recognition models.Here's an overview of my process and findings.
Initial Analysis and Split Organization
My first step was analyzing the dataset organization to understand its structure.Through analysis of filename stems as unique keys, I revealed and documented an important aspect of CommonVoice's design… See the full description on the dataset page: https://huggingface.co/datasets/ysdede/commonvoice_17_tr_fixed.gnl3_fixed_2Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
sqac_fixedfinepdfs_filtered_fixedfixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts
fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts
This dataset contains one compressed JSONL shard for every completed training
step. The step and rollout_index columns uniquely locate a rollout within
this training run. Run metadata and per-step row counts are recorded in
rollout_manifest.json.
GLOBE_V2_Fixed
A version of the GLOBE dataset that works with load_dataset
Important notice
Differences between V2 version and the version described in paper:
The V2 version provide audio in 44.1kHz sample rate. (Supersampling)
The V2 versionn removed some samples (~5%) due to the volumn and text aligment issues.
Globe
The full paper can be accessed here: arXiv
An online demo can be accessed here: Github
Abstract
This paper introduces GLOBE, a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/GLOBE_V2_Fixed.MTARSI-fixed
MTARSI dataset, but actually correctly labeled
The Multi-type Aircraft of Remote Sensing Images (MTARSI) original dataset https://zenodo.org/records/3464319
The original MTARSI dataset has a number of aircraft sorted into incorrect categories, and multiple categories completely mislabeled
This fixes that issue, by re-defining labels based on the aicraft actually in the dataset, and ensuring all aircraft are actually in their correct categories
Removed some images due to poor… See the full description on the dataset page: https://huggingface.co/datasets/amistele/MTARSI-fixed.TheBlueScrubs-v1-fixed
openmed-community/TheBlueScrubs-v1-fixed
What is this?
TheBlueScrubs-v1-fixed is a maintenance fork of the upstream TheBlueScrubs/TheBlueScrubs-v1 train split that resolves a schema bug in the meta column.In the original train files, some rows serialized meta incorrectly (appearing as the literal string "dict"). This fork re-exports the entire train split without meta column, preserving text field and values.
Document count: 11,080,331 texts (train)
Tokens (upstream… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/TheBlueScrubs-v1-fixed.lora-text-weight-sonnet5-fixed-a-r1-layer20-3epoch
Sonnet source descriptions → fixed-A rank-1 LoRA weights
This dataset contains 52,548 aligned examples for raw
text-to-LoRA-weight reconstruction with Qwen3-14B. Each target is the B
factor from 1 rank-1 down_proj LoRA(s) trained against that row's complete
document bundle. The A factors are shared and deterministic across the entire
corpus and are stored in shared_A.safetensors.
The primary text input is source_description_text, generated from the complete
source documents with… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/lora-text-weight-sonnet5-fixed-a-r1-layer20-3epoch.targets_fixed_filtered_latexDrafterBench-fixed-trajectories
AgentSuite/DrafterBench-fixed-trajectories
Per-model agent trajectory data for DrafterBench-fixed (public release).
Models: 30
Tasks per model: 1,920
One file per model: {model}.jsonl, one JSON object per line.
Fields: model_path, user_model_path, benchmark_name, task_name, sampling_params, user_sampling_params, messages, eval_result, meta.
sampling_params reflect each benchmark's own implementation; values the benchmark leaves unset are recorded as null (provider default).… See the full description on the dataset page: https://huggingface.co/datasets/AgentSuite/DrafterBench-fixed-trajectories.
