datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PCMind-2.1-Kaiyuan-2B-phase1-part1-2
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-2.PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323.cl3410-phase1
CL3410 Phase 1 — Malayalam and Assamese language-model corpora
Two independently built pretraining corpora with their own tokenizers:
Malayalam as the higher-resource language and Assamese as the
lower-resource one. Nothing is shared between them — separate sources,
separate cleaning thresholds, separate vocabularies, separate models.
Only the language-agnostic pipeline code is common, parameterised per
language.
Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.browser-agent-phase1-sft-action-only
Browser Agent Phase 1 SFT Action-Only
What this is
Action-only step-level chat SFT data for browser-agent training.
Each example teaches the model to predict the next BrowserGym action from:
the original generation-time system prompt used for data collection
task goal and URL
short recent history
current observation text and diagnostics
Assistant targets contain only the next action.
Why this format
This is the primary training format for small-model SFT… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-action-only.browser-agent-phase1-sft-reasoning-action
Browser Agent Phase 1 SFT Reasoning+Action
What this is
Reasoning-plus-action step-level chat SFT data for browser-agent training.
Each example uses the original generation-time system prompt, then appends a short instruction to reason first and output the final action.
Assistant targets contain:
one <think>...</think> block
then one BrowserGym action
Why this format
This is an experimental variant for comparing whether explicit reasoning supervision helps or… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-reasoning-action.RedPajamas_EN_Phase1
RedPajamas EN Phase1
This dataset contains Phase 1 logic and information-extraction annotations for English RedPajama text chunks.
Each row was generated only when the source document supported all required fields:
3-5 atomic_facts, each with subject, relation, object, supporting context, direct explicit_opposite, and implicit_terms
2-3 potential_unanswerable_entities, each present in the text but missing a specific attribute, constraint, or causal link
Documents that could not… See the full description on the dataset page: https://huggingface.co/datasets/canho/RedPajamas_EN_Phase1.
