datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
temppgc-ptsd
PGC Post-Traumatic Stress Disorder — GWAS Summary Statistics
Dataset Description
Genome-wide association study (GWAS) summary statistics for Post-Traumatic Stress Disorder phenotypes from the Psychiatric Genomics Consortium (PGC).
Usage
from datasets import load_dataset
ds = load_dataset("OpenMed/pgc-ptsd", "maltreatment2020")
print(ds)
Subsets
Config
Phenotype
Journal
Year
PubMed
Rows
maltreatment2020
Childhood… See the full description on the dataset page: https://huggingface.co/datasets/tunahanf/pgc-ptsd.bigtom_trainpgc-ptsd
PGC Post-Traumatic Stress Disorder — GWAS Summary Statistics
Dataset Description
Genome-wide association study (GWAS) summary statistics for Post-Traumatic Stress Disorder phenotypes from the Psychiatric Genomics Consortium (PGC).
Usage
from datasets import load_dataset
ds = load_dataset("OpenMed/pgc-ptsd", "maltreatment2020")
print(ds)
Subsets
Config
Phenotype
Journal
Year
PubMed
Rows
maltreatment2020
Childhood Maltreatment… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/pgc-ptsd.Maths-Grade-SchoolMaths-Grade-School
I am releasing large Grade School level Mathematics datatset.
This extensive dataset, comprising nearly one million instructions in JSON format, encapsulates a diverse array of topics fundamental to building a strong mathematical foundation.
This dataset is in instruction format so that model developers, researchers etc. can easily use this dataset.
Following Fields & sub Fields are covered:
Calculus
Probability
Algebra
Liner Algebra
Trigonometry
Differential Equations… See the full description on the dataset page: https://huggingface.co/datasets/pt-sk/Maths-Grade-School.Vision-COTsynth_docs_pts_10kpretraining-datasetTokens are prepared using TikToken's "cl100k_base" tokenizer, which is used for GPT3.5 and GPT4.
Aria_Dataset-UnsortedQwen3-0.6B-pts
Qwen/Qwen3-0.6B — Pivotal Token Search
Pivotal reasoning events for Qwen/Qwen3-0.6B, at three representational scales in one
file, produced with PTS.
latent meta-token / workspace event (Latent PTS) ← J-lens readout
↓
emitted pivotal token (Token PTS) ← Phi-4 PTS
↓
sentence-level thought anchor (Sentence PTS) ← Thought Anchors
↓
success / failure probability shift
All three are CausalReasoningEvent records — one schema… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts.PTSyn_v1fineweb_edu_10BThis is the finewebedu dataset 10 Billion tokens. Tokenized using gpt-2 tokenizer. Each shards are a numpy file and contains 100 million tokens.
toxic_classificationcombination of SetFit/toxic_conversations_50k, Arsive/toxicity_classification_jigsaw
pts-taigi
PTS Taigi
Collection of Taiwanese-language YouTube talk-show/documentary content, staged from COS
ahead of a full transfer. Each COS source gets its own config (schemas differ, so they
can't share one), each with one split named after the source folder.
Configs / Splits
ptv_ts_ds_test_zhtw (367 rows) -- schema normalized to this project's conventions
(see rule.md): tw/zh were split into clean text/mandarin plus
text_timestamped/mandarin_timestamped (original… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/pts-taigi.Qwen3-0.6B-pts-steering-vectors
PTS Steering Vectors Dataset
A dataset of activation-based steering vectors created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Dataset Structure
This dataset contains:
steering_vectors.jsonl: The main file with token-level steering vectors
Usage
These steering vectors can be used for activation-based steering during inference to guide language models toward particular… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-steering-vectors.Face-Emotion-DetectionQwen3-0.6B-pts-thought-anchors
PTS Thought Anchors Dataset
A dataset of thought anchors - critical reasoning steps - identified using the Thought Anchors technique from the PTS tool.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Tags: pts, thought-anchors, reasoning, llm-analysis
Dataset Structure
This dataset contains thought anchors identified from reasoning traces. Each anchor represents a sentence that significantly impacts the success probability of the reasoning… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-thought-anchors.PTSyn_v2DeepSeek-R1-Distill-Qwen-1.5B-pts-thought-anchors
PTS Thought Anchors Dataset
A dataset of thought anchors - critical reasoning steps - identified using the Thought Anchors technique from the PTS tool.
Details
Source: Generated using the PTS tool
Model: deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
Tags: pts, thought-anchors, reasoning, llm-analysis
Dataset Structure
This dataset contains thought anchors identified from reasoning traces. Each anchor represents a sentence that significantly impacts the success… See the full description on the dataset page: https://huggingface.co/datasets/codelion/DeepSeek-R1-Distill-Qwen-1.5B-pts-thought-anchors.tinystories_upsampled_tom_500kptsft150
Model Card for Model ID
Model Details
Model Description
Developed by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Model type: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/ptsft150.tinystories_baselineDeepSeek-R1-Distill-Qwen-1.5B-pts
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B — Pivotal Token Search
Pivotal reasoning events for deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B, at three representational scales in one
file, produced with PTS.
latent meta-token / workspace event (Latent PTS) ← J-lens readout
↓
emitted pivotal token (Token PTS) ← Phi-4 PTS
↓
sentence-level thought anchor (Sentence PTS) ← Thought Anchors
↓
success / failure probability shift
All… See the full description on the dataset page: https://huggingface.co/datasets/codelion/DeepSeek-R1-Distill-Qwen-1.5B-pts.zh-tw-pts-articles-sm
zh-tw-pts-articles-sm
🐣English • 🇹🇼 繁體中文
This dataset contains articles scraped from PNN News.
It's a news provider verified by the vast majority.
Note: some keys like conclusion may be None.
Dataset({
features: ['image', 'title', 'conclusion', 'content', 'timestamp', 'category', 'link'],
num_rows: 1400
})
Use The Dataset
Use 🤗 Datasets to download, use or modify this dataset.
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/AWeirdDev/zh-tw-pts-articles-sm.PTSyn_v1_valresearch_papers_short
Dataset Card
This is a dataset containing ML ArXiv papers. The dataset is a version of the original one from CShorten, which is a part of the ArXiv papers dataset from Kaggle.
Three steps are made to process the source data:
useless columns removal;
train-test split;
'\n' removal and trimming spaces on sides of the text.
Qwen3-0.6B-pts-dpo-pairs
PTS DPO Dataset
A Direct Preference Optimization (DPO) dataset created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Format
Each example in the dataset consists of:
prompt: The context leading up to the pivotal token
chosen: The preferred token that increases success probability
rejected: The alternative token that decreases success probability
metadata: Additional information about the… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-dpo-pairs.DeepSeek-R1-Distill-Qwen-1.5B-pts-dpo-pairs
PTS DPO Dataset
A Direct Preference Optimization (DPO) dataset created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
Format
Each example in the dataset consists of:
prompt: The context leading up to the pivotal token
chosen: The preferred token that increases success probability
rejected: The alternative token that decreases success probability
metadata:… See the full description on the dataset page: https://huggingface.co/datasets/codelion/DeepSeek-R1-Distill-Qwen-1.5B-pts-dpo-pairs.dino_touch_and_go_3_ptsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 90,
"total_frames": 26782,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sach088/dino_touch_and_go_3_pts.ptsft50
Model Card for Model ID
Model Details
Model Description
Developed by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Model type: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Finetuned from model [optional]: [More Information Needed]
Model Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/cjfcsjt/ptsft50.
