datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
s2orc-cs-enriched
S2ORC CS Enriched
A Computer Science subset of the Semantic Scholar Open Research Corpus (S2ORC) enriched with LLM-generated structured metadata. Contains 1.1 million CS papers with extracted methods, models, datasets, metrics, compute estimates, and summaries.
Dataset Summary
Statistic
Value
Total papers
1,117,706
Total size
54.7 GB
Parquet files
1,118
Split
train
Dataset Structure
Base Columns
Content: parsed_title, abstract… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-cs-enriched.s2orc-safety
S2ORC Safety
This dataset is a filtered and enriched subset of an S2ORC computer science paper corpus, focused on AI safety and adjacent safety-relevant research.
It contains 16,806 papers selected through:
local embedding generation
clustering
GPT-5.4 mini cluster-level screening
GPT-5.4 mini paper-level labeling
a rescue relabel pass on suspicious exclusions
structured metadata extraction over the accepted paper set
filtering out 304 rows that were missing both parsed_title and… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/s2orc-safety.s2orc-academic-papers-augmented-v0
