datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
modular-s2orc-parquetModularRSI_2000_Instances
ModularRSI 2000 Instances
Dataset summary
This dataset is the independent evolution pool released with ModularRSI. It
contains Harbor-formatted tasks for studying whether an agent can improve its
harness from execution experience and transfer those improvements to unseen
tasks.
Companion code: IQuestLab/ModularRSI
Domain
Instances
train
pool
Terminal tasks
1,000
120
880
Software-engineering tasks
1,000
120
880
The 120-task training subset is not… See the full description on the dataset page: https://huggingface.co/datasets/IQuestLab/ModularRSI_2000_Instances.modular-s2orc
Dataset Card for Modular S2ORC
Topically and temporally partitioned data from S2ORC, used for continued pre-training experiments in "Scalable Data Ablation Approximations for Language Models through Modular Training and Merging" to be presented at EMNLP 2024.
Validation and test splits are determined by 'sha1' values in metadata.
More details to come.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/claran/modular-s2orc.modular_charactersmodular-diffusers-blogmodular_characters_largeModularNodeIntegrator_119_4modular_charactersv2ModularCacheBridge_714_8CQmodular-pretrainingmodular-s2orc
Dataset Card for Modular S2ORC
Topically and temporally partitioned data from S2ORC, used for continued pre-training experiments in "Scalable Data Ablation Approximations for Language Models through Modular Training and Merging" to be presented at EMNLP 2024.
Validation and test splits are determined by 'sha1' values in metadata.
More details to come.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/usersaico/modular-s2orc.modular-sentence-encoders-paraphrase
Multi-parallel paraphrase corpus (23 languages)
The contrastive training data of Modular Sentence Encoders: Separating Language
Specialization from Cross-Lingual Alignment
(ACL 2025). Five English paraphrase datasets, each translated into 22 further
languages, published so that the paper's sentence-encoder and alignment stages
can be reproduced without re-running the translation.
Most of this corpus is machine-translated. Only the English columns are
original human-written text;… See the full description on the dataset page: https://huggingface.co/datasets/yoh/modular-sentence-encoders-paraphrase.friendship-graph-modular-edge-irregularity-proof
Defect Conservation and Exact Modular Edge-Irregularity Strength of Friendship Graphs
Public AI-friendly research release · candidate proof · independently verifiable artifacts
This repository contains a complete candidate resolution of Open Problem 3.3 from Koam, Ahmad, Bača, and Semaničová-Feňovčíková, AIMS Mathematics 8(1), 2023, concerning the modular edge irregularity strength of friendship graphs.
Main candidate theorem
For the friendship graph (F_n=K_1\vee… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/friendship-graph-modular-edge-irregularity-proof.modular-sentence-encoders-sts
STS / STR evaluation sets, resolved (23 languages)
The semantic textual similarity and relatedness evaluation data of Modular
Sentence Encoders: Separating Language Specialization from Cross-Lingual
Alignment (ACL 2025), in one flat
schema.
This repository contains no new data. It is a redistribution of six existing
benchmarks, with the alignment work already applied: pairing monolingual sets
into cross-lingual ones by row index, intersecting SICK pair IDs across
translations… See the full description on the dataset page: https://huggingface.co/datasets/yoh/modular-sentence-encoders-sts.LLM_Modularity_attribution
Per-task neuron attribution scores for Modular Cognitive Architecture Emerges in Large Language Models
Code and analysis pipeline: https://github.com/Pengrui-Han/LLM_Modularity
This dataset contains the raw attribution-patching tensors that the GitHub
release omits (they are ~6 GB). With these files you can run every
overlap / ablation / statistics script in the repo without re-running
attribution patching on a GPU.
Layout
results/<model>/<domain>/<task>/… See the full description on the dataset page: https://huggingface.co/datasets/barryhpr/LLM_Modularity_attribution.repro-gram-modular-pretraining-traces
Agent traces
Agent sessions published from a Trackio Logbook.
SynthCoNL-neardedup
SynthCoNL-neardedup
SynthCoNL-neardedup corpus is a dataset of (comment, code, code) triplets generated starting from CodeSearchNet for the human data.
We then generated the code in a secondary language using Qwen 2.5 Coder-7B-Instruct.
SynthCoNL-neardedup has been used to finetune ModularStarEncoder-finetuned.
This dataset followed the near-deduplication process in ''MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings'', by processing the SynthCoNL raw dataset.… See the full description on the dataset page: https://huggingface.co/datasets/modularStarEncoder/SynthCoNL-neardedup.quirky_modularaddition_increment0
Dataset Card for "quirky_modularaddition_increment0"
More Information needed
ModularFusionBridge_98_9MHmodular_characters_small_RGBmodular-additionquirky_modularaddition_increment0_alice
Dataset Card for "quirky_modularaddition_increment0_alice"
More Information needed
quirky_modularaddition_increment0_bob_hard
Dataset Card for "quirky_modularaddition_increment0_bob_hard"
More Information needed
quirky_modularaddition_rawquirky_modularaddition_increment0_bob
Dataset Card for "quirky_modularaddition_increment0_bob"
More Information needed
hotpotqa_four_agents_pipeline-preference_modular_model_prior-bakSynthCoNL
SynthCode2Code2NL
SynthCoNL is a dataset of (comment, code, code) triplets generated starting from CodeSearchNet for the human data.
We then generated the code in a secondary language using Qwen 2.5 Coder-7B-Instruct.
This dataset is the non near deduplicated version of SynthCoNL-neardedup
Paper: MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
Languages
Go programming language
Java programming language
Javascript programming language
PHP… See the full description on the dataset page: https://huggingface.co/datasets/modularStarEncoder/SynthCoNL.quirky_modularaddition_increment0_alice_easy
Dataset Card for "quirky_modularaddition_increment0_alice_easy"
More Information needed
quirky_modularaddition_increment0_alice_hard
Dataset Card for "quirky_modularaddition_increment0_alice_hard"
More Information needed
why-minimalist-modular-kitchens-are-replacing-traditional-designs
