datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
modular-s2orc-parquetmodular_charactersmodular_characters_largemodular_charactersv2modular-sentence-encoders-paraphrase
Multi-parallel paraphrase corpus (23 languages)
The contrastive training data of Modular Sentence Encoders: Separating Language
Specialization from Cross-Lingual Alignment
(ACL 2025). Five English paraphrase datasets, each translated into 22 further
languages, published so that the paper's sentence-encoder and alignment stages
can be reproduced without re-running the translation.
Most of this corpus is machine-translated. Only the English columns are
original human-written text;… See the full description on the dataset page: https://huggingface.co/datasets/yoh/modular-sentence-encoders-paraphrase.modular-sentence-encoders-sts
STS / STR evaluation sets, resolved (23 languages)
The semantic textual similarity and relatedness evaluation data of Modular
Sentence Encoders: Separating Language Specialization from Cross-Lingual
Alignment (ACL 2025), in one flat
schema.
This repository contains no new data. It is a redistribution of six existing
benchmarks, with the alignment work already applied: pairing monolingual sets
into cross-lingual ones by row index, intersecting SICK pair IDs across
translations… See the full description on the dataset page: https://huggingface.co/datasets/yoh/modular-sentence-encoders-sts.SynthCoNL-neardedup
SynthCoNL-neardedup
SynthCoNL-neardedup corpus is a dataset of (comment, code, code) triplets generated starting from CodeSearchNet for the human data.
We then generated the code in a secondary language using Qwen 2.5 Coder-7B-Instruct.
SynthCoNL-neardedup has been used to finetune ModularStarEncoder-finetuned.
This dataset followed the near-deduplication process in ''MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings'', by processing the SynthCoNL raw dataset.… See the full description on the dataset page: https://huggingface.co/datasets/modularStarEncoder/SynthCoNL-neardedup.modular_characters_small_RGBquirky_modularaddition_increment0
Dataset Card for "quirky_modularaddition_increment0"
More Information needed
hotpotqa_four_agents_pipeline-preference_modular_model_prior-bakquirky_modularaddition_increment0_alice
Dataset Card for "quirky_modularaddition_increment0_alice"
More Information needed
quirky_modularaddition_increment0_bob_hard
Dataset Card for "quirky_modularaddition_increment0_bob_hard"
More Information needed
SynthCoNL
SynthCode2Code2NL
SynthCoNL is a dataset of (comment, code, code) triplets generated starting from CodeSearchNet for the human data.
We then generated the code in a secondary language using Qwen 2.5 Coder-7B-Instruct.
This dataset is the non near deduplicated version of SynthCoNL-neardedup
Paper: MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
Languages
Go programming language
Java programming language
Javascript programming language
PHP… See the full description on the dataset page: https://huggingface.co/datasets/modularStarEncoder/SynthCoNL.quirky_modularaddition_rawquirky_modularaddition_increment0_bob
Dataset Card for "quirky_modularaddition_increment0_bob"
More Information needed
quirky_modularaddition_increment0_alice_easy
Dataset Card for "quirky_modularaddition_increment0_alice_easy"
More Information needed
quirky_modularaddition_increment0_alice_hard
Dataset Card for "quirky_modularaddition_increment0_alice_hard"
More Information needed
modulargotoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 70,
"total_frames": 8235,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:70"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ckirk/modulargoto.modular-additionquirky_modularaddition_increment0_bob_easy
Dataset Card for "quirky_modularaddition_increment0_bob_easy"
More Information needed
bankless_What_is_Celestia__TIA__Unpacking_Modular_Blockchainsmodular_characters_hairsmodular_characters_medium_RGBanchor-pairs-ubuntu-modularmodular_characters_hairs_RGBmodular_characters_smallgrading_modularjenny-tts-tags-6hjenny-tts-6h-tagged
