paired-data
Paired_Boussinesq_Compressible_Dataset
Paired Boussinesq / Compressible Natural Convection — 10,000 Simulations (version 2)
📄 Paper: A Neural Surrogate Approach for Simulating Natural Convection
Problems (arXiv:2606.25259) — Nurshat Menglik,
Alex Shao, David Hyde.
Version 2 (September 2026) replaces the original release. All 10,000 pairs were
regenerated; see What changed in version 2. The original
files remain available unchanged under the tag v1.0-legacy
(snapshot_download(..., revision="v1.0-legacy")). New work… See the full description on the dataset page: https://huggingface.co/datasets/NurshatMenglik/Paired_Boussinesq_Compressible_Dataset.oas-paired-sequence-data
Dataset Card for OAS Paired Sequence Data
Dataset Summary
Paired heavy- and light-chain sequence information from the Observed Antibody Space (OAS) database, downloaded on September 9, 2023.
lab_data_paired_64This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 60,
"total_frames": 19056,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ceilingfan456/lab_data_paired_64.lab_data_orange_cube_single_point_paired_25This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 50,
"total_frames": 12001,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ceilingfan456/lab_data_orange_cube_single_point_paired_25.soda-vec-data-full_pmc_title_abstract_paired
SODA-VEC Paired Dataset for Negative Sampling
This is a paired version of the SODA-VEC dataset, specifically formatted for negative sampling training with MultipleNegativesRankingLoss.
Dataset Overview
Total examples: 26,573,900
Format: Paired (anchor-positive) for contrastive learning
Source: EMBO/soda-vec-data-full_pmc_title_abstract
Purpose: Training sentence transformers with negative sampling
Data Format
Each example contains:
anchor (string): The title… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract_paired.uiclip_human_data-paired_hf
