join
Datasets
All datasets matching “join”the-join
The Join
A broad collection of relational databases spanning many domains (academic,
e-commerce, finance, sports, biomedical, government, text2sql, and more), ported to the
RelBench manifest format. The Join is built for pretraining relational/tabular foundation
models: each database is self-describing and tasks ship labels as-is for large-scale
pretraining rather than held-out benchmarking.
Each dataset lives in its own subdirectory in the self-describing manifest layout (plain… See the full description on the dataset page: https://huggingface.co/datasets/stanford-star/the-join.0XoLemon
Localization Pipeline Assets
Purpose
Repository for storing assets, intermediate outputs, dictionaries, and binaries used in game localization workflows.
Included Data
extracted resources
translation cache
intermediate processed files
rebuild artifacts
tool dependencies
Intended Use
This repository is used for:
reproducible localization workflows
asset synchronization across machines
versioned storage of large artifacts… See the full description on the dataset page: https://huggingface.co/datasets/JOINCANE/0XoLemon.atlas-35-joint-math-and-code-training-data
35. One training set of mathematics and code: OpenMathReasoning and OpenCodeReasoning-2 together
1. Question and links
Read this first. The reading copy of this directory is t2ance/atlas-experiments under 35-joint-math-and-code-training-data/; the saved training steps and the per-token training arrays are on the Hugging Face repository t2ance/atlas-35-joint-math-and-code-training-data only.
Can OpenCodeReasoning-2 (OCR-2) give code rows with the properties… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-35-joint-math-and-code-training-data.icl-dataset-joint-space
icl-dataset-fixed-action
Derived from adityx23/icl-dataset
(lerobot v2.1 format). Every existing column, task, episode flag
(success/valid/keep), and episode_uid is carried through unchanged.
What's added
One new feature, action.q_target (float32, shape [14], names
lj0..lj6, rj0..rj6): the joint-space reconstruction of each frame's
action.left_ee / action.right_ee cartesian targets, via the mink-based
IK procedure documented in vr_teleop_ik.md (available in the… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/icl-dataset-joint-space.texas-42-joint-world-corpus-v2
Texas 42 Joint-World Corpus v2 — the clean deck
Regeneration of
texas-42-joint-world-corpus
on the repaired world sampler, with write-time validity assertions and
per-world posterior weights. Same schema, same teacher, same seed plan
(minus declaration 8 — see below).
Why a v2 repo
The original corpus (April 2026) was generated with a world sampler whose
no-candidate branch injected domino 0-0 instead of rejecting, so 27–67% of
stored worlds per decision are not… See the full description on the dataset page: https://huggingface.co/datasets/jasonyandell/texas-42-joint-world-corpus-v2.the-join-preprocessed
