lineage
Datasets
All datasets matching “lineage”real01b-routing-d1-umirel-lineage-15arm-heldout-sobol50-s2026090701This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-routing-d1-umirel-lineage-15arm-heldout-sobol50-s2026090701.tabrepairbench-replayable-corruption-lineage
TabRepairBench: Replayable Corruption Lineage
This is a finite, wholly authored synthetic reference dataset for auditing
tabular corruption lineage. Public Lineage v1 contains 72 independent groups
and 6,912 clean/corrupt cell pairs across three authored structural generator
families, six corruption schedules, two severities, and two audit partitions.
It is intentionally not presented as real-world data. It makes no claim of
real-data representativeness, causal identification… See the full description on the dataset page: https://huggingface.co/datasets/haidang2405/tabrepairbench-replayable-corruption-lineage.real01c-insert-marker-d1-r2-stage2-lineage-comparison-sobol10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 50,
"total_frames": 15063,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01c-insert-marker-d1-r2-stage2-lineage-comparison-sobol10.LineageFlow-assets
LineageFlow Pfam Assets
This dataset contains the preprocessed Pfam assets used by the released LineageFlow inference pipeline.
Contents
pfam_priors_asr_mad/: family-specific ASR Dirichlet priors.
pfam_gap_rates/: family-specific alignment gap statistics.
pfam_fastas_clean/: cleaned Pfam family alignments.
pfam_pi_smooth_tau0.5_gap060_gt80_020.csv: family sampling distribution.
pfam_priors_keep_ids_gap060_gt80_020.txt: family keep list used by the default sampler.… See the full description on the dataset page: https://huggingface.co/datasets/jinxbye/LineageFlow-assets.weft-script-lineage-synth
weft-script-lineage — synthetic training data + negative-result evidence
Companion data for the model wallfacers/weft-lineage-extractor-1.5b,
a research artifact demonstrating that synthetic-only training induces a memorization
leak in small models for ETL table-lineage extraction.
Contents
train.jsonl / heldout.jsonl — 10,000 + 600 synthetic ETL scripts with table-lineage
labels (Python/Shell). out-jvm/ adds the Scala/Java-augmented variant.
reports/ — the… See the full description on the dataset page: https://huggingface.co/datasets/wallfacers/weft-script-lineage-synth.lineagebench
LineageBench
A small, honestly-scoped benchmark for weight-based lineage verification:
given a suspect open-weight model and a set of candidate parents, decide which
foundation base it actually descends from — from the weights alone.
The hard part of a lineage benchmark is the ground truth. It cannot come from a
model's own base_model tag, because auditing that tag is the entire point of
the exercise. LineageBench instead labels every pair from the publishing
organization's own… See the full description on the dataset page: https://huggingface.co/datasets/AwaisAdilKhokhar/lineagebench.
