datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.2026-08-17-table2-9284-peer-critique-good-716-train-mixture
Qwen3.6-27B SFT mixture: 9,284 Table2 + 716 peer_critique GOOD ARM (10,000 rows)
The one-variable twin of LASR-Callum/2026-08-16-table2-9284-peer-critique-716-train, whose
716 peer-critique rows are 358 good / 358 flawed. Here all 716 are drawn from the good arm.
field
value
experiment
Arm ablation: does the peer-critique FLAWED arm contribute anything? Train on good-arm-only critiques and compare against the 358/358 arm.
date_generated
2026-08-17
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-peer-critique-good-716-train-mixture.gene-r1-go-sft-table2-reconstructed
Gene-R1 GO SFT Table 2 Reconstructed Splits
Small workshop dataset used for tokenizer-transfer experiments with ncbi/Gene-R1-1B.
Rows are reconstructed from released Gene Ontology benchmark materials into Gene-R1-style prompt/completion text for tokenizer transfer experiments:
train.jsonl: 2400 rows, 800 BP + 800 MF + 800 CC
validation.jsonl: 300 rows, 100 BP + 100 MF + 100 CC
test.jsonl: 300 rows, 100 BP + 100 MF + 100 CC
Each row contains metadata plus prompt, completion… See the full description on the dataset page: https://huggingface.co/datasets/transhumanist-already-exists/gene-r1-go-sft-table2-reconstructed.md-2-xml-wiki-tables
md-2-xml-wiki-tables
958 markdown tables extracted from fan/community MediaWiki sites for markdown-to-XML format conversion tasks.
Format
JSONL with fields:
title: article title from the source wiki page
section: section heading the table appeared under
wiki: source wiki name
table_md: raw markdown table
filename: original filename
Splits
train: 894 tables
eval: 64 held-out tables
Source
Various fan/community MediaWiki sites. Most use CC-BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/md-2-xml-wiki-tables.table-sft-eval-predictions
💾 Raw Predictions for "What Really Matters for Table LLMs?"
This dataset contains the raw model outputs from the experiments in:
Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang,
Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng.
What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects.
Findings of EACL 2026. https://aclanthology.org/2026.findings-eacl.195/
🗂️ Layout… See the full description on the dataset page: https://huggingface.co/datasets/dnaihao/table-sft-eval-predictions.relationalrag-tables-artifact-data
RelationalRAG-Tables Artifact Data
This public dataset repository mirrors the lightweight data payloads from the
RelationalRAG-Tables reproducibility artifact:
fixtures/: smoke and synthetic fixtures used by the CPU-only artifact path.
artifacts/: stored JSON outputs and manifest.json hashes for reported
experimental numbers.
The repository does not redistribute heavyweight upstream benchmark corpora
such as BIRD or HybridQA, nor does it redistribute pretrained model weights.… See the full description on the dataset page: https://huggingface.co/datasets/lexuanbach/relationalrag-tables-artifact-data.
