datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gene-r1-go-sft-table2-reconstructed
Gene-R1 GO SFT Table 2 Reconstructed Splits
Small workshop dataset used for tokenizer-transfer experiments with ncbi/Gene-R1-1B.
Rows are reconstructed from released Gene Ontology benchmark materials into Gene-R1-style prompt/completion text for tokenizer transfer experiments:
train.jsonl: 2400 rows, 800 BP + 800 MF + 800 CC
validation.jsonl: 300 rows, 100 BP + 100 MF + 100 CC
test.jsonl: 300 rows, 100 BP + 100 MF + 100 CC
Each row contains metadata plus prompt, completion… See the full description on the dataset page: https://huggingface.co/datasets/transhumanist-already-exists/gene-r1-go-sft-table2-reconstructed.FinTagging1000_table
FinTagging Table Context Extraction Split
For HTML table inputs, the target is a JSON list of numeric entity/datatype pairs enriched with deterministic row and column context. Missing row or column context is represented as null.
The dataset is derived from FinTagging_800_200_HF and preserves the original
train/test assignment by source_sample_idx and context_id. The XBRL concept
tag is intentionally omitted from the target.
Splits
Split
Samples
Output… See the full description on the dataset page: https://huggingface.co/datasets/lm2445/FinTagging1000_table.
