markdown-table
GreenNode-Table-Markdown-Retrieval-VN
GreenNodeTableMarkdownRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
GreenNodeTable documents
Task category
t2t
Domains
Financial, Encyclopaedic, Non-fiction
Reference
https://huggingface.co/GreenNode
Source datasets:
GreenNode/GreenNode-Table-Markdown-Retrieval-VN
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("GreenNodeTableMarkdownRetrieval")… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/GreenNode-Table-Markdown-Retrieval-VN.markdown-table-qa
Markdown Table QA Dataset
A synthetic dataset of 11,000 (instruction, input, response) triples (10,000 train + 1,000 validation) for training and evaluating language models on structured table understanding and computational reasoning.
What's in it
Each sample contains a markdown table paired with a natural language question and a conversational answer:
Field
Description
instruction
Natural language question about the table
input
The markdown table… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-qa.Viet-Table-Markdownmarkdown-table-qa-19
Markdown Table QA Dataset — Part 19/20
Part 19 of a 20-dataset collection for training and evaluating language models on structured table understanding and computational reasoning. Each part contains 2,200 samples (2,000 train + 200 validation) with step-by-step reasoning traces.
See the full collection: cetusian/markdown-table-qa-01 through cetusian/markdown-table-qa-20
Parent dataset: cetusian/markdown-table-qa (11,000 samples)
What's in it
Each sample contains a… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-qa-19.markdown-table-qa-01
Markdown Table QA Dataset — Part 01/20
Part 1 of a 20-dataset collection for training and evaluating language models on structured table understanding and computational reasoning. Each part contains 2,200 samples (2,000 train + 200 validation) with step-by-step reasoning traces.
See the full collection: cetusian/markdown-table-qa-01 through cetusian/markdown-table-qa-20
Parent dataset: cetusian/markdown-table-qa (11,000 samples)
What's in it
Each sample contains a… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-qa-01.markdown-table-qa-05
Markdown Table QA Dataset — Part 05/20
Part 5 of a 20-dataset collection for training and evaluating language models on structured table understanding and computational reasoning. Each part contains 2,200 samples (2,000 train + 200 validation) with step-by-step reasoning traces.
See the full collection: cetusian/markdown-table-qa-01 through cetusian/markdown-table-qa-20
Parent dataset: cetusian/markdown-table-qa (11,000 samples)
What's in it
Each sample contains a… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/markdown-table-qa-05.
