datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CROHME-fullCROHME-full-clone-fixedCROHME_RS_think
CROHME — CROHME_RS_think
Rejection-sampled from the CROHME train split. This split holds the accepted items, with the model's reasoning trace.
rows
4,233
QA pairs
4,233
shards
11
accepted / rejected (whole family)
4,233 / 2,436
accept rate
63.5%
verifier
latex
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier described… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/CROHME_RS_think.CROHME_rejected
CROHME — CROHME_rejected
Rejection-sampled from the CROHME train split. This split holds the rejected items — the answer field holds the official ground truth.
rows
2,436
QA pairs
2,436
shards
7
accepted / rejected (whole family)
4,233 / 2,436
accept rate
63.5%
verifier
latex
The rejected split is training data, not just diagnostics: answer is the official ground truth, and wrong_vlm records what the model said instead.
How the data was… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/CROHME_rejected.CROHME_RS_nothink
CROHME — CROHME_RS_nothink
Rejection-sampled from the CROHME train split. This split holds the accepted items, answer only.
rows
4,233
QA pairs
4,233
shards
11
accepted / rejected (whole family)
4,233 / 2,436
accept rate
63.5%
verifier
latex
How the data was produced
A VLM answers every question at temperature 0 with reasoning enabled. Its answer is compared with
the official ground truth by the verifier described below; matches go to… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/CROHME_RS_nothink.crohme_train_cleaned
crohme_train_cleaned
The crohme_train family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
8,708
QA turns
22,610
answers rewritten by the cleaning pass
790
QA created by the cleaning pass (new_qa)
15,140 (67.0%)
shards
2
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong but… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/crohme_train_cleaned.test2_CROHME2014test2_CROHME2019test2_CROHME2023-offlinetest2_CROHME2016test2_CROHME2023-onlineCROHME_longformCROHME-full-cloneCROHME-chat-format
