manifesta/scientific-chart-qa-17k
Scientific Chart QA, 17,070 rows A multimodal chart-interpretation dataset built around one idea: teaching a model when not to answer matters as much as teaching it to answer. One in seven questions here cannot be answered from its figure, and the correct response is cannot be determined. Baseline vision-language models overwhelmingly guess a plausible-looking number instead. That is the behaviour this set targets. The four things worth… See the full description on the dataset page: https://huggingface.co/datasets/manifesta/scientific-chart-qa-17k.
Add public build pipeline + verifier repo, CI badge
Add 'Verify this card', per-source provenance links, published build pipeline
build/README: run order and expected outputs
Publish the build pipeline that produced this dataset
update card: honest results section, decontam evidence, live interface
Upload assets/samples.png with huggingface_hub
Upload assets/task_mix.png with huggingface_hub
Upload assets/banner.png with huggingface_hub
Upload README.md with huggingface_hub
Upload extras/figure_license_manifest.jsonl with huggingface_hub
Upload extras/BLUEPRINT.md with huggingface_hub
Upload extras/build_manifest.json with huggingface_hub
Upload decontam_report.json with huggingface_hub
Upload data/train-00008-of-00009.parquet with huggingface_hub
Upload data/train-00007-of-00009.parquet with huggingface_hub
Upload data/train-00006-of-00009.parquet with huggingface_hub
Upload data/train-00005-of-00009.parquet with huggingface_hub
Upload data/train-00004-of-00009.parquet with huggingface_hub
Upload data/train-00003-of-00009.parquet with huggingface_hub
Upload data/train-00002-of-00009.parquet with huggingface_hub
Upload data/train-00001-of-00009.parquet with huggingface_hub
Upload data/train-00000-of-00009.parquet with huggingface_hub
Upload README.md with huggingface_hub
initial commit
