CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TableSenseAI /TabMWPtext1K<n<10K0 likes2.8k downloads1y agoHugging Face02Arietem /tabmwptabulartable-question-answering1K<n<10K2 likes763 downloads2y agoHugging Face03bevaya /TABMEpp Dataset Card for TABME++ The TABME dataset is a synthetic collection of business document folders generated from the Truth Tobacco Industry Documents archive, with preprocessing and OCR results included, designed to simulate real-world digitization tasks. TABME++ extends TABME by enriching it with commercial-quality OCR (Microsoft OCR). Dataset Details Dataset Description The TABME dataset is a synthetic collection created to simulate the… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/TABMEpp.texttext-classification100K<n<1M5 likes348 downloads2y agoHugging Face04sionic-ai /tabmwpimage10K<n<100K0 likes310 downloads1y agoHugging Face05TabMaven /arxiv-rawtext1M<n<10M0 likes235 downloads1y agoHugging Face06TableSenseAI /TabMWPSelectionThis dataset is a high-fidelity selection from the Tabular Math Word Problems (TabMWP) benchmark (Lu et al., 2023). TabMWP is a leading resource for evaluating mathematical reasoning over heterogeneous tabular and textual data. To address potential noise and ensure the highest standards of logical grounding, this curated version consists of 100 hand-verified examples. Each entry has been audited to confirm that the multi-step reasoning chains—including information look-up and numerical… See the full description on the dataset page: https://huggingface.co/datasets/TableSenseAI/TabMWPSelection.textquestion-answeringn<1K0 likes189 downloads7mo agoHugging Face07germane /Tab-MIA Tab-MIA: A Benchmark for Membership Inference Attacks on Tabular Data Tab-MIA is a benchmark dataset designed to evaluate the privacy risks of fine-tuning large language models (LLMs) on structured tabular data. It enables reproducible and systematic testing of Membership Inference Attacks (MIAs) across diverse datasets and six different serialization formats. 📋 Overview Datasets: WTQ (WikiTableQuestions) WikiSQL TabFact Adult Census California Housing… See the full description on the dataset page: https://huggingface.co/datasets/germane/Tab-MIA.texttext-classification100K<n<1M0 likes186 downloads1y agoHugging Face08theblackcat102 /tabmwp-cleantabular10K<n<100K0 likes134 downloads6mo agoHugging Face09JiaerX /TabMWPimage10K<n<100K0 likes100 downloads1y agoHugging Face10nimapourjafar /mm_tabmwpimage10K<n<100K3 likes70 downloads2y agoHugging Face11EvalData /TabMI-Bench TabMI-Bench A protocol benchmark for mechanistic interpretability (MI) of tabular foundation models (TFMs). NeurIPS 2026 Evaluations & Datasets Track submission. What's in this dataset This Hugging Face repository hosts the frozen aggregated artifacts that drive every numbered table and figure in the paper. Bundling these allows reviewers to verify the paper's key numerics without re-running 40 GPU-hours of experiments. File Source experiment Used by… See the full description on the dataset page: https://huggingface.co/datasets/EvalData/TabMI-Bench.tabularn<1K0 likes64 downloads5mo agoHugging Face12pingzhili /tabmwptabular1K<n<10K0 likes60 downloads1y agoHugging Face13Raywithyou /TabMCQtext1K<n<10K0 likes55 downloads1y agoHugging Face14lbourdois /VQA-tabmwp Description This dataset is a processed version of the TabMWP dataset by Lu et al.We converted the images to PIL and translated question and answers from Englist to French. Citation TabMWP @misc{lu2023dynamicpromptlearningpolicy, title={Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning}, author={Pan Lu and Liang Qiu and Kai-Wei Chang and Ying Nian Wu and Song-Chun Zhu and Tanmay Rajpurohit and Peter Clark and… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/VQA-tabmwp.imagevisual-question-answering10K<n<100K0 likes48 downloads8mo agoHugging Face15alckasoc /tabmwp_expel_train_100tabularn<1K0 likes25 downloads2y agoHugging Face16Sing0402 /tabmwp_200textn<1K0 likes24 downloads2y agoHugging Face17theblackcat102 /tabmwp-hard-verifiedtabularn<1K0 likes22 downloads5mo agoHugging Face18elliot-mllm /tabmwp_cleanedgated tabmwp_cleaned The tabmwp__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning. images 24,966 QA turns 143,078 answers rewritten by the cleaning pass 6,128 QA created by the cleaning pass (new_qa) 93,315 (65.2%) shards 1 How this was cleaned A vision-language model read each image together with its QA and judged the item. The pass is not a filter that only removes rows — it rewrites answers it finds wrong but… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/tabmwp_cleaned.imagevisual-question-answering10K<n<100K0 likes22 downloads23d agoHugging Face19TabMaven /fineweb-sample-10BT-completion-augmented-v0tabular10K<n<100K0 likes21 downloads1y agoHugging Face20TabMaven /all-the-newstabular1M<n<10M0 likes20 downloads1y agoHugging Face21TabMaven /s2orc-academic-papers-augmented-v0tabularn<1K0 likes20 downloads1y agoHugging Face22alckasoc /tabmwp_200tabularn<1K0 likes18 downloads1y agoHugging Face23TabMaven /fineweb-sample-10BT-augmented-v0tabular1K<n<10K0 likes16 downloads1y agoHugging Face24alckasoc /tabmwp_200_traintabularn<1K0 likes15 downloads1y agoHugging Face25TabMaven /all_the_news_augmented-v0tabular1K<n<10K0 likes14 downloads1y agoHugging Face26TabMaven /tabmaven-270925tabular1K<n<10K0 likes8 downloads1y agoHugging Face27TabMaven /enron-mail-dataset-rawtext100K<n<1M0 likes8 downloads1y agoHugging Face28Hypodossier /tabme_smallimagen<1K0 likes7 downloads2y agoHugging Face29TabMaven /finepdfs-eng_Latn-augmented-v0tabular1K<n<10K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.