CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01riotu-lab /SARD SARD: Synthetic Arabic Recognition Dataset Overview SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts. Key Features… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/SARD.image-to-text100K<n<1M14 likes33k downloads4mo agoHugging Face02sardinelab /MF2tabularvisual-question-answeringn<1K9 likes670 downloads1y agoHugging Face03vrinda2712 /SARD SARD: Synthetic Arabic Recognition Dataset Overview SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts. Key Features Massive… See the full description on the dataset page: https://huggingface.co/datasets/vrinda2712/SARD.image-to-text100K<n<1M0 likes442 downloads8mo agoHugging Face04caoxuhao /SARD SARD: Synthetic Arabic Recognition Dataset Overview SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts. Key Features Massive… See the full description on the dataset page: https://huggingface.co/datasets/caoxuhao/SARD.image-to-text100K<n<1M0 likes295 downloads4mo agoHugging Face05riotu-lab /SARD-Extended SARD: Synthetic Arabic Recognition Dataset Overview SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts. Key Features Massive… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/SARD-Extended.image-to-text9 likes290 downloads4mo agoHugging Face06sardinelab /DocBlocks Dataset Card for DocBlocks DocBlocks is a high-quality, multilingual document-level machine translation (MT) dataset designed to fine-tune large language models (LLMs) on long-context translation tasks. Unlike traditional sentence-level datasets, it contains full documents with natural discourse structures and contextual alignment, helping models maintain coherence, consistency, and high translation quality across longer texts. Curated by: Instituto Superior Técnico, Instituto de… See the full description on the dataset page: https://huggingface.co/datasets/sardinelab/DocBlocks.texttranslation100K<n<1M4 likes133 downloads1y agoHugging Face07pauljngr /sardi-data SARDI — Evaluation Data Test splits and prebuilt BM25 indices for Self-Augmenting Retrieval for Diffusion Language Models (ICML 2026). Paper · Code · Model Download hf download pauljngr/sardi-data --repo-type dataset --local-dir data Contents dataset questions passages size 2WikiMultiHopQA 6,253 406,822 308 MB HotpotQA 3,701 5,239,002 2.7 GB MuSiQue 2,417 103,035 92 MB CofCA 900 3,156 6 MB SynthWorlds-SM 1,200 8,055 16 MB… See the full description on the dataset page: https://huggingface.co/datasets/pauljngr/sardi-data.textquestion-answering10K<n<100K0 likes121 downloads1mo agoHugging Face08average-developer /stocks-SARDAEN-1D-candlesn<1K0 likes115 downloads1d agoHugging Face09growan /SAR-Duty-Cycle Dataset Card for SAR-Duty-Cycle project This is a preliminary training and validation dataset for sub-aperture reconstruction using generative AI. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/growan/SAR-Duty-Cycle.10K<n<100K0 likes96 downloads1y agoHugging Face10IridescentOwl /SARD-Vulnerability-Datasettext100K<n<1M0 likes70 downloads1y agoHugging Face11sardinelab /MT-preftabular10K<n<100K5 likes37 downloads2y agoHugging Face12nicksteiner /sardine-demo-data SARdine demo data — Pacaya-Samiria full-frame NISAR COGs Full-frame Cloud Optimized GeoTIFFs backing the SARdine README hero demo. SARdine streams these directly in the browser via HTTP Range reads — the hero link renders in seconds while reading only the tiles in view. File Contents pacaya_full_hh.tif HHHH gamma0 backscatter, raw power float32 (~500 MB) pacaya_full_hv.tif HVHV gamma0 backscatter, raw power float32 (~495 MB) Source granule:… See the full description on the dataset page: https://huggingface.co/datasets/nicksteiner/sardine-demo-data.imagen<1K0 likes36 downloads1mo agoHugging Face13baebee /Sardiustextn<1K0 likes32 downloads3y agoHugging Face14sardukar /physiology-mcqa-8kThis dataset is a subset of MedMCQA textquestion-answering1K<n<10K2 likes32 downloads2y agoHugging Face15MattiaSangermano /SardiStance SardiStance Disclaimer: This dataset is not the official SardiStance repository from EVALITA. For the official dataset and more information, please visit the EVALITA SardiStance page and the SardiStance repository SardiStance is a unique dataset designed for the task of stance detection in Italian tweets. It consists of tweets related to the Sardines movement, providing a valuable resource for researchers and practitioners in the field of NLP. This dataset was curated for the… See the full description on the dataset page: https://huggingface.co/datasets/MattiaSangermano/SardiStance.texttext-classification1K<n<10K0 likes32 downloads2y agoHugging Face16sardinelab /bilingual_25k_filestext100K<n<1M0 likes28 downloads3y agoHugging Face17omlab /SARDet_REC6-FSimagen<1K0 likes28 downloads8mo agoHugging Face18Nicolas-BZRD /SARDE_opendata SARDE (Système d'Aide à la Recherche Documentaire Elaborée) SARDE is a repository designed to provide a thematic search mode for the majority of legislative and regulatory texts in force. The texts referenced are those published in the "Laws and Decrees" edition of the Journal officiel and in the Bulletins officiels distributed by the DILA. text100K<n<1M0 likes23 downloads3y agoHugging Face19sardinelab /flores-for-Towertext1K<n<10K0 likes17 downloads3y agoHugging Face20sardinelab /tico19-for-Towertext1K<n<10K0 likes17 downloads3y agoHugging Face21xya22er /test_sard0 likes16 downloads5mo agoHugging Face22sardinelab /wmt23-for-Towertext1K<n<10K0 likes13 downloads3y agoHugging Face23AI4DM /SARD0 likes13 downloads10mo agoHugging Face24FrancophonIA /dictionnaire_sarde_francais_italien_anglais_allemand [!NOTE] Dataset origin: https://web.archive.org/web/20121116095151/http:/www.sardegnacultura.it/documenti/7_81_20080107092727.pdf translation0 likes10 downloads1y agoHugging Face25sardor123 /biruniy-tts-dataaudio1K<n<10K0 likes10 downloads2mo agoHugging Face26sardinelab /MT-pref-humantabularn<1K0 likes9 downloads2y agoHugging Face27sardinelab /mt-align-study-w-idiom-1203tabular10K<n<100K0 likes8 downloads3y agoHugging Face28omlab /SARDet_REC6_NORM-FSimagen<1K0 likes7 downloads8mo agoHugging Face29ChaseCheng /SAR-DRG Dataset Card for SAR-DRG Dataset Details SAR-DRG is a scaffold-pocket dataset for realistic R-chain generation in lead optimization. It provides affinity-labeled R-chain samples organized by shared scaffold-pocket contexts, supporting SAR-informed molecular generation and evaluation. The dataset contains 96,158 samples across 29,387 scaffold-pocket groups. Dataset Architecture data_parquet/: Contains the processed Parquet files organized according to the… See the full description on the dataset page: https://huggingface.co/datasets/ChaseCheng/SAR-DRG.tabular10K<n<100K1 likes7 downloads5mo agoHugging Face30ReginaFoley /sar_data_512image1K<n<10K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.