CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmritaBha /mscoco-controlnet-cannyimage100K<n<1M4 likes620 downloads2y agoHugging Face02amrosama /al-kawakib-magazine-ocr Al-Kawakib Magazine OCR Pages This dataset contains rendered Arabic magazine page images paired with page-level text and line-level bounding boxes. It is intended for OCR, document understanding, and VLM fine-tuning experiments. Fine-Tuning Notebook A standalone Google Colab notebook for DeepSeek-OCR 3B + TRL SFT is available here: Open the fine-tuning notebook in Colab The notebook can train on this dataset alone or on all three Arabic magazine OCR… See the full description on the dataset page: https://huggingface.co/datasets/amrosama/al-kawakib-magazine-ocr.imageimage-to-text10K<n<100K3 likes603 downloads2mo agoHugging Face03amrachraf /arXiv-full-text-chunked Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.texttext-generation100K<n<1M1 likes447 downloads2y agoHugging Face04AMR-KELEG /DLAMA-v1 DLAMA-v1 A representative benchmark of factual triples curated from Wikidata and Wikipedia. Predicate Template P17 (Country) [X] is located in [Y] . P19 (Place of birth) [X] was born in [Y] . P20 (Place of death) [X] died in [Y] . P27 (Country of citizenship) [X] is [Y] citizen . P30 (Continent) [X] is located in [Y] . P36 (Capital) The capital of [X] is [Y] . P37 (Official language) The official language of [X] is [Y] . P47 (Shares border with) [X] shares… See the full description on the dataset page: https://huggingface.co/datasets/AMR-KELEG/DLAMA-v1.text100K<n<1M0 likes442 downloads1y agoHugging Face05amrosama /al-lataif-al-musawwara-magazine-ocr Al-Lataif Al-Musawwara Magazine OCR Pages This dataset contains rendered Arabic magazine page images from Al-Lataif Al-Musawwara paired with page-level text and line-level bounding boxes. It is intended for OCR, document understanding, and VLM fine-tuning experiments. Fine-Tuning Notebook A standalone Google Colab notebook for DeepSeek-OCR 3B + TRL SFT is available here: Open the fine-tuning notebook in Colab The notebook can train on this dataset alone or… See the full description on the dataset page: https://huggingface.co/datasets/amrosama/al-lataif-al-musawwara-magazine-ocr.imageimage-to-text1K<n<10K8 likes196 downloads2mo agoHugging Face06amritha27 /cl3410-phase1 CL3410 Phase 1 — Malayalam and Assamese language-model corpora Two independently built pretraining corpora with their own tokenizers: Malayalam as the higher-resource language and Assamese as the lower-resource one. Nothing is shared between them — separate sources, separate cleaning thresholds, separate vocabularies, separate models. Only the language-agnostic pipeline code is common, parameterised per language. Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.texttext-generation100M<n<1B0 likes156 downloads8d agoHugging Face07AmrahMaryam /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/AmrahMaryam/cnn_dailymail.textsummarization100K<n<1M0 likes140 downloads8mo agoHugging Face08amrithagk /capstone_sakuga_simple_description_mlm_hstabular10K<n<100K0 likes121 downloads2y agoHugging Face09AmritaBha /mscoco-colour_masksimage10K<n<100K0 likes110 downloads2y agoHugging Face10amrachraf /arXiv-full-text-chunked-qatext100K<n<1M0 likes109 downloads2y agoHugging Face11amrithagk /capstone_sakuga_simple_descriptiontabular10K<n<100K0 likes92 downloads2y agoHugging Face12amrithagk /sakuga_preprocessedtabular100K<n<1M0 likes89 downloads2y agoHugging Face13Tverous /mnli-amr Dataset Card for "mnli-amr" More Information needed text100K<n<1M0 likes81 downloads3y agoHugging Face14amrithanandini /phishing-email-rich-dataset-v2tabular10K<n<100K1 likes53 downloads8mo agoHugging Face15electricsheepafrica /africa-synth-antibiotic-quality-amr-all Antibiotic Quality & AMR Acceleration (SSA) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-antibiotic-quality-amr-all.imagetabular-classificationn<1K0 likes49 downloads1mo agoHugging Face16Chaymaa /UFPR-AMR Dataset Card for "UFPR-AMR" More Information needed image1K<n<10K0 likes48 downloads3y agoHugging Face17DataMuncher-Labs /AMReasoning-100000 All Mathematical Reasoning-100000 Built via a python script Contents add sub mul div linear_eq two_step_eq fraction exponent inequality word_algebra quadratic system_2x2 abs_eq percent mod simplify_expr mixed_fraction neg_div linear_fraction_eq rational_eq quadratic_nonunit cubic_int_root system_3x3 diophantine exponential_eq log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-100000.text100K<n<1M0 likes45 downloads9mo agoHugging Face18helena-balabin /vg_coco_overlap_for_graphormer_processed_amr_graphsimage10K<n<100K0 likes43 downloads1y agoHugging Face19amrithagk /capstone_sakuga_mlm_text_outputtabular100K<n<1M0 likes41 downloads2y agoHugging Face20ayates /amr_portal AMR Portal — Multi-Dataset Release (Phenotype + Genotype) This repository contains multiple datasets from the EMBL-EBI AMR Portal, distributed in Apache Parquet format: phenotype.parquet – phenotypic antimicrobial susceptibility data genotype.parquet – AMR genes and mutations from in silico methods All datasets are released under CC-BY-4.0. Source documentation: https://www.ebi.ac.uk/amr/developers/ Dataset Summary This dataset contains phenotypic antimicrobial… See the full description on the dataset page: https://huggingface.co/datasets/ayates/amr_portal.tabular1M<n<10M0 likes41 downloads10mo agoHugging Face21Amr04 /UnitedNations-ParagraphsAlligned-ar-en-datasettext1K<n<10K0 likes40 downloads28d agoHugging Face22DataMuncher-Labs /AMReasoning-750000 All Mathematical Reasoning-750000 Built via a python script Contents add sub mul div linear_eq two_step_eq fraction exponent inequality word_algebra quadratic system_2x2 abs_eq percent mod simplify_expr mixed_fraction neg_div linear_fraction_eq rational_eq quadratic_nonunit cubic_int_root system_3x3 diophantine exponential_eq log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-750000.text100K<n<1M0 likes39 downloads9mo agoHugging Face23AMR-KELEG /PTCCThe Parallel Tunisian Constitution Corpus (PTCC) corpus is a corpus of 149 articles written in Modern Standard Arabic and Tunisian Arabic. Tesseract was used to transform the constitution's pdf files into text files. Afterward, alignment of the parallel articles was achieved by a simple Python script. More details can be found in: https://amr-keleg.github.io/projects/digitalizing_dialectal_arabic/ Sources: Tunisian Arabic translation of the 2014 Tunisian Constitution 2014… See the full description on the dataset page: https://huggingface.co/datasets/AMR-KELEG/PTCC.texttext-generationn<1K0 likes37 downloads3y agoHugging Face24amrachraf /arXiv-full-text-chunked-testtextn<1K0 likes36 downloads2y agoHugging Face25DataMuncher-Labs /AMReasoning-50000 All Mathematical Reasoning-50000 Built via a python script Contents add sub mul div linear_eq two_step_eq fraction exponent inequality word_algebra quadratic system_2x2 abs_eq percent mod simplify_expr mixed_fraction neg_div linear_fraction_eq rational_eq quadratic_nonunit cubic_int_root system_3x3 diophantine exponential_eq log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-50000.text10K<n<100K0 likes31 downloads9mo agoHugging Face26DataMuncher-Labs /AMReasoning-2500000 All Mathematical Reasoning-2500000 Built via a python script Contents add sub mul div linear_eq two_step_eq fraction exponent inequality word_algebra quadratic system_2x2 abs_eq percent mod simplify_expr mixed_fraction neg_div linear_fraction_eq rational_eq quadratic_nonunit cubic_int_root system_3x3 diophantine exponential_eq log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-2500000.text1M<n<10M0 likes31 downloads9mo agoHugging Face27DataMuncher-Labs /AMReasoning-250000 All Mathematical Reasoning-250000 Built via a python script Contents add sub mul div linear_eq two_step_eq fraction exponent inequality word_algebra quadratic system_2x2 abs_eq percent mod simplify_expr mixed_fraction neg_div linear_fraction_eq rational_eq quadratic_nonunit cubic_int_root system_3x3 diophantine exponential_eq log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-250000.text100K<n<1M0 likes30 downloads9mo agoHugging Face28Nam-toon-studio /AMRIT-Punjabi-Clinical-Dialogue-Corpus ੴ AMRIT Punjabi Clinical Dialogue & Medical Diagnosis Corpus ☬ ਅੰਮ੍ਰਿਤ ਪੰਜਾਬੀ ਕਲੀਨਿਕਲ ਸੰਵਾਦ ਅਤੇ ਡਾਕਟਰੀ ਨਿਦਾਨ ਡਾਟਾਸੈੱਟ (v1.0) 👨‍💻 Research & Medical AI Architecture Lead Developer: Gurpreet Singh Dhillon (Nam-toon Studio) Mission: Free Autonomous AI Doctor for Humanity (ਦੁਨੀਆਂ ਦੇ ਲੋੜਵੰਦ ਲੋਕਾਂ ਲਈ ਮੁਫ਼ਤ AI ਡਾਕਟਰ) Flagship Platform: AMRIT Research OS (100% Local Medical Intelligence) 📖 Dataset Overview / ਸੰਖੇਪ The… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/AMRIT-Punjabi-Clinical-Dialogue-Corpus.textquestion-answeringn<1K0 likes30 downloads3d agoHugging Face29hoshuhan /amr-3-parsed Dataset Card for AMR 3.0 Parsed Dataset Summary This dataset contains parsed Abstract Meaning Representation (AMR) annotations from the LDC2020T02 release, formatted as instruction-following conversations. Each example consists of a sentence and its corresponding AMR graph representation. Supported Tasks and Leaderboards Tasks: Semantic parsing, specifically generating AMR graphs from English sentences Leaderboards: AMR Parsing Languages The… See the full description on the dataset page: https://huggingface.co/datasets/hoshuhan/amr-3-parsed.text10K<n<100K0 likes29 downloads2y agoHugging Face30DataMuncher-Labs /AMReasoning-300000 All Mathematical Reasoning-300000 Built via a python script Contents add sub mul div linear_eq two_step_eq fraction exponent inequality word_algebra quadratic system_2x2 abs_eq percent mod simplify_expr mixed_fraction neg_div linear_fraction_eq rational_eq quadratic_nonunit cubic_int_root system_3x3 diophantine exponential_eq log_eq… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/AMReasoning-300000.text100K<n<1M0 likes29 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.