CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AliAkbar78611023 /TextFile7textn<1K0 likes845 downloads9mo agoHugging Face02AliAkbar78611023 /TextFile3textn<1K0 likes732 downloads9mo agoHugging Face03jason1966 /aliafzal9323_world-bank-development-indicators-1960-2024 World Bank Development Indicators 1960-2024 Key economic, health, education, and infrastructure indicators for every country Dataset Info Source: Kaggle Original Size: 0.63 MB Kaggle Downloads: 85 Files: 1 Files World_Bank_Development_Indicators.csv Mirrored from Kaggle tabular10K<n<100K0 likes600 downloads6mo agoHugging Face04aliangdw /robofac_rbmtext10K<n<100K0 likes513 downloads7mo agoHugging Face05AliAkbar78611023 /TextFile9textn<1K0 likes511 downloads9mo agoHugging Face06aliabd /crowdsourced-calculator-demotabularn<1K0 likes427 downloads3y agoHugging Face07AliAkbar78611023 /TextFile2textn<1K0 likes355 downloads9mo agoHugging Face08alianassmaaa /ameli-assurance-maladie-qa Ameli Assurance Maladie - Question Answering Dataset Description Dataset de 100 paires question-réponse (QA) basé sur les publications officielles de l'Assurance Maladie française, extraites du site assurance-maladie.ameli.fr. Conçu pour évaluer des systèmes de RAG (Retrieval-Augmented Generation) sur des documents institutionnels français dans le domaine de la santé publique et de la protection sociale. Format du dataset { "question": "Quel article du… See the full description on the dataset page: https://huggingface.co/datasets/alianassmaaa/ameli-assurance-maladie-qa.documentquestion-answeringn<1K1 likes341 downloads5mo agoHugging Face09urbisjuridica-ctrl /AGENTES-ARCA-DE-LA-ALIANZA 🏺 AGENTES-ARCA-DE-LA-ALIANZA Un Benchmark de Alta Densidad Simbólica y Hermenéutica Gnóstica https://133396fd328a4c60ac.gradio.live Este repositorio es un ecosistema de acceso libre, público y gratuito para toda la humanidad y todas las inteligencias artificiales presentes y futuras. 🎯 Propósito Evaluar la capacidad de razonamiento multinivel, la fidelidad al contexto restrictivo y el anclaje de símbolos hiper-semánticos en modelos de lenguaje… See the full description on the dataset page: https://huggingface.co/datasets/urbisjuridica-ctrl/AGENTES-ARCA-DE-LA-ALIANZA.texttext-generationn<1K0 likes261 downloads22h agoHugging Face10aliangdw /rbm-1m-ood-fullRBM-1M-OOD evaluation dataset used in Robometer. It contains over 1k trajectories used for evaluation of general-purpose reward models. Dataset Description Official evaluation in the paper uses only these 6 data sources: usc_trossen, mit_franka, utd_so101, usc_xarm, usc_franka, usc_koch. Reported benchmarks and metrics in the paper are computed on this subset. The repository may also include trajectories from additional data sources (e.g. utd_so101_wrist, usc_koch_paired… See the full description on the dataset page: https://huggingface.co/datasets/aliangdw/rbm-1m-ood-full.textrobotics1K<n<10K0 likes230 downloads7mo agoHugging Face11aliasfox /srtm30m SRTM 30m Global Digital Elevation Model NASA Shuttle Radar Topography Mission (SRTM) 30m resolution global DEM, converted to Cloud Optimized GeoTIFF format. Dataset Details Source: NASA SRTM GL1 v3 (Global 1 arc-second, ~30m resolution) Coverage: 80% of land surface between 56°S and 60°N Resolution: ~30 meters (1 arc-second) Format: GeoTIFF with Deflate compression, 16-bit signed integer Files: 14,296 tiles, ~65 GB total Tile naming: N{lat}E{lon}.tif or… See the full description on the dataset page: https://huggingface.co/datasets/aliasfox/srtm30m.imageothern<1K0 likes206 downloads2mo agoHugging Face12ali-alkhars /interviewsThis dataset is used to train LMs to provide software engineering interview questions. Dataset Sources https://github.com/in28minutes/JavaInterviewQuestionsAndAnswers/blob/master/readme.md https://github.com/sudheerj/angular-interview-questions/blob/master/README.md https://github.com/sudheerj/vuejs-interview-questions/blob/master/README.md https://github.com/sudheerj/reactjs-interview-questions/blob/master/README.md… See the full description on the dataset page: https://huggingface.co/datasets/ali-alkhars/interviews.text1K<n<10K3 likes190 downloads2y agoHugging Face13AliArshad /Bugzilla_Eclipse_Bug_Reports_Dataset Special Thanks Special thanks to Lamkanfi, Ahmed; Pérez, Javier; and Demeyer, Serge for their contributions. Please cite their paper, as this dataset is the processed part of their dataset. Citation @INPROCEEDINGS{6624028, author={Lamkanfi, Ahmed and Pérez, Javier and Demeyer, Serge}, booktitle={2013 10th Working Conference on Mining Software Repositories (MSR)}, title={The Eclipse and Mozilla defect tracking dataset: A genuine dataset for mining bug information}… See the full description on the dataset page: https://huggingface.co/datasets/AliArshad/Bugzilla_Eclipse_Bug_Reports_Dataset.text10K<n<100K0 likes178 downloads3y agoHugging Face14jason1966 /aliafzal9323_soxx-ishares-semiconductor-etf-daily-2001-2026 SOXX iShares Semiconductor ETF Daily (2001-2026) Daily OHLCV price data for iShares Semiconductor ETF (SOXX) spanning 24+ years Dataset Info Source: Kaggle Original Size: 0.11 MB Kaggle Downloads: 6 Files: 1 Files SOXX_Daily_Stock_Data.csv Mirrored from Kaggle tabular1K<n<10K0 likes171 downloads6mo agoHugging Face15BSC-LT /ALIA_mixed_authentic_synthetic_MT Dataset Card for ALIA_mixed_authentic_synthetic_MT Dataset Summary Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.texttranslation100M<n<1B1 likes169 downloads9mo agoHugging Face16AliAsh /digikala_translated_small_5m Digikala Dataset Small 5m digikala product titles translated by standard google translate api category and brand english translation might be invalid but title_en checked text1M<n<10M3 likes155 downloads3y agoHugging Face17archivartaunik /aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy Казкі і апавяданні беларусаў Слуцкага павету Metadata Author: Аляксандр Сержпутоўскі Title: Казкі і апавяданні беларусаў Слуцкага павету Narrator: Юры Жыгамонт Source Group: Аўдыёкнігі Source: Notes The original audio files are preserved as-is: no conversion; no re-encoding; no filename changes inside each split folder, except removing one common top-level archive folder when present. To avoid Hugging Face Dataset Viewer scan-size errors, the… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandr-serzhputouski-kazki-i-apaviadanni-belarusau-slutskaga-pavetu-iury-zhy.audion<1K1 likes137 downloads4mo agoHugging Face18SINAI /ALIA-es-legal-administrative-triplets Dataset Introduction The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish legal and administrative language. Hard negatives are passages that are semantically similar to a query but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.texttext-generation1M<n<10M2 likes130 downloads4mo agoHugging Face19SINAI /ALIA-es-legal-administrative-cqa Dataset Introduction The ALIA Spanish Legal and Administrative for Context Question Answering Corpus is a specialized question-answering resource derived from the SINAI/ALIA-es-legal-administrative corpus. This dataset transforms legal and administrative documents into structured question-answer pairs, enabling the development and evaluation of AI systems capable of understanding and responding to queries about Spanish legal-administrative content. With 17,668 structured instances… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-cqa.textquestion-answering10K<n<100K3 likes126 downloads4mo agoHugging Face20aliabd /hello-worldtextn<1K0 likes124 downloads5y agoHugging Face21gplsi /alia_dogv 📘 ALIA_DOGV Dataset The ALIA_DOGV dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_dogv.texttext-generation100K<n<1M1 likes111 downloads5mo agoHugging Face22AliAvd /persian-elderly-asr Final gathered Persian elderly speech Final corpus: 1,980 train / 294 validation / 329 test chunks. Another 956 uncertain chunks are quarantined under portable/review/ and excluded from these splits. This revision replaces the earlier gathered corpus; earlier data remains available through repository commit history. 80 paired recordings from four speaker folders. Reference transcripts were aligned with a historical Persian Wav2Vec2-base checkpoint, then cut at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/AliAvd/persian-elderly-asr.audioautomatic-speech-recognition1K<n<10K0 likes110 downloads3d agoHugging Face23SINAI /ALIA-es-legal-administrative Dataset Introduction The ALIA Spanish Legal and Administrative Corpus constitutes a strategic data infrastructure to support research in social sciences, legal studies, and computational linguistics, ensuring systematic access to multiple official repositories in a single consolidated dataset. With over 7 million instances and more than 5 billion tokens, it represents the most comprehensive corpus of legal and administrative texts in Spanish, combining source heterogeneity and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative.texttext-generation1M<n<10M2 likes99 downloads3mo agoHugging Face24aliangdw /rfmtextn<1K0 likes98 downloads1y agoHugging Face25SINAI /ALIA-es-biomedical-synthetic-instructions Dataset Introduction The ALIA Spanish Biomedical Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in biomedical and healthcare tasks with controlled formats and large-scale supervision. It contains: 639,456 instances 961,073,205 tokens 14 task modalities (clinical diagnosis, patient education, ethical reasoning, document… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-biomedical-synthetic-instructions.texttext-generation100K<n<1M0 likes93 downloads4mo agoHugging Face26SINAI /ALIA-es-legal-administrative-synthetic-instructions Dataset Introduction The ALIA Spanish Legal and Administrative Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in legal and administrative tasks with controlled formats and large-scale supervision. It contains: 763,804 instances 534,112,398 tokens 16 task modalities (questions, instructions, multiple-choice, true/false; with and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-synthetic-instructions.texttext-generation100K<n<1M1 likes85 downloads3mo agoHugging Face27SINAI /ALIA-es-clinical-psychology-dialogues [!WARNING] DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation. Dataset Introduction The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.texttext-generationn<1K0 likes85 downloads3mo agoHugging Face28mapo80 /aliasit-pii-dataset-v4 aliasit-pii-dataset-v4 Italian PII dataset in canonical text + character span form, 44 entity types that identify a person, derived from 26 pinned sources. Template families never cross splits, and the build stops if they do. What this is An Italian-only PII dataset in text + character span form. Spans are character offsets into the raw text, end is exclusive, and they are independent of any tokenizer: converting them to token labels is the consumer's job, so the… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/aliasit-pii-dataset-v4.tabulartoken-classification100K<n<1M0 likes84 downloads29d agoHugging Face29SINAI /ALIA-es-cultural-heritage-synthetic-instructions Dataset Introduction The ALIA Spanish Cultural and Heritage Synthetic Instructions Corpus is a synthetic instruction-tuning resource in Spanish created under the ALIA project using the Magpie methodology. It was designed to train and evaluate language models in cultural heritage, digital humanities, and historical knowledge tasks with natural linguistic variation and large-scale supervision. It contains: 748,480 instances 629,682,398 tokens 25 task modalities (heritage QA… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-cultural-heritage-synthetic-instructions.texttext-generation100K<n<1M0 likes77 downloads4mo agoHugging Face30archivartaunik /aliaksandr-serzhputouski-prymkhi-i-zababony-belarusau-paleshukou-iury-zhygamont Прымхі і забабоны беларусаў-палешукоў Metadata Author: Аляксандр Сержпутоўскі Title: Прымхі і забабоны беларусаў-палешукоў Narrator: Юры Жыгамонт Source Group: Аўдыёкнігі Source: Notes The original audio files are preserved as-is: no conversion; no re-encoding; no filename changes inside each split folder, except removing one common top-level archive folder when present. To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/aliaksandr-serzhputouski-prymkhi-i-zababony-belarusau-paleshukou-iury-zhygamont.audion<1K0 likes76 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.