CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tarekmasryo /youtube-tiktok-trends-dataset-2025 🎬 YouTube Shorts & TikTok Trends (2025) Author: Tarek MasryoLicense: CC BY 4.0 A structured snapshot of short-form video activity across YouTube Shorts and TikTok during 2025 (Jan–Aug).Built for content intelligence, analytics dashboards, and ML baselines (classification/regression). What’s inside This repository ships: Two loadable dataset configs (via datasets.load_dataset): default → ML-ready table (cleaned + modeling-friendly) raw → raw video-level table (wider… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/youtube-tiktok-trends-dataset-2025.tabulartabular-regression10K<n<100K7 likes553 downloads8mo agoHugging Face02taresco /afri-dict Afri-Dict Dataset Summary afri-dict is a bilingual dictionary dataset for four major African languages: Hausa, Igbo, Swahili, and Yoruba. Entries include a headword, part-of-speech tag, and definition in English or the target African language. This dataset can serve as a foundational resource for machine translation systems, language learning tools, spell checkers, cross-lingual search, and other NLP applications for African languages. Languages… See the full description on the dataset page: https://huggingface.co/datasets/taresco/afri-dict.texttranslation10K<n<100K0 likes436 downloads1mo agoHugging Face03tarekmasryo /football-matches-2025-dataset ⚽ European Football Matches 2024/2025 Season Author: Tarek Masryo · KaggleLicense: CC BY 4.0 (Attribution) 📌 Dataset Summary Clean and structured dataset with 1,941 matches from the 2024/2025 European football season across 6 competitions: Premier League (England) La Liga (Spain) Serie A (Italy) Bundesliga (Germany) Ligue 1 (France) UEFA Champions League (Europe) Each match includes results, dates, referees, and detailed score breakdowns (full-time &… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/football-matches-2025-dataset.tabulartabular-classification1K<n<10K2 likes415 downloads8mo agoHugging Face04Alexhe101 /tartanair_videotext1K<n<10K0 likes274 downloads1y agoHugging Face05Tc-43 /SAMTOR_Novel_Target_Designs_GA-II SAMTOR — SAM-competitive de novo designs (Technetium GA-II) 196 small molecules across two generations, generated de novo by the Technetium TC-43.ai engine (GA-II) and conditioned on the S-adenosylmethionine (SAM) pocket of human SAMTOR. Each molecule was constructed against this pocket rather than selected from a compound library — docking (AutoDock Vina) came afterwards, to place and score the generated molecules in the site. With no approved drug, clinical candidate or… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/SAMTOR_Novel_Target_Designs_GA-II.tabulartabular-regressionn<1K0 likes263 downloads7d agoHugging Face06FuhaiLiAiLab /Target-QA 🎯 Target-QA: The First QA Dataset Benchmarking Target Priorization Based on DepMap 📑 Dataset Summary Target-QA is derived from the DepMap multi-omics and CRISPR screening cohorts, harmonized via BioMedGraphica.It enables multi-modal reasoning by combining numeric evidence, topological knowledge and language context for CRISPR target prioritization. This dataset supports the training and benchmarking of… See the full description on the dataset page: https://huggingface.co/datasets/FuhaiLiAiLab/Target-QA.tabularquestion-answeringn<1K0 likes253 downloads1y agoHugging Face07tarekmasryo /global-ev-infra-dataset 🌍 Global EV Charging Stations & EV Models (2025) Author: Tarek MasryoLicense: CC BY 4.0Version: v1.0 (2025-09-15) A clean, analysis-ready snapshot of global EV infrastructure: Main stations table: 242,417 rows (charging sites) Companion summaries: country + world rollups EV models table for enrichment 📦 What’s inside (files) All CSVs live under data/: data/charging_station.csv — charging stations (main table) data/charging_station_ml.csv — ML-oriented derived… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/global-ev-infra-dataset.tabulartabular-regression100K<n<1M2 likes170 downloads8mo agoHugging Face08tarudesu /ViHealthQA Disclaimer: The dataset may contain personal information crawled along with the contents of various sources. Please make a filter in pre-processing data before starting your research training. SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts This is the official repository for the ViHealthQA dataset from the paper SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts, which was… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/ViHealthQA.textquestion-answering10K<n<100K17 likes138 downloads3y agoHugging Face09tarudesu /ViCTSD Constructive and Toxic Speech Detection for Open-domain Social Media Comments in Vietnamese This is the official repository for the UIT-ViCTSD dataset from the paper Constructive and Toxic Speech Detection for Open-domain Social Media Comments in Vietnamese, which was accepted at the IEA/AIE 2021. Citation Information The provided dataset is only used for research purposes! @InProceedings{nguyen2021victsd, author="Nguyen, Luan Thanh and Van Nguyen, Kiet and Nguyen… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/ViCTSD.tabulartext-classification10K<n<100K2 likes135 downloads3y agoHugging Face10tarudesu /gendec-dataset Gendec: Gender Dection from Japanese Names with Machine Learning This is the official repository for the Gendec framework from the paper Gendec: Gender Dection from Japanese Names with Machine Learning, which was accepted at the ISDA'23. Citation Information The provided dataset is only used for research purposes! @misc{pham2023gendec, title={Gendec: A Machine Learning-based Framework for Gender Detection from Japanese Names}, author={Duong Tien Pham and Luan… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/gendec-dataset.texttext-classification10K<n<100K3 likes117 downloads3y agoHugging Face11Geralt-Targaryen /MELASee the GitHub repo for details. texttext-classification10K<n<100K0 likes113 downloads2y agoHugging Face12tarekmasryo /hospital-deterioration-dataset 🏥 Hospital Deterioration — Simulated Early Warning Clinical Time-Series Benchmark for Early Warning Models A fully simulated hospital cohort for building and testing early warning models and clinical deterioration risk scores.Each admission includes up to 72 hours of hourly data: vitals, labs, patient context, and multiple deterioration outcomes — with a main label for “deterioration in the next 12 hours”. All records are fully simulated, internally consistent, and… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/hospital-deterioration-dataset.tabulartabular-classification1M<n<10M1 likes94 downloads8mo agoHugging Face13emre /TARA_Turkish_LLM_Benchmark TARA: Turkish Advanced Reasoning Assessment Veri Seti *Img Credit: Open AI ChatGPT **English version is given below.** Evaluation Notebook / Değerlendirme Not Defteri Dataset Summary TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.textquestion-answeringn<1K28 likes89 downloads1y agoHugging Face14dreamland4dnam /targettabular1M<n<10M0 likes89 downloads29d agoHugging Face15tarekmasryo /rag-qa-logs-corpus-data 🧠📚 RAG QA Logs & Corpus (Synthetic) 🧪 Multi-table synthetic RAG telemetry for quality, hallucinations, latency, and cost A production-style, privacy-safe synthetic dataset that mimics telemetry exported from a real RAG system — from corpus → chunks → retrieval events → eval runs. ✅ Fully synthetic (no real users / orgs / PII). ⚡ Quick facts Total rows: 103,255 across 6 linked tables Labels (in eval_runs): is_correct, hallucination_flag, faithfulness_label… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/rag-qa-logs-corpus-data.tabularquestion-answering100K<n<1M2 likes86 downloads8mo agoHugging Face16tarudesu /VOZ-HSD ViHateT5: Enhancing Hate Speech Detection in Vietnamese with A Unified Text-to-Text Transformer Model This is the official repository for the VOZ-HSD dataset from the paper ViHateT5: Enhancing Hate Speech Detection in Vietnamese with A Unified Text-to-Text Transformer Model, which was accepted at the ACL'2024. Citation Information The provided dataset is only used for research purposes! @inproceedings{thanh-nguyen-2024-vihatet5, title = "{V}i{H}ate{T}5: Enhancing… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/VOZ-HSD.tabulartext-classification10M<n<100M13 likes83 downloads2y agoHugging Face17tarungupta83 /MidJourney_v5_Prompt_datasetDataset contain raw prompts from Mid Journey v5 Total Records : 4245117 Sample Data AuthorID Author Date Content Attachments Reactions 936929561302675456 Midjourney Bot#9282 04/20/2023 12:00 AM benjamin frankling with rayban sunglasses reflecting a usa flag walking on a side of penguin, whit... Link 936929561302675456 Midjourney Bot#9282 04/20/2023 12:00 AM Street vendor robot in 80's Poland, meat market, fruit stall, communist style, real photo, real ph... Link… See the full description on the dataset page: https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset.text1M<n<10M31 likes77 downloads3y agoHugging Face18tarudesu /ViOCD Vietnamese Open-Domain Complaint Detection in E-commerce Websites This is the official repository for the ViOCD dataset from the paper Vietnamese Open-Domain Complaint Detection in E-commerce Websites, which was accepted at the SoMeT 2021. Citation Information The provided dataset is only used for research purposes! @misc{nguyen2021vietnamese, title={Vietnamese Complaint Detection on E-Commerce Websites}, author={Nhung Thi-Hong Nguyen and Phuong Phan-Dieu Ha… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/ViOCD.tabulartext-classification1K<n<10K0 likes72 downloads3y agoHugging Face19TTA01 /redteaming-attack-target Annotated version of DEFCON 31 Generative AI Red Teaming dataset with additional labels for attack targets. This dataset is an extended version of the DEFCON31 Generative AI Red Teaming dataset, released by Humane Intelligence. Our team conducted additional labeling on the accepted attack samples to annotate: Attack Targets (e.g., gender, race, age, political orientation) Attack Types (e.g., question, request, build-up, scenario assumption, misinformation injection) →… See the full description on the dataset page: https://huggingface.co/datasets/TTA01/redteaming-attack-target.textn<1K1 likes72 downloads1y agoHugging Face20tarekmasryo /llm-system-ops-production-telemetry-sft-data 🤖📈 LLM System Ops Telemetry (Synthetic) A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments. It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level, with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension. Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.tabulartabular-classification10K<n<100K1 likes72 downloads8mo agoHugging Face21nhatnguyet /22-an-chinh-tarot 22 lá Ẩn Chính Tarot The 22 Major Arcana of the Tarot 1. Mô tả · Description Bộ Major Arcana kèm tên Việt và Anh, từ khoá, nghĩa xuôi và nghĩa ngược. The Major Arcana with Vietnamese and English names, keywords, and upright and reversed meanings. Số dòng · Rows: 22 Phiên bản · Version: 1.0.0 (2026-09-16) Mã hoá · Encoding: UTF-8 không BOM 2. Cấu trúc · Structure Cột · Column Kiểu · Type Ý nghĩa · Meaning id string Định danh ổn định của… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/22-an-chinh-tarot.tabularn<1K0 likes61 downloads3d agoHugging Face22tarekmasryo /cancer-risk-factors-data 🧬 Cancer Risk Factors & Types (2,000 Rows) Author: Tarek Masryo · KaggleLicense: CC BY 4.0 (Attribution) — Free for research, education, and commercial use 📌 Dataset Summary Clean, standardized tabular dataset linking lifestyle, environmental, and genetic factors to five cancer types. 2,000 rows × 21 columns Encodings: ordinal exposure indices (0–10), demographics (Age, BMI, Gender), binary flags (0/1) for family/genetics/infection Includes engineered fields:… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/cancer-risk-factors-data.tabulartabular-classification1K<n<10K4 likes59 downloads8mo agoHugging Face23tarekmasryo /digital-lifestyle-benchmark-dataset 📱 Digital Lifestyle Benchmark Dataset (2025) Author: Tarek Masryo · KaggleLicense: CC BY 4.0 (Attribution) A structured tabular benchmark of 3,500 synthetic digital-lifestyle records with 24 columns. It captures device usage patterns, attention signals, sleep/activity habits, and mental well-being scores, with a binary risk flag. 📦 What’s inside Canonical file: data/digital_lifestyle_benchmark_2025.csvUnit of analysis: 1 row = 1 synthetic participant recordRows: 3… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/digital-lifestyle-benchmark-dataset.tabulartabular-classification1K<n<10K1 likes53 downloads8mo agoHugging Face24Blacik /deckaura-tarot-card-meanings Tarot Card Meanings — Complete 78-Card Deck This dataset contains the complete 78-card tarot deck with structured interpretations across upright, reversed, love, career, and yes/no dimensions. Maintained by Deckaura, a US-based oracle & tarot knowledge project. Hugging Face: Blacik/deckaura-tarot-card-meanings DOI (Zenodo): 10.5281/zenodo.19475329 Canonical source: deckaura.com/blogs/guide/tarot-card-meanings Dataset Summary Total rows: 78 (22 Major Arcana + 56… See the full description on the dataset page: https://huggingface.co/datasets/Blacik/deckaura-tarot-card-meanings.texttext-classificationn<1K0 likes52 downloads3mo agoHugging Face25stilletto /peptide_target_residues_affinitytext10K<n<100K1 likes50 downloads2y agoHugging Face26taresco /piqa_yoruba_pidgin Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin Dataset Summary This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures. It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios. The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.textquestion-answeringn<1K2 likes45 downloads10mo agoHugging Face27BioDockify /alzheimers-multi-target-10k-dataset 📚 BioDockify: Multi-Target Alzheimer's Chemical Space & Virtual Screening Dataset (10,000 Verified Compounds) Principal Investigator: Tajuddin Shaik (tajo9128@gmail.com)Affiliation: Faculty of Pharmacy, Bharath Institute of Higher Education and Research (BIHER), Chennai, IndiaPlatform: www.biodockify.com | ai.biodockify.com 📌 Dataset Summary This repository contains the complete 10,000 curated, literature-grounded chemical space dataset for Alzheimer's… See the full description on the dataset page: https://huggingface.co/datasets/BioDockify/alzheimers-multi-target-10k-dataset.texttabular-classification10K<n<100K0 likes45 downloads24d agoHugging Face28tarun115027 /Telangana_time_series_2023-2025The dataset was retrieved from Open Data Telangana, from February 1, 2023, to January 31, 2025 with daily granularity. The dataset contains various fields such as District, Mandal, Date, rainfall (in millimeters), minimum and maximum temperature (in Celsius), minimum and maximum wind speed, and humidity. It provides a District and Mandal wise distribution as well. Total Rows - 4,45,213 Total Columns - 10 tabular100K<n<1M0 likes43 downloads6d agoHugging Face29tarekmasryo /blood-donation-registry-dataset 🩸 Blood Donation Registry — Synthetic Donors, Prevalence & Compatibility Synthetic, decision-focused tables for blood donation operations: donor eligibility/deferrals, donation history, rare blood types, country-level prevalence, and RBC transfusion compatibility (ABO/Rh). Synthetic data (safe for experimentation and teaching)Not clinical/medical ground truth — do not use for real-world medical decisions. 📦 What’s inside This repo provides four loadable dataset… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/blood-donation-registry-dataset.tabulartabular-classification10K<n<100K1 likes41 downloads8mo agoHugging Face30electricsheepafrica /africa-synth-energy-tariff-subsidy-africa-niger Africa Synth Energy Tariff Subsidy Africa Niger | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-tariff-subsidy-africa-niger.tabulartabular-classification10K<n<100K0 likes40 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.