datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube-tiktok-trends-dataset-2025
🎬 YouTube Shorts & TikTok Trends (2025)
Author: Tarek MasryoLicense: CC BY 4.0
A structured snapshot of short-form video activity across YouTube Shorts and TikTok during 2025 (Jan–Aug).Built for content intelligence, analytics dashboards, and ML baselines (classification/regression).
What’s inside
This repository ships:
Two loadable dataset configs (via datasets.load_dataset):
default → ML-ready table (cleaned + modeling-friendly)
raw → raw video-level table (wider… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/youtube-tiktok-trends-dataset-2025.afri-dict
Afri-Dict
Dataset Summary
afri-dict is a bilingual dictionary dataset for four major African languages: Hausa, Igbo, Swahili, and Yoruba.
Entries include a headword, part-of-speech tag, and definition in English or the target African language.
This dataset can serve as a foundational resource for machine translation systems, language learning tools, spell checkers, cross-lingual search, and other NLP applications for African languages.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/taresco/afri-dict.football-matches-2025-dataset
⚽ European Football Matches 2024/2025 Season
Author: Tarek Masryo · KaggleLicense: CC BY 4.0 (Attribution)
📌 Dataset Summary
Clean and structured dataset with 1,941 matches from the 2024/2025 European football season across 6 competitions:
Premier League (England)
La Liga (Spain)
Serie A (Italy)
Bundesliga (Germany)
Ligue 1 (France)
UEFA Champions League (Europe)
Each match includes results, dates, referees, and detailed score breakdowns (full-time &… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/football-matches-2025-dataset.tartanair_videoSAMTOR_Novel_Target_Designs_GA-II
SAMTOR — SAM-competitive de novo designs (Technetium GA-II)
196 small molecules across two generations, generated de novo by the Technetium TC-43.ai engine (GA-II) and conditioned on the S-adenosylmethionine (SAM) pocket of human SAMTOR.
Each molecule was constructed against this pocket rather than selected from a compound library — docking (AutoDock Vina) came afterwards, to place and score the generated molecules in the site. With no approved drug, clinical candidate or… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/SAMTOR_Novel_Target_Designs_GA-II.Target-QA
🎯 Target-QA: The First QA Dataset Benchmarking Target Priorization Based on DepMap
📑 Dataset Summary
Target-QA is derived from the DepMap multi-omics and CRISPR screening cohorts, harmonized via BioMedGraphica.It enables multi-modal reasoning by combining numeric evidence, topological knowledge and language context for CRISPR target prioritization.
This dataset supports the training and benchmarking of… See the full description on the dataset page: https://huggingface.co/datasets/FuhaiLiAiLab/Target-QA.global-ev-infra-dataset
🌍 Global EV Charging Stations & EV Models (2025)
Author: Tarek MasryoLicense: CC BY 4.0Version: v1.0 (2025-09-15)
A clean, analysis-ready snapshot of global EV infrastructure:
Main stations table: 242,417 rows (charging sites)
Companion summaries: country + world rollups
EV models table for enrichment
📦 What’s inside (files)
All CSVs live under data/:
data/charging_station.csv — charging stations (main table)
data/charging_station_ml.csv — ML-oriented derived… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/global-ev-infra-dataset.ViHealthQA
Disclaimer:
The dataset may contain personal information crawled along with the contents of various sources. Please make a filter in pre-processing data before starting your research training.
SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts
This is the official repository for the ViHealthQA dataset from the paper SPBERTQA: A Two-Stage Question Answering System Based on Sentence Transformers for Medical Texts, which was… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/ViHealthQA.ViCTSD
Constructive and Toxic Speech Detection for Open-domain Social Media Comments in Vietnamese
This is the official repository for the UIT-ViCTSD dataset from the paper Constructive and Toxic Speech Detection for Open-domain Social Media Comments in Vietnamese, which was accepted at the IEA/AIE 2021.
Citation Information
The provided dataset is only used for research purposes!
@InProceedings{nguyen2021victsd,
author="Nguyen, Luan Thanh and Van Nguyen, Kiet and Nguyen… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/ViCTSD.gendec-dataset
Gendec: Gender Dection from Japanese Names with Machine Learning
This is the official repository for the Gendec framework from the paper Gendec: Gender Dection from Japanese Names with Machine Learning, which was accepted at the ISDA'23.
Citation Information
The provided dataset is only used for research purposes!
@misc{pham2023gendec,
title={Gendec: A Machine Learning-based Framework for Gender Detection from Japanese Names},
author={Duong Tien Pham and Luan… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/gendec-dataset.MELASee the GitHub repo for details.
hospital-deterioration-dataset
🏥 Hospital Deterioration — Simulated Early Warning
Clinical Time-Series Benchmark for Early Warning Models
A fully simulated hospital cohort for building and testing early warning models and clinical deterioration risk scores.Each admission includes up to 72 hours of hourly data: vitals, labs, patient context, and multiple deterioration outcomes — with a main label for “deterioration in the next 12 hours”.
All records are fully simulated, internally consistent, and… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/hospital-deterioration-dataset.TARA_Turkish_LLM_Benchmark
TARA: Turkish Advanced Reasoning Assessment Veri Seti
*Img Credit: Open AI ChatGPT
**English version is given below.**
Evaluation Notebook / Değerlendirme Not Defteri
Dataset Summary
TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.targetrag-qa-logs-corpus-data
🧠📚 RAG QA Logs & Corpus (Synthetic)
🧪 Multi-table synthetic RAG telemetry for quality, hallucinations, latency, and cost
A production-style, privacy-safe synthetic dataset that mimics telemetry exported from a real RAG system — from corpus → chunks → retrieval events → eval runs.
✅ Fully synthetic (no real users / orgs / PII).
⚡ Quick facts
Total rows: 103,255 across 6 linked tables
Labels (in eval_runs): is_correct, hallucination_flag, faithfulness_label… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/rag-qa-logs-corpus-data.VOZ-HSD
ViHateT5: Enhancing Hate Speech Detection in Vietnamese with A Unified Text-to-Text Transformer Model
This is the official repository for the VOZ-HSD dataset from the paper ViHateT5: Enhancing Hate Speech Detection in Vietnamese with A Unified Text-to-Text Transformer Model, which was accepted at the ACL'2024.
Citation Information
The provided dataset is only used for research purposes!
@inproceedings{thanh-nguyen-2024-vihatet5,
title = "{V}i{H}ate{T}5: Enhancing… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/VOZ-HSD.MidJourney_v5_Prompt_datasetDataset contain raw prompts from Mid Journey v5
Total Records : 4245117
Sample Data
AuthorID
Author
Date
Content
Attachments
Reactions
936929561302675456
Midjourney Bot#9282
04/20/2023 12:00 AM
benjamin frankling with rayban sunglasses reflecting a usa flag walking on a side of penguin, whit...
Link
936929561302675456
Midjourney Bot#9282
04/20/2023 12:00 AM
Street vendor robot in 80's Poland, meat market, fruit stall, communist style, real photo, real ph...
Link… See the full description on the dataset page: https://huggingface.co/datasets/tarungupta83/MidJourney_v5_Prompt_dataset.ViOCD
Vietnamese Open-Domain Complaint Detection in E-commerce Websites
This is the official repository for the ViOCD dataset from the paper Vietnamese Open-Domain Complaint Detection in E-commerce Websites, which was accepted at the SoMeT 2021.
Citation Information
The provided dataset is only used for research purposes!
@misc{nguyen2021vietnamese,
title={Vietnamese Complaint Detection on E-Commerce Websites},
author={Nhung Thi-Hong Nguyen and Phuong Phan-Dieu Ha… See the full description on the dataset page: https://huggingface.co/datasets/tarudesu/ViOCD.redteaming-attack-target
Annotated version of DEFCON 31 Generative AI Red Teaming dataset with additional labels for attack targets.
This dataset is an extended version of the DEFCON31 Generative AI Red Teaming dataset, released by Humane Intelligence.
Our team conducted additional labeling on the accepted attack samples to annotate:
Attack Targets (e.g., gender, race, age, political orientation)
Attack Types (e.g., question, request, build-up, scenario assumption, misinformation injection) →… See the full description on the dataset page: https://huggingface.co/datasets/TTA01/redteaming-attack-target.llm-system-ops-production-telemetry-sft-data
🤖📈 LLM System Ops Telemetry (Synthetic)
A synthetic, production-style, multi-table LLM telemetry dataset designed for LLMOps analytics and decision-grade experiments.
It supports monitoring cost, latency, tokens, failures, safety flags, tool usage, and user feedback at the interaction level,
with rollups at the session and user levels — plus an SFT table aligned 1:1 with interactions and a prompt/config dimension.
Synthetic data (safe for teaching, prototyping, and portfolio… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/llm-system-ops-production-telemetry-sft-data.22-an-chinh-tarot
22 lá Ẩn Chính Tarot
The 22 Major Arcana of the Tarot
1. Mô tả · Description
Bộ Major Arcana kèm tên Việt và Anh, từ khoá, nghĩa xuôi và nghĩa ngược.
The Major Arcana with Vietnamese and English names, keywords, and upright and reversed meanings.
Số dòng · Rows: 22
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type
Ý nghĩa · Meaning
id
string
Định danh ổn định của… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/22-an-chinh-tarot.cancer-risk-factors-data
🧬 Cancer Risk Factors & Types (2,000 Rows)
Author: Tarek Masryo · KaggleLicense: CC BY 4.0 (Attribution) — Free for research, education, and commercial use
📌 Dataset Summary
Clean, standardized tabular dataset linking lifestyle, environmental, and genetic factors to five cancer types.
2,000 rows × 21 columns
Encodings: ordinal exposure indices (0–10), demographics (Age, BMI, Gender), binary flags (0/1) for family/genetics/infection
Includes engineered fields:… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/cancer-risk-factors-data.digital-lifestyle-benchmark-dataset
📱 Digital Lifestyle Benchmark Dataset (2025)
Author: Tarek Masryo · KaggleLicense: CC BY 4.0 (Attribution)
A structured tabular benchmark of 3,500 synthetic digital-lifestyle records with 24 columns.
It captures device usage patterns, attention signals, sleep/activity habits, and mental well-being scores, with a binary risk flag.
📦 What’s inside
Canonical file: data/digital_lifestyle_benchmark_2025.csvUnit of analysis: 1 row = 1 synthetic participant recordRows: 3… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/digital-lifestyle-benchmark-dataset.deckaura-tarot-card-meanings
Tarot Card Meanings — Complete 78-Card Deck
This dataset contains the complete 78-card tarot deck with structured interpretations across upright, reversed, love, career, and yes/no dimensions. Maintained by Deckaura, a US-based oracle & tarot knowledge project.
Hugging Face: Blacik/deckaura-tarot-card-meanings
DOI (Zenodo): 10.5281/zenodo.19475329
Canonical source: deckaura.com/blogs/guide/tarot-card-meanings
Dataset Summary
Total rows: 78 (22 Major Arcana + 56… See the full description on the dataset page: https://huggingface.co/datasets/Blacik/deckaura-tarot-card-meanings.peptide_target_residues_affinitypiqa_yoruba_pidgin
Physical Commonsense Reasoning for Yorùbá and Nigerian Pidgin
Dataset Summary
This dataset was developed for the MRL 2025 Shared Task on Multilingual Physical Reasoning. For more details, see Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures.
It provides a test collection for evaluating physical commonsense reasoning, that is, a model's ability to understand how objects, actions, and outcomes relate in everyday scenarios.
The… See the full description on the dataset page: https://huggingface.co/datasets/taresco/piqa_yoruba_pidgin.alzheimers-multi-target-10k-dataset
📚 BioDockify: Multi-Target Alzheimer's Chemical Space & Virtual Screening Dataset (10,000 Verified Compounds)
Principal Investigator: Tajuddin Shaik (tajo9128@gmail.com)Affiliation: Faculty of Pharmacy, Bharath Institute of Higher Education and Research (BIHER), Chennai, IndiaPlatform: www.biodockify.com | ai.biodockify.com
📌 Dataset Summary
This repository contains the complete 10,000 curated, literature-grounded chemical space dataset for Alzheimer's… See the full description on the dataset page: https://huggingface.co/datasets/BioDockify/alzheimers-multi-target-10k-dataset.Telangana_time_series_2023-2025The dataset was retrieved from Open Data Telangana, from February 1, 2023, to January 31, 2025 with daily granularity. The dataset contains various fields such as District, Mandal, Date, rainfall (in millimeters), minimum and maximum temperature (in Celsius), minimum and maximum wind speed, and humidity. It provides a District and Mandal wise distribution as well.
Total Rows - 4,45,213
Total Columns - 10
blood-donation-registry-dataset
🩸 Blood Donation Registry — Synthetic Donors, Prevalence & Compatibility
Synthetic, decision-focused tables for blood donation operations: donor eligibility/deferrals, donation history, rare blood types, country-level prevalence, and RBC transfusion compatibility (ABO/Rh).
Synthetic data (safe for experimentation and teaching)Not clinical/medical ground truth — do not use for real-world medical decisions.
📦 What’s inside
This repo provides four loadable dataset… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/blood-donation-registry-dataset.africa-synth-energy-tariff-subsidy-africa-niger
Africa Synth Energy Tariff Subsidy Africa Niger | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-tariff-subsidy-africa-niger.
