datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hoaxpedia
license: mit
task_categories:
- text-classification
language:
- en
pretty_name: hoaxpedia
size_categories:
- 10K<n<100K
HOAXPEDIA: A Unified Wikipedia Hoax Articles Dataset
Hoaxpedia is a Dataset containing Hoax articles collected from Wikipedia and semantically similar Legitimate article in 2 settings: Fulltext and Definition and in 3 splits based on Hoax:Legit ratio (1:2,1:10,1:100).
Dataset Details
Dataset Description
We introduce H OAXPEDIA, a… See the full description on the dataset page: https://huggingface.co/datasets/hsuvaskakoty/hoaxpedia.decoding-robustness-results
Decoding Robustness Results
Mechanistic robustness evaluation results for language models under six input
perturbations: character replacement, BPE-token replacement, word replacement,
local token shuffle, typographical corruption, and synonym replacement.
The repository is organized by model and perturbation:
models/<model>/<perturbation>/<percentage>/evals.csv
The qwen2.5_1.5b/adversarial directory contains the separate adversarial
evaluation outputs and manifest. Failed or… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/decoding-robustness-results.fusion-image-to-latex-datasets
Collects and builds the largest dataset to date from online sources, creating a robust and generalizable dataset. This dataset includes approximately 3.4 million image-text pairs, including both handwritten mathematical expressions (200,330 examples) and printed mathematical expressions (3,237,250 examples). Due to the large dataset and the fact that the same mathematical formula can be represented in different LaTeX string formats in an image, it is easy to cause polymorphic ambiguity. To… See the full description on the dataset page: https://huggingface.co/datasets/hoang-quoc-trung/fusion-image-to-latex-datasets.60-hoa-giap
Bảng 60 Hoa Giáp
Sexagenary cycle table (60 Hoa Giap)
1. Mô tả · Description
Đủ 60 cặp Can Chi kèm ngũ hành nạp âm, tên nạp âm và các năm âm lịch tương ứng trong khoảng 1924 tới 2043.
All 60 stem-branch pairs with their Na Yin five-element attribution and the lunar years they cover between 1924 and 2043.
Số dòng · Rows: 60
Phiên bản · Version: 1.0.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu ·… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/60-hoa-giap.BiomedKeyRefine
BiomedKeyRefine: Refined Biomedical Keyword Extraction Benchmark
Version 1.0Last updated: February 2026Author: Christian Hoang (@christhoang04)License: MIT
🌟 Dataset Summary
BiomedKeyRefine is a high-quality, refined subset of the PubMedAKE benchmark (CIKM 2022), specifically curated for biomedical keyword extraction research.
The original PubMedAKE contains 843k+ articles with author-assigned keywords (extractive + abstractive). This version:
Filters for… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/BiomedKeyRefine.mcqa_hoax_1h10r_def_bigbenchmy-dataset-test222CTI-to-MITRE-datasetindonesian_hoax_news_oriNutuk_QA📌 Overview
This dataset consists of question-answer pairs automatically generated from the Turkish historical speech "Nutuk" by Mustafa Kemal Atatürk. Each entry includes a question, its answer, and a question type label. The dataset is intended for research, educational purposes, and for fine-tuning Turkish language models in various NLP tasks such as question answering, information retrieval, and instruction tuning.
📂 Dataset Structure
Each row in the dataset has the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/hoatac/Nutuk_QA.hoa-giap-ghep-cheo
Ghép chéo 60 Hoa Giáp: nạp âm, hợp khắc và quan hệ ngũ hành
The sixty stem-branch pairs crossed: sounds, harmony, clash and element relation
1. Mô tả · Description
Mọi cặp trong sáu mươi Hoa Giáp, ba nghìn sáu trăm dòng, cho biết quan hệ ngũ hành nạp âm giữa hai tuổi và cách xếp nhóm theo địa chi.
Every pair among the sixty stem-branch combinations, three thousand six hundred rows, giving the element relation between their sounds and the branch grouping.
Số dòng… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/hoa-giap-ghep-cheo.gio-hoang-dao-60-ngay
Giờ hoàng đạo theo 60 ngày can chi
Auspicious hours across the sixty-day stem-branch cycle
1. Mô tả · Description
Mười hai khung giờ của từng ngày trong vòng sáu mươi ngày can chi, cho biết khung nào là giờ hoàng đạo. Sáu mươi nhân mười hai là bảy trăm hai mươi dòng, tức đủ một chu kỳ.
The twelve double-hours of each day across the sixty-day stem-branch cycle, marking which are auspicious. Sixty times twelve is seven hundred and twenty rows, one full cycle.
Số… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/gio-hoang-dao-60-ngay.tam-tai-kim-lau-hoang-oc
Tam Tai, Kim Lâu, Hoang Ốc theo tuổi và năm xem
Tam Tai, Kim Lau and Hoang Oc by birth year and year in question
1. Mô tả · Description
Ghép chéo mọi năm sinh từ 1940 tới 2010 với mọi năm xem từ 2020 tới 2050, cho biết cặp ấy có phạm Tam Tai, Kim Lâu hay Hoang Ốc không, kèm phép tính để đối chiếu.
Every birth year from 1940 to 2010 crossed with every year in question from 2020 to 2050, showing whether the pair falls under Tam Tai, Kim Lau or Hoang Oc, with the… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/tam-tai-kim-lau-hoang-oc.bien-cung-hoang-dao-1950-2050
Biên ngày 12 cung hoàng đạo theo từng năm, 1950 tới 2050
Zodiac sign boundaries year by year, 1950 to 2050
1. Mô tả · Description
Ngày và giờ mặt trời vào từng cung hoàng đạo, tính riêng cho mỗi năm thay vì dùng một biên cố định.
The date and time the sun enters each zodiac sign, computed per year instead of using a fixed boundary.
Số dòng · Rows: 1,212
Phiên bản · Version: 1.0.0 (2026-09-21)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/bien-cung-hoang-dao-1950-2050.Legal-Documentindonesian_hoax_news_datasethoa-condo-accounting-glossary
HOA & Condominium Association Accounting Glossary
A structured reference glossary of accounting, financial-reporting and
governance terminology used by homeowners' associations (HOAs), condominium
associations, and cooperative associations — with a focus on Florida
community-association context (Chapters 718, 719 and 720, Florida Statutes).
48 terms across 7 categories: Fund Accounting, Financial Reporting,
Assessments & Collections, Reserves, Budgeting, Common Expenses, and… See the full description on the dataset page: https://huggingface.co/datasets/NODARISHUB/hoa-condo-accounting-glossary.data_spam_emailviet-robot-schedule-templates
viet-robot-schedule-templates
A small dataset of weekly cleaning schedule templates for a home robot.
Columns:
id: row id
template_id: short template identifier
description_vi: description in Vietnamese
description_en: description in English
days: comma-separated days (mon,tue,wed,thu,fri,sat,sun)
time_local: time of day in HH:MM (24h)
zone_id: generic zone identifier (e.g., LR_CENTER)
Manually authored for demos.
License
MIT
id-hoax-report-merge-v3We do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
The dataset is taken from nlp-brin-id/id-hoax-report-merge-v2 by filtering out null samples.
mmistralviet-robot-laundry-reminder-log
viet-robot-laundry-reminder-log
A synthetic dataset for laundry-related reminders generated by a home robot.
Columns:
id: row id
timestamp_local: local time string
basket_location: where the laundry basket is
fill_level_pct: estimated fill level 0–100
days_since_last_wash: integer days
reminder_sent: yes/no
user_action: did_wash, snoozed, ignored
note_en: short English note
License
MIT
viet-robot-nav-instructions
viet-robot-nav-instructions
A small dataset of Vietnamese navigation instructions for home robots.
Each row contains:
id: row id
instruction_vi: navigation instruction in Vietnamese
instruction_en: English translation of the instruction
target_room: target room or area
path_type: type of path (direct, multi_step, shortest, around, wall_follow, tour)
has_obstacle: whether the instruction mentions obstacles (yes/no)
The dataset is manually written for educational and… See the full description on the dataset page: https://huggingface.co/datasets/hoangs/viet-robot-nav-instructions.vnmese_sentimentsviet-robot-energy-usage-log-v2
viet-robot-energy-usage-log
A small synthetic dataset with coarse energy usage estimates for different
robot activities.
Columns:
id: row id
activity: cleaning, mapping, idle_docked, patrolling, voice_only
duration_min: duration in minutes
energy_wh: estimated energy used in Wh
time_of_day: morning, afternoon, evening, night
note_en: short English context
Values are made up for demo purposes and not based on real hardware.
License
MIT
testGPT-demonews-online-vnviet-robot-music-scenes
viet-robot-music-scenes
A small synthetic dataset of music scenes and preferences for a home robot.
Columns:
id: row id
scene_id: identifier of the scene
description_vi: description in Vietnamese
description_en: description in English
time_of_day: morning, afternoon, evening, night
mood: calm, focused, energetic, sleep
volume_level: 1–10 suggested volume
Manually authored for demo purposes.
License
MIT
viet-robot-security-events
viet-robot-security-events
A synthetic dataset of simple security-related events in a smart home.
Columns:
id: row id
event_type: door_open, window_open, motion_detected, noise_high, device_disconnected
location: where the event happened
time_hour: integer hour of day (0-23)
is_night: yes/no
human_confirmed: yes/no (whether a human was seen)
note: short English note
For demos only, not real security logs.
License
MIT
