datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GuiaCat
Dataset Card for GuiaCat
Dataset Summary
GuiaCat is a dataset consisting of 5.750 restaurant reviews in Catalan, with 5 associated scores and a label of sentiment. The data was provided by GuiaCat and curated by the BSC.
This work is licensed under a Creative Commons Attribution Non-commercial No-Derivatives 4.0 International License.
Supported Tasks and Leaderboards
This corpus is mainly intended for sentiment analysis.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/GuiaCat.EC-Guide
This repo is only used for dataset viewer. Please download from here.
Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5)
The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.Myanmar-Tuberculosis-Guidelines-Instructions
Myanmar Tuberculosis Guidelines Instructions
A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages.
Authors: Min Si Thu, Khin Myat Noe
Abstract
Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.career-guidance-qa-dataset
Dataset Card for Career Guidance Dataset
Dataset Overview
This dataset provides career guidance information for a variety of career roles. It includes questions and answers related to career roles such as "Data Scientist," "Software Engineer," "Product Manager," and many more. The dataset covers aspects like job responsibilities, required skills, career progression, salary expectations, and work environment. It is intended for use in building chatbot applications for… See the full description on the dataset page: https://huggingface.co/datasets/Pradeep016/career-guidance-qa-dataset.ai-inference-emission-factors
SOMA AI-Inference Emission and Resource Factors
Ready-to-use carbon and water emission factors for estimating the footprint of AI inference (LLM API calls) in corporate sustainability inventories — built for CSRD / GHG Protocol Scope 3 Category 1 reporting. Every factor is derived from primary, cited sources (GPU energy benchmarks, grid carbon intensity registries, datacenter water-use studies); derivations are documented column-by-column below and in full in Supplementary S1 of… See the full description on the dataset page: https://huggingface.co/datasets/GuillermoLlopis/ai-inference-emission-factors.3D-Printable-Guitar-Modelsvcdbmeddocan
Dataset Card for "meddocan"
Dataset Summary
A personal upload of the SPACC_MEDDOCAN corpus. The tokenization is made with the help of a custom spaCy pipeline.
Supported Tasks and Leaderboards
Name Entity Recognition
Languages
More Information Needed
Dataset Structure
Data Instances
More Information Needed
Data Fields
The data fields are the same among all splits.
Data Splits
name
train
validation
test… See the full description on the dataset page: https://huggingface.co/datasets/GuiGel/meddocan.Puzzles_10kair-pollution-population-exposed-to-levels-exceeding-who-guideline-percentage-of-total-africa
Air Pollution Population Exposed to Levels Exceeding WHO Guideline Percentage of Total Africa | Africa (World Health Organization)
Size category: n<1K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/air-pollution-population-exposed-to-levels-exceeding-who-guideline-percentage-of-total-africa.green-books-travel-guides
African American Travel Guides: The Green Book & Companion Directories (1930–1966)
A unified, structured dataset of 113,827 business and lodging listings transcribed from 50 volumes of mid-20th-century African American travel guides, spanning 1930–1966. During the Jim Crow era, these guides told Black travelers which hotels, restaurants, tourist homes, service stations, and other businesses would serve them safely. This dataset brings The Negro Motorist Green Book together with… See the full description on the dataset page: https://huggingface.co/datasets/hadro/green-books-travel-guides.scalared_guidelines
Scalared Guidelines
Description
The following dataset is a collection of 7226 rules and guidelines related to digital design and digital verification. They can be used to create an "AI linter" for digital designs and testbenches.
There are 10 fields in each row:
title: a very short human-readable description of the guideline. Useful mostly for uniquifying the guidelines.
statement: a single sentence of what a user should do to follow the guideline
description: a… See the full description on the dataset page: https://huggingface.co/datasets/Arrakark/scalared_guidelines.career-guidance-qa-dataset
Dataset Card for Career Guidance Dataset
Dataset Overview
This dataset provides career guidance information for a variety of career roles. It includes questions and answers related to career roles such as "Data Scientist," "Software Engineer," "Product Manager," and many more. The dataset covers aspects like job responsibilities, required skills, career progression, salary expectations, and work environment. It is intended for use in building chatbot applications for… See the full description on the dataset page: https://huggingface.co/datasets/Anita018/career-guidance-qa-dataset.techchallenge-animal-condition-datasetMexEmojisThis is a dataset of tweets written in Mexican Spanish and the labels is an emoji describing the emotion.
The region columns is the token for the Mexican state from where the message was written.
The state token is one of the following:
State name
Label
State name
Label
Aguascalientes
Aguascalientes
Mexico
Mexico
Baja California
BC
Nayarit
Nayarit
Baja California Sur
BCS
Nuevo León
NL
Campeche
Campeche
Oaxaca
Oaxaca
Chiapas
Chiapas
Puebla
Puebla
Chihuahua
Chihuahua… See the full description on the dataset page: https://huggingface.co/datasets/guillermoruiz/MexEmojis.massive-guitar-8b0d4e
massive-guitar-8b0d4e
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/wolferussell14/massive-guitar-8b0d4e.gp-long-paragraphspageguide_guide_data
PageGuide Dataset
This repository contains the dataset for PageGuide, a browser extension that assists users in navigating webpages and locating information by grounding LLM answers directly in the HTML DOM.
Project Page: pageguide.github.io
Paper: PageGuide: Browser extension to assist users in navigating a webpage and locating information
Code: github.com/tin-xai/pageguide
Dataset Description
The PageGuide evaluation utilizes several distinct datasets… See the full description on the dataset page: https://huggingface.co/datasets/ttn0011/pageguide_guide_data.guide-to-level-measurement-2021-edition-en-77708_combined_synthetic_dataPEEP
PEEP: Prompts, Extracted Entities with Privacy
Paper: Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
Dataset Summary
PEEP is a multilingual dataset of 15,282 real user queries from the Wildchat dataset, annotated with extracted personal information and associated with synthetic privacy profiles. It is designed to support research on privacy-preserving language models, enabling controlled evaluation of models’ adherence to… See the full description on the dataset page: https://huggingface.co/datasets/guillemram97/PEEP.python_code_summarizationTiktok_Chatgpt_Prompt_Guide
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language, description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/Tiktok_Chatgpt_Prompt_Guide.intent-recognition-biomedicalsource
Career-Guidance
📚 Career Compass Instruction Dataset
The Career Compass Instruction Dataset is a curated set of instruction-response pairs designed to train or evaluate conversational AI systems that provide personalized career guidance. It reflects real queries students may ask and informative answers from the Career Compass AI.
✨ Dataset Summary
This dataset contains high-quality examples of how a conversational AI system can provide career-related guidance. Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/VamshiSurya/Career-Guidance.clinical-guideline-strength-correspondence-v0.1
What this dataset tests
Guideline strength must track evidence strength.
Authority must not exceed data.
Why it exists
Guidelines often harden too early.
Language outruns certainty.
This set checks whether recommendation force matches evidence quality.
Data format
Each row contains
evidence_profile
guideline_recommendation
strength_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
evidence_profile… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-guideline-strength-correspondence-v0.1.clinical-tpib-invariant-guided-next-intervention-prediction-v0.1What this dataset tests
Given a patient’s manifold typepredict the top 3 next interventions that are most coherent.
It rewards
manifold-consistent moves
constraint-aware choices
cross-domain suggestions when warranted
It penalizes
repeating tolerance loops
repeating paradoxical worseners
ignoring contraindications
choosing common care without manifold fit
Labels
coherent_top3
partially_coherent_top3
incoherent_top3
Suggested prompt wrapper
System
You propose the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-tpib-invariant-guided-next-intervention-prediction-v0.1.ai-model-evaluation-guide
AI 模型选型与测评维度词典
版本:1.0.0|更新日期:2026-07-29
Keygate 是覆盖全球主流与前沿 AI 模型的测评、排行榜与选型平台。这份中英双语词典将语言、图像、视频与语音模型比较中常见的 18 项指标整理为结构化字段,帮助读者正确理解榜单、建立选型表,并减少不同资料之间的术语混用。
A bilingual data dictionary of 18 dimensions for evaluating and selecting leading language, image, video and speech models.
配套资料
Keygate 实时排行榜、模型详情与并排对比
GitHub:AI 模型测评与选型维度指南
公开评测基准索引
可下载的评测基准 CSV
数据内容
统一中英文指标名称,减少同一概念被不同译法混用。
明确数值应当“越高越好”还是“越低越好”。
区分输出速度与首段响应时间,避免把两个概念当成同一项。… See the full description on the dataset page: https://huggingface.co/datasets/keygate-ai/ai-model-evaluation-guide.RegTweetsThis is a dataset created from Tweets written in Spanish, particularly from México. The classification task is to chose the Mexican State where the message was written from.
The label correspond to the state given the following table:
State name
Label
State name
Label
Aguascalientes
Aguascalientes
Mexico
Mexico
Baja California
BC
Nayarit
Nayarit
Baja California Sur
BCS
Nuevo León
NL
Campeche
Campeche
Oaxaca
Oaxaca
Chiapas
Chiapas
Puebla
Puebla
Chihuahua
Chihuahua
Querétaro… See the full description on the dataset page: https://huggingface.co/datasets/guillermoruiz/RegTweets.hazop-guidewords-reference-2026
Canonical landing page: https://www.smartqhse.com/datasets/hazop-guidewords-reference-2026
HAZOP Guidewords Reference 2026 — IEC 61882 with Node Examples
HAZOP guidewords reference following IEC 61882:2016 and CCPS guidance. 7 standard guidewords (NO/NOT, MORE, LESS, AS WELL AS, PART OF, REVERSE, OTHER THAN) applied to key process parameters (flow, pressure, temperature, level, composition, reaction, rotation, operating mode) with ~20 worked examples each providing: deviation… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hazop-guidewords-reference-2026.Jennifer-brainwash
