datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
energy-consumption-hourly-spainenergy-consumption-weather-hourly-spainCLIP-ViT-L-14-336-L20-features
OpenAI/CLIP-ViT-L/14@336 Layer 20 features, CLIP+BLIP labels
Feature activation max visualization of the 4096 Features @ L20
CLIP+BLIP labels (may or may not describe what a neuron truly encodes!)
⚠️ May contain sensitive images, albeit abstract. Use responsibly!
Examples:
osworld_tasks_filessynthetic-fraud-detectionbridgev2-vita-toykitchen-manifests
BridgeV2 VITA ToyKitchen-like Manifests
This repository contains manifest files for a reconstructed VITA-style BridgeV2 ToyKitchen-like pick-and-place subset.
Source dataset
The source dataset is:
Gaugou/BridgeV2
This repository does not duplicate the original BridgeV2 videos. It provides episode IDs and metadata for selecting the subset from the source dataset.
Split
Train: 2,986 episodes
Test: 287 episodes
Total selected: 3,273 episodes
Selection… See the full description on the dataset page: https://huggingface.co/datasets/praedico/bridgev2-vita-toykitchen-manifests.csgoheloc
HELOC (Home Equity Line of Credit)
The HELOC dataset from FICO.
Each entry in the dataset is a line of credit, typically offered by a bank as a percentage of home equity (the difference between the current market value of a home and its purchase price).
The customers in this dataset have requested a credit line in the range of $5,000 - $150,000.
The fundamental task is to use the information about the applicant in their credit report to predict whether they will repay their HELOC… See the full description on the dataset page: https://huggingface.co/datasets/vitaliykinakh/heloc.sick
Thyroid Disease Dataset
Please refer to original source for more details.
Dataset Description
The Thyroid Disease dataset comprises medical records related to thyroid conditions. Supplied by the Garavan Institute and J. Ross Quinlan from the New South Wales Institute, Sydney, Australia in 1987, this dataset is widely used for diagnosing thyroid disorders. It contains 3,772 instances with 30 attributes, including both continuous and discrete features.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/vitaliykinakh/sick.clinical-vital-sign-escalation-response-coherence-risk-v0.1What this repo is for
Detect when
vital signs deteriorate
but escalation and response
fail
Common breaks
high NEWS not escalated
doctor notified late
review done but no action
action taken too late
Used for
deterioration detection
ward safety
ICU outreach
rapid response auditing
africa-synth-cancer-cancer-mortality-vital-registration-all
Cancer Mortality Vital Registration | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cancer-cancer-mortality-vital-registration-all.fmcw-vital-signswikipedia-id-qna
Wikipedia ID Synthetic QnA
This dataset contains synthetic question-answer pairs (QnA) generated using DeepSeek from Indonesian Wikipedia articles. The data has been sourced from this Wikipedia dataset, which contains a subset of Indonesian Wikipedia articles. Each entry includes a context, a related question and answer pair, and an unrelated question.
Dataset Structure
The dataset contains the following columns:
id: A unique identifier for each row.
context: A… See the full description on the dataset page: https://huggingface.co/datasets/vitoghif/wikipedia-id-qna.VITHSD
Dataset Card for VITHSD
1. Dataset Summary
VITHSD (Vietnamese Targeted Hate Speech Detection) contains 10,000 Vietnamese social‐media comments annotated for hate toward five target categories:
individual
groups
religion/creed
race/ethnicity
politics
Each target is labeled on a 3‐point scale (e.g., 0 = no hate, 1 = offensive, 2 = hateful). In this unified version, all splits are combined into one CSV with an extra type column indicating train / dev / test.
2.… See the full description on the dataset page: https://huggingface.co/datasets/visolex/VITHSD.ko_gpt4omini_note_15.4k
한국어 메모 데이터셋
GPT-4o-mini를 통해 생성된 한국어 메모 데이터셋입니다.
대주제(main_topic), 소주제(sub_topic)를 통해 메모처럼 보이는 데이터를 생성하였습니다.
vithsd
Dataset Card for Dataset Name
ViTHSD: Vietnamese Targeted Hate Speech Detection
Dataset Details
A new version of Vietnamese Hate Speech Detection with target-oriented labels. Each comment can have multiple targets, each target has a hatred level indicating the hateful: CLEAN, OFFENSIVE, and HATE.
The targets are: Individuals, Groups, Religion/creed, Race/ethnicity, and Politics
Uses
Directly load and use the dataset from hugging face:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/sonlam1102/vithsd.Squad_PT
Dataset Card para o SQuAD 1.1 em Português Brasil
O conjunto de dados "Stanford Question Answering Dataset" (SQuAD),
para tarefa de perguntas e respostas extrativas, foi desenvolvido em 2016. Ele utiliza perguntas geradas a partir de
536 artigos da Wikipedia* com mais de 100.000 linhas de dados. É construído na forma de uma pergunta e um contexto dos artigos da
Wikipedia contendo a resposta à pergunta. [1]Originalmente este dataset foi construído no idioma inglês, contudo, o grupo… See the full description on the dataset page: https://huggingface.co/datasets/vitorandrade/Squad_PT.clinical-deterioration-vitals-escalation-coherence-risk-v0.1What this repo is for
Detect when
vital signs show deterioration
but escalation
does not happen
or happens too late
before
ICU transfer
cardiac arrest
or avoidable harm.
gen_image_wordnet_preferencesThis dataset contains generated images. See the associated Hugging Face Collection for examples and additional details: Generated Image Wordnet
TextToCodevi_term_definitionuser-prompt-responsetravelSource
TestAIFB
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/VitoVoelker/TestAIFB.clinical_trial_patient_vitals
