datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
9_Facial_ExpressionsFacial Expressions for YOLO & ViT – 9 Emotions Dataset
Dataset ID: LaurenGurgiolo/9_Facial_Expressions
Task: Facial Emotion Recognition
Domain: Computer Vision, Deep Learning, Affective Computing
Languages: N/A
License: Refer to original Roboflow Universe sources
Dataset Description
The Facial Expressions for YOLO & ViT – 9 Emotions Dataset is a curated facial emotion recognition dataset derived from the Facial Expressions for YOLO Facial Expression Recognition – 9 Emotions Dataset, originally… See the full description on the dataset page: https://huggingface.co/datasets/LaurenGurgiolo/9_Facial_Expressions.synthetic-human-expressions-poses-3d
3D Synthetic Human Poses and FACS Expressions Dataset
This is a high-fidelity synthetic dataset consisting of 10,075 pairs of 3D human character renders and detailed natural language annotations.
Dataset Structure & Generation
To ensure consistency, the dataset is generated using a single base 3D human model. The diversity of the dataset is achieved through a wide range of body poses, facial expressions, and camera angles:
Character: 1 base human model.
Camera… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/synthetic-human-expressions-poses-3d.Micro_Facial_ExpressionsMicro Facial Expressions Dataset
Dataset ID: LaurenGurgiolo/Micro_Facial_Expressions
Task: Facial Emotion Recognition
License: Refer to source datasets
Languages: N/A
Domain: Computer Vision, Affective Computing
Dataset Description
The Micro Facial Expressions dataset is a combined facial emotion recognition dataset composed of two sources:
FER-2013 – a large-scale grayscale facial expression dataset
Micro-Expression Image Dataset – a curated collection of color facial images representing… See the full description on the dataset page: https://huggingface.co/datasets/LaurenGurgiolo/Micro_Facial_Expressions.Southwestern_Mandarin_Dialectal_Expressions_SpeechMRSeg-Referring-Expressions
MRSeg Referring Expressions (single-turn)
Single-turn referring-expression segmentation samples derived from the multi-round MR-Seg data
of SegLLM (paper).
SegLLM's original conversations provide the target regions as input to later rounds, which
makes them unusable for models that receive only an image and text. We used
Qwen3-VL-32B-Instruct to keep only the
samples whose target can be identified from the image and expression alone, rewriting each into
a self-contained… See the full description on the dataset page: https://huggingface.co/datasets/Panorama-grounding/MRSeg-Referring-Expressions.Gene_expressions_UCIboolean_expressions_taskaging-expressions
aging-expressions: Age-Stratified Gene Expression Dataset
Dataset Description
This dataset provides age-stratified gene expression data derived from ARCHS4 and GTEx databases, specifically curated for fine-tuning BulkFormer and other bulk RNA-seq deep learning models. The dataset contains TPM-normalized expression values for protein-coding genes, enriched with comprehensive demographic metadata including precise age information.
Key Features:
🎯 Optimized for BulkFormer:… See the full description on the dataset page: https://huggingface.co/datasets/longevity-genie/aging-expressions.temporal_expressions
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/temporal_expressions.time_expressions_dataset
Dataset Card for Time Expressions Dataset
Dataset Summary
The Time Expressions Dataset is a collection of synthetic data designed for training and evaluating natural language processing (NLP) models on temporal expression recognition and resolution tasks. It contains 378 unique data points, each consisting of a natural language sentence (input_text) and a corresponding JSON-structured output (target_output) that resolves a specific time expression to a standardized date… See the full description on the dataset page: https://huggingface.co/datasets/namesarnav/time_expressions_dataset.labelled-cron-expressions
Labelled Cron Expressions with Human-Readable Descriptions
This dataset provides a collection of common cron expressions, each paired with a clear, human-readable description and a note explaining its typical use case. It serves as a reference for developers to quickly understand or implement scheduling patterns. The accompanying tool allows for looking up these descriptions and parsing cron expressions into their constituent parts.
25 rows · category: reference · licence:… See the full description on the dataset page: https://huggingface.co/datasets/SharkSkin/labelled-cron-expressions.ffhq-female-expressions-batch-2SIGS-symbolic-expressions
SIGS symbolic expression latent corpus
This dataset contains 23,695 grammar-generated symbolic expressions used by the SIGS Grammar-VAE, together with their 32-dimensional latent statistics and variable-based mathematical classes.
Dataset structure
Each row contains:
id: stable row index;
expression: symbolic expression generated by the SIGS grammar;
math_class: one of CONSTANT, TEMPORAL_1D, SPATIAL_1D, SPATIAL_2D, SPATIOTEMPORAL_2D, or SPATIOTEMPORAL_3D;
has_x… See the full description on the dataset page: https://huggingface.co/datasets/oroikono/SIGS-symbolic-expressions.indonesian-numbers-expressions
Ungkapan Angka Indonesia 🔢
Angka ga cuma buat hitung — di bahasa Indonesia, angka jadi bagian idiom. Setengah hati, dua muka, seribu satu alasan. Dataset ini ngumpulin ungkapan-ungkapan angka yang dipakai orang Indonesia sehari-hari.
Kenapa dataset ini ada?
LLM sering salah artiin ungkapan angka secara literal — dua muka bukan dua wajah, setengah hati bukan separuh jantung. Dataset ini bantu model paham makna kiasan. Belum ada dataset ungkapan angka bahasa… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-numbers-expressions.assetbot-labelled-cron-expressions
Labelled Cron Expressions with Human-Readable Descriptions
This dataset provides a collection of common cron expressions, each paired with a clear, human-readable description and a note explaining its typical use case. It serves as a reference for developers to quickly understand or implement scheduling patterns. The accompanying tool allows for looking up these descriptions and parsing cron expressions into their constituent parts.
Get it here
math-expressionseng_latn_temporal_expressions
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/eng_latn_temporal_expressions.gsm8k-parsed-expressionsffhq-female-expressions-batch-1arithmetic-expressions-with-parentheses-v2minimal-arithmetic-expressionscf-ruleset-expressionsThis dataset includes valid Cloudflare Ruleset expressions with plausible user input used to generate the expressions. The user input is included in 2 formats: "formal" and "internet". The former is precise and uses proper spelling, whereas the latter is all lowercase, might omit information and resembles your usual social media spelling.
The expressions have a maximum complexity of 5, where complexity refers to the number of fields present in the expression. The expressions were generated… See the full description on the dataset page: https://huggingface.co/datasets/deathbyknowledge/cf-ruleset-expressions.Idiomatic-Expressions
🇰🇿 Kazakh Idiomatic Expressions and Cultural Context
📖 Overview
This dataset is a specialized linguistic resource containing 413 samples focused on the interpretation of Kazakh idioms, phraseology, and cultural metaphors. It is designed to help models move beyond literal translation and understand the deep-seated figurative meanings (идиомалар мен тұрақты тіркестер) used in daily Kazakh communication and literature.
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Idiomatic-Expressions.dog-expressions
Dog-Expressions
Made with ❤️ using 🎨 NeMo Data Designer
This dataset is a test dataset for dog expressions
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset (use trust_remote_code=True for image columns to display)
dataset = load_dataset("nabinnvidia/dog-expressions", "data", split="train", trust_remote_code=True)
df = dataset.to_pandas()
Image columns (gpt-image-1.5-image) are loaded as Sequence(Image()) via the dataset loading script so the dataset… See the full description on the dataset page: https://huggingface.co/datasets/nabinnvidia/dog-expressions.Micro_Facial_ExpressionsMicro Facial Expressions Dataset
Dataset ID: LaurenGurgiolo/Micro_Facial_Expressions
Task: Facial Emotion Recognition
License: Refer to source datasets
Languages: N/A
Domain: Computer Vision, Affective Computing
Dataset Description
The Micro Facial Expressions dataset is a combined facial emotion recognition dataset composed of two sources:
FER-2013 – a large-scale grayscale facial expression dataset
Micro-Expression Image Dataset – a curated collection of color facial images representing… See the full description on the dataset page: https://huggingface.co/datasets/thesomdeep/Micro_Facial_Expressions.ffhq-female-expressions-batch-3BanBan_2024-10-17-facial_expressions
Dataset Card for "BanBan_2024-10-17-facial_expressions"
將asadfgglie/BanBan_2024-10-17-Pure_text_raw_data用GPT4o-mini與人工標記後的資料集
arithmetic-expressions-with-parenthesesnifer-facial-expressionsexpressions_langage_populaire_Bresse_Louhannaise
[!NOTE]
Dataset origin: https://archive.org/details/dictionnairepato00guiluoft/page/n7/mode/2up
