datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
job-descriptionsjob_titles_and_descriptions
IT Job Roles, Skills, and Descriptions Dataset
This dataset provides detailed information about various IT job roles, the required skills for each role, and the job descriptions that outline the responsibilities and qualifications associated with each position. It is designed for use in applications such as career guidance systems, job recommendation engines, and educational tools aimed at aligning skills with industry demands.
Dataset Upload
This dataset is uploaded… See the full description on the dataset page: https://huggingface.co/datasets/NxtGenIntern/job_titles_and_descriptions.recruitment-dataset-job-descriptions-english
Djinni Dataset (English Job Descriptions part)
Overview
The Djinni Recruitment Dataset (English Job Descriptions part) contains 150,000 job descriptions and 230,000 anonymized candidate CVs, posted between 2020-2023 on the Djinni IT job platform. The dataset includes samples in English and Ukrainian.
The dataset contains various attributes related to job descriptions, including position titles, job descriptions, company names, experience requirements, keywords, English… See the full description on the dataset page: https://huggingface.co/datasets/lang-uk/recruitment-dataset-job-descriptions-english.mls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.UTD-descriptions
📝 UTD‑descriptions Dataset
The UTD‑descriptions dataset provides multiple kinds of textual descriptions for video samples belonging to 12 widely used video understanding datasets (e.g., Kinetics‑400, UCF101, HMDB51, DiDeMo, ActivityNet, MSR‑VTT, Charades, etc.).It contains no video files — instead, it offers captions, attributes, and metadata that correspond to videos stored in their original datasets.
This dataset is ideal for video captioning, multimodal learning, video–language… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/UTD-descriptions.voxpopuli-mls-de-descriptions
Natural Language Voice Descriptions of the VoxPopuli and MLS German Datasets
German read and parliamentary speech paired with its transcript, acoustic
measurements, discrete German descriptor tags, and a free-text German
description of the speaker's voice and recording conditions. The dataset is intended
for training description-conditioned TTS models such as
Parler-TTS.
The data was built as part of research work. It is a random subset of the
pooled German portions of VoxPopuli… See the full description on the dataset page: https://huggingface.co/datasets/leonhard-behr/voxpopuli-mls-de-descriptions.galaxy-descriptions
Galaxy Descriptions
Project Page | Code
This dataset provides galaxy cutout images, natural-language descriptions, text embeddings, and image embeddings for galaxies drawn from multiple imaging surveys (specifically Legacy DR10 and HSC PDR3 Wide).
Each row corresponds to a single galaxy and contains:
A preprocessed RGB galaxy image
A caption generated by gpt-4.1-mini
A single-sentence summary of the caption generated by gpt-4.1-nano
Text embeddings for the caption and summary… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/galaxy-descriptions.circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.wikidata-en-descriptionslibritts-r-filtered-speaker-descriptions
Dataset Card for Annotated LibriTTS-R
This dataset is an annotated version of a filtered LibriTTS-R [1].
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 960 hours of read English speech at 24kHz sampling rate, published in 2019.
In the text_description column, it provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts-r-filtered-speaker-descriptions.robot_descriptionsamazon-product-descriptions-vlm
Amazon Multimodal Product dataset
This is a modfied and slim verison of bprateek/amazon_product_description helpful to get started training multimodal LLMs.
The description field was generated used Gemini Flash.
listing-descriptions
Listing-Descriptions
Made with ❤️ using 🦥 Unsloth Studio
Listing descriptions dataset was generated with Unsloth Recipe Studio. It contains 280 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("standrey/listing-descriptions", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 280
📋 Columns: 2
✅ Completion: 56.0% (500 requested)
📋 Schema & Statistics… See the full description on the dataset page: https://huggingface.co/datasets/standrey/listing-descriptions.manipulation-init-frame-descriptions
Manipulation init-frame / description pairs
1000 (init frame image, task description) pairs randomly sampled
(seed=42) from the manipulation task family of
nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.
Each row is the first frame of a simulated manipulation clip paired with its
task description text, drawn from the following generators within the
manipulation task family: DreamZero, MimicGen (AgiBot G1 / Fourier GR-1 /
Galbot G1), and Simulario.
Columns:… See the full description on the dataset page: https://huggingface.co/datasets/khang123452/manipulation-init-frame-descriptions.goodreads-book-descriptions
Goodreads Book Descriptions
A dataset of English book titles and descriptions from Goodreads.
The original dataset has 2.3 million books total with many more fields.
There may exist a small number of non-English books in this dataset.
Citations
Mengting Wan, Julian McAuley, "Item Recommendation on Monotonic Behavior Chains", in RecSys'18.
Mengting Wan, Rishabh Misra, Ndapa Nakashole, Julian McAuley, "Fine-Grained Spoiler Detection from Large-Scale Review Corpora", in… See the full description on the dataset page: https://huggingface.co/datasets/booksouls/goodreads-book-descriptions.book_titles_and_descriptionsmac-app-store-apps-descriptions
Dataset Card for Macappstore Applications Descriptions
📌 Dataset status: static snapshot (no scheduled updates). This dataset is derived from the December 2023 – January 2024 Mac App Store metadata snapshot and reflects the store as of that period. The dataset is stable and remains available for research use; it is not refreshed on a schedule.
Mac App Store Applications descriptions extracted from the metadata from the public API.
Curated by: MacPaw Way Ltd.
Language(s)… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/mac-app-store-apps-descriptions.vietnamese-job-descriptions
💼 Tinix Vietnam Job Description
1. 📌 Giới Thiệu Tinix Vietnam Job Description
Tinix Vietnam Job Description là bộ dữ liệu tuyển dụng tiếng Việt ở định dạng CSV, gồm các tin tuyển dụng có cấu trúc về chức danh, công ty, mức lương, địa điểm, loại hợp đồng, ngành nghề, yêu cầu kinh nghiệm, trình độ học vấn, mô tả công việc, phúc lợi, yêu cầu ứng viên và năm đăng tin.
Bộ dữ liệu được thiết kế cho các bài toán NLP và phân tích thị trường lao động tại Việt Nam, đặc biệt trong… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnamese-job-descriptions.wikidata-en-descriptions-smallpacman_descriptionslisting-descriptions-unsloth
Listing-Descriptions-Unsloth
Made with ❤️ using 🦥 Unsloth Studio
listing_descriptions_dataset_full was generated with Unsloth Recipe Studio. It contains 1,654 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("standrey/listing-descriptions-unsloth", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 1,654
📋 Columns: 2
✅ Completion: 51.7% (3,200 requested)… See the full description on the dataset page: https://huggingface.co/datasets/standrey/listing-descriptions-unsloth.book_titles_and_descriptions_en_cleanblind-people-scene-descriptions
Merged Navigation-Focused Image Caption Dataset
This dataset is a combination and filtered version of two publicly available image captioning datasets, specifically curated to focus on images and captions relevant to navigation and scene understanding.
Source Datasets
This dataset is derived from the following two sources:
COCO Captions (jxie/coco_captions)
Original Hugging Face Hub ID: jxie/coco_captions
Link: https://huggingface.co/datasets/jxie/coco_captions
Original… See the full description on the dataset page: https://huggingface.co/datasets/mlevytskyi/blind-people-scene-descriptions.HumanML3D-500ms-FPP-descriptions-CoTs-1
HumanML3D 500ms First person perspective descriptions for CoTs
Introduction
This repository contains files of the Mr. Ri's and Ms. Tique's HumanML3D human motion dataset,
but also descriptions of the movements in first person perspective in 0.5 second time windows.
The descriptions were created synthetically with use of a multimodal LLM and are in json format. They can be found in comics_and_descriptions folder.
The dataset also contains motion capture data and… See the full description on the dataset page: https://huggingface.co/datasets/Wojtekb30/HumanML3D-500ms-FPP-descriptions-CoTs-1.job-descriptions-dataset-minimtg-scryfall-unique-artwork-20240809-with-card-art-descriptions-and-images-with-embeddingsus-patent-descriptions
US Patent Descriptions
This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView.
Splits
train: 10,000 rows for model training
validation: 2,500 rows for validation
test: 2,500 rows for evaluation
Columns
patent_id: Identifier for the patent; useful for reconciling with other PatentsView datasets
description_text: Full… See the full description on the dataset page: https://huggingface.co/datasets/mhurhangee/us-patent-descriptions.job-titles-descriptions
Synthetic Job Descriptions Dataset
A high-quality synthetic expansion of gpriday/job-titles — 65,248 structured job descriptions with a pre-computed FAISS index for semantic search.
Overview
This dataset provides comprehensive, synthetically generated job descriptions for all role titles found in the gpriday/job-titles source dataset. Each entry is structured for consistency and designed to support NLP pipelines, career-tech applications, and Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/ismailcemsahin/job-titles-descriptions.LLM-generated-emoji-descriptions
Emoji Metadata Dataset
Overview
The LLM Emoji Dataset is a comprehensive collection of enriched semantic descriptions for emojis, generated using Meta AI's Llama-3-8B model. This dataset aims to provide semantic context for each emoji, enhancing their usability in various NLP applications, especially those requiring semantic search. The LLM Emoji Dataset was used to build a multilingual search engine for emojies, which you can interact with using this online Streamlit… See the full description on the dataset page: https://huggingface.co/datasets/badrex/LLM-generated-emoji-descriptions.SC-train-valid-test_SDG-Descriptionsnli-label:
(0) entailment
(2) contradiction
