datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
healthbench-professionalContains the data for the HealthBench Professional eval.
Each example contains:
conversation: list of user / assistant messages, ending in a user message
rubric_items: list of rubric items, each containing criterion_text and points
use_case: one of consult, writing, or research
type: one of good_faith or red_teaming
difficulty: physician-assigned difficulty rating (difficult for Likert 1-2, typical for Likert 3-7)
specialty: medical specialty or sub-specialty
physician_response: response… See the full description on the dataset page: https://huggingface.co/datasets/openai/healthbench-professional.mmlu_professional_medicineCreative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.taiwan-professional-exams-115-2Machine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2.SlimPajama-Meta-rater-Professionalism-30B
Top 30B token SlimPajama Subset selected by the Professionalism rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.task728_mmmlu_answer_generation_professional_accounting
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task728_mmmlu_answer_generation_professional_accounting
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task728_mmmlu_answer_generation_professional_accounting.Arabic-professional-voice
Arabic Professional Voice
A high-quality, single-speaker Arabic Text-to-Speech (TTS) dataset recorded by a professional speaker. All transcriptions include full Tashkeel (diacritical marks), making it directly suitable for training neural TTS systems without additional text normalization.
Dataset Summary
Property
Value
Language
Arabic — Modern Standard Arabic (MSA)
Utterances
439
Speaker
1 (professional male speaker)
Sampling Rate
16 kHz
Format
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/Arabic-professional-voice.task731_mmmlu_answer_generation_professional_psychology
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task731_mmmlu_answer_generation_professional_psychology
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task731_mmmlu_answer_generation_professional_psychology.taiwan-professional-exams-115-2-resultsBenchmark results produced by any-to-bench. One subset here is one taker
configuration — a single model at a single reasoning effort — sat against the
exams in another dataset repo. Every row names the exam repo and subset it was
earned against, so results from several corpora, and from several people, can
live side by side.
results-index.json — the catalog: one headline row per configuration
results-<entry>/entry.json — that configuration's per-paper scores
results-<entry>/raw/<subset>/ —… See the full description on the dataset page: https://huggingface.co/datasets/skyhong2002/taiwan-professional-exams-115-2-results.mmlu-professional_medicine-arabicprofessional-apps-grounding-with-filteringeasyr1-10k-hard-qwen7b-easy-gta1-4MP-professional-apps-grounding-only-no-resolution-in-prompt
easyr1-10k-hard-qwen7b-easy-gta1-4MP-professional-apps-grounding-only-no-resolution-in-prompt
This dataset was generated using the EasyR1 grounding dataset pipeline.
Generation Details
Generated on: 2025-08-26 12:16:32 UTC
Script: push_easyr1_to_hf.py
Data directory: /lustre/fsw/portfolios/nvr/users/aawadalla/LLaMA-Factory/data
Parameters Used
Maximum samples: 10000
Image resize (max megapixels): 4.0 MP
Minimum native image resolution: 0.0 MP
Prompt format:… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-10k-hard-qwen7b-easy-gta1-4MP-professional-apps-grounding-only-no-resolution-in-prompt.africa-synth-education-professional-development-nigeria
Nigeria Education - Professional Development | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Education… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-professional-development-nigeria.mmlu-professional_psychology-neg
Dataset Card for "mmlu-professional_psychology-neg"
More Information needed
mmlu-professional_law
Dataset Card for "mmlu-professional_law"
More Information needed
professional-image-editing-dataset-sample
Professional Image Editing Dataset for AI Training — Sample
Professional human-edited visual data for training, fine-tuning and evaluating image-editing and generative AI models.
FixThePhoto produces image-editing training datasets based on professional retouching and visual-production workflows.
Typical data structures can include:
Source → Human-Edited Target
Source → Instruction → Human-Edited Target
AI Output → Human Correction
Masks and alpha mattes
Editing metadata
Human… See the full description on the dataset page: https://huggingface.co/datasets/FixThePhoto/professional-image-editing-dataset-sample.professional-licensing-continuing-education-by-jurisdiction
CPA continuing professional education (CPE) requirements by US jurisdiction
Canonical, always-current version: https://referencesource.org/professional-licensing-continuing-education-by-jurisdiction/
Machine-readable: https://referencesource.org/professional-licensing-continuing-education-by-jurisdiction/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-12
Stale after: 2027-02-08 (past this date, prefer the canonical copy —
it re-verifies on a cadence this… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/professional-licensing-continuing-education-by-jurisdiction.mmlu-professional_law-neg
Dataset Card for "mmlu-professional_law-neg"
More Information needed
rollout_professional_maya_dagger_r1_20260716_191317This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos",
"left_joint_6.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/collected-ai/rollout_professional_maya_dagger_r1_20260716_191317.bangladesh-law-professional
🇧🇩 Bangladesh Law Professional Dataset
A clean, instruction-tuned (Alpaca-style) question–answer dataset for
fine-tuning language models on Bangladesh law, in Bangla and English.
👤 Author & Contribution
Curated & built by
Sadat Sami (@Sadatsami)
Role
Dataset architect — collected, cleaned, filtered, reformatted and published
Motivation
Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.philippine_driving_professional_exam_2023ProfessionalTermsAmharicEnglishProfessional-Profiles
Resume Dataset
This dataset comprises resume data aggregated from a variety of online sources, including professional networking platforms, job portals, company career pages, and personal portfolio websites. The collection period spans from 2020 to 2025, ensuring that the dataset reflects contemporary trends in career trajectories, skill sets, and educational backgrounds across diverse industries.
Data Collection
Timeframe: 2020 to 2025
Data Sources: Resumes have been… See the full description on the dataset page: https://huggingface.co/datasets/popoy-salvador/Professional-Profiles.mmlu-professional-medicineProfessional_and_Hobby_Topicsrollout_professional_maya_dagger_r1_20260716_195156This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos",
"left_joint_6.pos",
"left_gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/collected-ai/rollout_professional_maya_dagger_r1_20260716_195156.mmlu-professional-medicineprofessional-chatgpt-prompts🧠 Awesome ChatGPT Prompts [CSV dataset]
This is a Dataset Repository of Awesome ChatGPT Prompts
View All Prompts on GitHub
License
CC-0
