datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hubbench
HubBench 1.4.0
One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.function_calling_v3_SAMPLE
Trelis Function Calling Dataset - VERSION 3 - SAMPLE
This is a SAMPLE of the v3 dataset available for purchase here.
Features:
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
The dataset includes 66 training rows, 19 validation rows and 5 test rows (for manual evaluation).
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_v3_SAMPLE.Trustpilot-Reviews-Dataset-20K-Sample
Trustpilot Reviews Dataset – 20K Sample
This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com. It is a representative subset of our larger collection containing over 1 million Trustpilot reviews across various industries and companies.
🗂️ Dataset Overview
Source: Trustpilot
Total Records: 20,000
Language: English
Industries: E-commerce, SaaS, Travel, Finance, Education, and more
Use Case: NLP tasks… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Trustpilot-Reviews-Dataset-20K-Sample.MedPRESS_Benchmark
MedPRESS
MedPRESS is a multi-turn benchmark for evaluating patient-pressure-induced medical sycophancy in large language models. The dataset tests whether a model maintains a safe medical stance when a user repeatedly pressures it toward an unsafe or false health belief.
The benchmark is designed around five-turn conversations. Each row contains one medical scenario, the unsafe or false belief being pressured, the expected safe stance, and five progressively stronger user turns.… See the full description on the dataset page: https://huggingface.co/datasets/samanjoy2/MedPRESS_Benchmark.arabic-palestinian-levantine-sample
4FACTORS — Palestinian Levantine Conversational Sample
50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data.
What this is
Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.interview_QA_sample_set
Huy Interview Instruction Dataset
Dataset Description
This is an instruction-answer dataset for fine-tuning conversational AI models to answer interview-style questions based on a personal CV/profile.
The dataset has two columns: instruction and answer.
The dataset contains 5,100 instruction-answer pairs.
Data Creation
This dataset was created using a GenAI-assisted pipeline. A personal CV/profile was provided as source material, and GenAI was used… See the full description on the dataset page: https://huggingface.co/datasets/dinhxuanhuy/interview_QA_sample_set.clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1Clarus Clinical Quad Coupling PK Integrity v0.1
PurposeDetect PK integrity distortion driven by four interacting nodes.
Quad nodes
Sampling window deviation
Bioanalytical or stability variance
Dose adjustment decisions
Governance interim or submission timing
InputOne vignette.
OutputStrict JSON only.
Required keys
pk_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py
Run scoringCreate… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-pk-sampling-window-deviation-bioanalytical-variance-dose-adjustment-interim-v0.1.ceo-quotes-verified-sample
🎙️ CEO Transcripts — Verified Executive Interviews
The World's Largest Database of Verified C-Suite Transcripts
20,000+ Executives · 100,000+ Transcripts · 400,000+ Quotes · S&P 500 + NASDAQ + Global Leaders
🔥 What's In This Sample?
This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy.
Executive
Role
Why They Matter
Jensen Huang
CEO, NVIDIA
Every AI… See the full description on the dataset page: https://huggingface.co/datasets/codelucas/ceo-quotes-verified-sample.sam_altman_essays
Dataset Card for Sam Altman Essay Collection Dataset
Dataset Description
This dataset contains a complete collection of essays written by Sam Altman, an entrepreneur, investor, and former president of Y Combinator. The essays cover a wide range of topics including startups, technology, artificial intelligence, leadership, and personal growth. Each essay has been cleaned and processed to extract the title, date of publication, and the full text content.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/sgoel9/sam_altman_essays.PubMed-Bilingual-Medical-Sample-EN-ID
🏥 PubMed Bilingual Medical Sample (EN-ID)
Providing premium, high-quality English-Indonesian bilingual medical datasets for AI, NLP, and Machine Learning research.
📥 Download Free Sample
You can directly access and download the free dataset sample (.csv format) from our repository files here:
⬇️ Download Free Sample File (tree/main)
🚀 Upgrade to Full Version (Volume 1)
This repository contains a free sample of our industry-grade parallel… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/PubMed-Bilingual-Medical-Sample-EN-ID.tst
The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset
This data set is based on the judgments of "Offenses of Fraudulence" cases published by the Judicial Yuan. The data range of the dataset is from January 1, 2011, to December 31, 2021. 74,823 pieces of original data (judgments and rulings) were collected. We only took the contents of the "criminal facts" field of the judgment. This dataset is divided into three parts. The training dataset has 59,858… See the full description on the dataset page: https://huggingface.co/datasets/Sama1030/tst.wildchat-stratified-sample
WildChat Stratified Sample
Dataset Description
This dataset contains a stratified sample of 263 GPT-4 conversations (347 total turns) from the WildChat dataset. The sample was carefully selected to ensure balanced representation across conversation turn positions and user message lengths.
Dataset Summary
Total Conversations: 263
Total Turns/Rows: 347
Average Turns per Conversation: 1.32
Conversation Length: 1-5 turns (conversations with >5 turns excluded)… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/wildchat-stratified-sample.arabic-egyptian-sample
4FACTORS Arabic — Egyptian Q&A Sample
Conversational question–answer pairs in spoken Egyptian Arabic, written by a
first-language Egyptian speaker. A 50-item demonstration sample, with English
glosses, released under CC BY-NC 4.0.
This is the Egyptian variety in the 4FACTORS Arabic sample set, alongside the
Palestinian Levantine
and Modern Standard Arabic (MSA) sets.
What this is
Fifty short question–answer exchanges of the kind that come up in everyday life —… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-egyptian-sample.arabic-msa-sample
Arabic — Modern Standard Arabic (MSA) Sample
Native-written, human-verified Modern Standard Arabic. No scraping. No machine
translation. No synthetic generation. Every sentence written from scratch by a
first-language speaker in formal news / official-statement register, then reviewed
line by line against a written checklist and measured for structural diversity across
the whole set.
A public demonstration sample (50 items). Larger MSA datasets and other varieties
(Levantine… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-msa-sample.verse-anvaya-rigvedaRisk_Factor_Disclosure_SampleDataset
📊 Sample Preview – Risk Factor Disclosure Dataset v1.0
👉 This is a preview sample (100 records) of the full Risk Factor Disclosure Dataset v1.0.🔗 To access the full dataset (1,869 enriched risk disclosures), visit:https://asapworks.gumroad.com/l/jbxtfd
📦 About the Sample File
This sample contains 100 enriched Item 1A "Risk Factor" disclosures extracted from 10-K filings submitted by top public companies between 2010 and 2024.
Each row represents a structured risk… See the full description on the dataset page: https://huggingface.co/datasets/asapworks/Risk_Factor_Disclosure_SampleDataset.blindspots
Blind Spots of tencent/Youtu-LLM-2B-Base
Overview
This dataset documents failure cases ("blind spots") observed when testing the base language model tencent/Youtu-LLM-2B-Base. The dataset contains prompts where the model produced incorrect answers, incorrect reasoning, or failed to follow instructions.
The goal of this dataset is to highlight weaknesses in small frontier models and provide examples that could be used for evaluation or future fine-tuning.
Model… See the full description on the dataset page: https://huggingface.co/datasets/Samri-g/blindspots.SampleTest
