datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
function_calling_v3_SAMPLE
Trelis Function Calling Dataset - VERSION 3 - SAMPLE
This is a SAMPLE of the v3 dataset available for purchase here.
Features:
Allows models to be fine-tuned for function-calling.
The dataset is human generated and does not make use of Llama 2 or OpenAI!
The dataset includes 66 training rows, 19 validation rows and 5 test rows (for manual evaluation).
Based on eight functions: search_bing, search_arxiv, save_chat, read_json_file, list_files, get_current_weather, delete_file… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/function_calling_v3_SAMPLE.Trustpilot-Reviews-Dataset-20K-Sample
Trustpilot Reviews Dataset – 20K Sample
This dataset contains a curated sample of 20,000 English-language user reviews sourced exclusively from Trustpilot.com. It is a representative subset of our larger collection containing over 1 million Trustpilot reviews across various industries and companies.
🗂️ Dataset Overview
Source: Trustpilot
Total Records: 20,000
Language: English
Industries: E-commerce, SaaS, Travel, Finance, Education, and more
Use Case: NLP tasks… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Trustpilot-Reviews-Dataset-20K-Sample.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.arabic-palestinian-levantine-sample
4FACTORS — Palestinian Levantine Conversational Sample
50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data.
What this is
Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.interview_QA_sample_set
Huy Interview Instruction Dataset
Dataset Description
This is an instruction-answer dataset for fine-tuning conversational AI models to answer interview-style questions based on a personal CV/profile.
The dataset has two columns: instruction and answer.
The dataset contains 5,100 instruction-answer pairs.
Data Creation
This dataset was created using a GenAI-assisted pipeline. A personal CV/profile was provided as source material, and GenAI was used… See the full description on the dataset page: https://huggingface.co/datasets/dinhxuanhuy/interview_QA_sample_set.ceo-quotes-verified-sample
🎙️ CEO Transcripts — Verified Executive Interviews
The World's Largest Database of Verified C-Suite Transcripts
20,000+ Executives · 100,000+ Transcripts · 400,000+ Quotes · S&P 500 + NASDAQ + Global Leaders
🔥 What's In This Sample?
This is a free evaluation sample from CEOInterviews.ai featuring 9 of the most market-moving voices in finance, tech, and policy.
Executive
Role
Why They Matter
Jensen Huang
CEO, NVIDIA
Every AI… See the full description on the dataset page: https://huggingface.co/datasets/codelucas/ceo-quotes-verified-sample.wildchat-stratified-sample
WildChat Stratified Sample
Dataset Description
This dataset contains a stratified sample of 263 GPT-4 conversations (347 total turns) from the WildChat dataset. The sample was carefully selected to ensure balanced representation across conversation turn positions and user message lengths.
Dataset Summary
Total Conversations: 263
Total Turns/Rows: 347
Average Turns per Conversation: 1.32
Conversation Length: 1-5 turns (conversations with >5 turns excluded)… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/wildchat-stratified-sample.PubMed-Bilingual-Medical-Sample-EN-ID
🏥 PubMed Bilingual Medical Sample (EN-ID)
Providing premium, high-quality English-Indonesian bilingual medical datasets for AI, NLP, and Machine Learning research.
📥 Download Free Sample
You can directly access and download the free dataset sample (.csv format) from our repository files here:
⬇️ Download Free Sample File (tree/main)
🚀 Upgrade to Full Version (Volume 1)
This repository contains a free sample of our industry-grade parallel… See the full description on the dataset page: https://huggingface.co/datasets/IndoHealth-NLP/PubMed-Bilingual-Medical-Sample-EN-ID.arabic-egyptian-sample
4FACTORS Arabic — Egyptian Q&A Sample
Conversational question–answer pairs in spoken Egyptian Arabic, written by a
first-language Egyptian speaker. A 50-item demonstration sample, with English
glosses, released under CC BY-NC 4.0.
This is the Egyptian variety in the 4FACTORS Arabic sample set, alongside the
Palestinian Levantine
and Modern Standard Arabic (MSA) sets.
What this is
Fifty short question–answer exchanges of the kind that come up in everyday life —… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-egyptian-sample.arabic-msa-sample
Arabic — Modern Standard Arabic (MSA) Sample
Native-written, human-verified Modern Standard Arabic. No scraping. No machine
translation. No synthetic generation. Every sentence written from scratch by a
first-language speaker in formal news / official-statement register, then reviewed
line by line against a written checklist and measured for structural diversity across
the whole set.
A public demonstration sample (50 items). Larger MSA datasets and other varieties
(Levantine… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-msa-sample.Risk_Factor_Disclosure_SampleDataset
📊 Sample Preview – Risk Factor Disclosure Dataset v1.0
👉 This is a preview sample (100 records) of the full Risk Factor Disclosure Dataset v1.0.🔗 To access the full dataset (1,869 enriched risk disclosures), visit:https://asapworks.gumroad.com/l/jbxtfd
📦 About the Sample File
This sample contains 100 enriched Item 1A "Risk Factor" disclosures extracted from 10-K filings submitted by top public companies between 2010 and 2024.
Each row represents a structured risk… See the full description on the dataset page: https://huggingface.co/datasets/asapworks/Risk_Factor_Disclosure_SampleDataset.SampleTest
