datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.ml-interview-examples-movielens-1mthai-investment-consultant-licensing-exams
Thai Public Investment Consultant (IC) Exams Dataset
Overview
This dataset comprises a collection of exam questions and answers from the Thai Public Investment Consultant (IC) Examinations. It's a valuable resource for developing and evaluating question-answering systems in the finance sector.
Dataset Source
The Stock Exchange of Thailand (SET)
Maintainer
Dr. Kobkrit Viriyayudhakorn
Email: kobkrit@iapp.co.th
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-investment-consultant-licensing-exams.ExamQAml-interview-examples-mm-imdbCrosscoder-Qwen2.5-1.5B-vs-DeepScaleR-1.5B_max_activating_examplesSee Files and versions for pickled dictionaries and database versions of of max activating examples organized per available layer, as well as dataframes of available features.
Taylor-Swift-ExampleNote: This is a copy of https://www.kaggle.com/datasets/thespacefreak/taylor-swift-song-lyrics-all-albums that I'm hosting over here for convenience for a workshop
spacr-example-import
spaCR — Import test data
The same four microscope fields written in every container format and filename convention the Import module of spaCR reads, each with its cell, nucleus and pathogen masks and the measurements of its cells. It is the data behind Load test data… on the Import screen: pick a variant, and spaCR fills the screen with it and previews the import, so you can see every file land on the well, field and channel it came from.
About 283 MB in one uncompressed archive… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/spacr-example-import.example_annotated_code_repo_dataA description of the fields:
Column
What it captures
Typical values
id
Row identifier
1-100
repo_name
Example repository label
repo_14
file_path
Path + filename with extension
src/utils/parsefile.py
language
Programming language
Python, Java…
function_name
Target symbol that was reviewed
validateSession
annotation_summary
Free-text note written by the annotator
“Added input validation…”
potential_bug
Did the annotator flag a likely bug? (Yes/No)… See the full description on the dataset page: https://huggingface.co/datasets/hackerrank/example_annotated_code_repo_data.max-activating-examples-gemma-2-2b-l13-ckissanehcm-examples-aug-2024Dataset of some examples with hallucinations before and after passing through Vectara's Hallucination Correction Model. See our blogpost for details.
FAQ_embeddings_exampleecg-examples
ECG Heartbeat Examples
This dataset contains example ECG heartbeats. Source:
Single heartbeats were taken from MIT-BIH dataset preprocess into heartbeat Python
The dataset above was derived from the Original MIT-BIH Arrhythmia Dataset on PhysioNet. That dataset contains half-hour annotated ECGs from 48 patients.
These examples are part of the Heart Arrhythmia Detection Tools (hadt) Project (hadt GitHub Repository) and are intended for educational and research purposes.
A demo… See the full description on the dataset page: https://huggingface.co/datasets/fabriciojm/ecg-examples.example-dataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/harpomaxx/example-dataset.student_exam_performance
Student Exam Performance Prediction Dataset
This dataset contains numerical data about 20 students including their study habits, sleep patterns, and social media usage, along with their attendance rate and final exam score.
Dataset Fields
StudyHours: Number of hours the student studies daily
SleepHours: Number of hours the student sleeps daily
SocialMediaHours: Daily social media usage in hours
AttendanceRate: Percentage of class attendance
FinalExamScore: Score of the… See the full description on the dataset page: https://huggingface.co/datasets/Zeyustun/student_exam_performance.legal-cross-examination-impeachment-map-coherence-v0.1What this dataset does
You receive
witness key claim
prior statement
document contradiction
cross questions
impeachment point
materiality
You decide
coherent
or
incoherent
Daily use
cross plan QC
weak linkage detection
materiality focus check
wrong document flag
OpenAI_PodcastSentiment_XTwitterScraper_Example
🔍 X-Twitter Scraper: Real-Time Tweet Search & Scrape Tool
Search and scrape X-Twitter for posts by keyword, account, or trending topics.A simple, no-code tool to pull real-time, relevant content in LLM-ready JSON format — perfect for agents, RAG systems, or content workflows.
👉 Start Searching & Scraping on Hugging Face
✨ Features
⚡ Real-Time FetchStream the latest tweets as they’re posted — no delay.
🎯 Flexible SearchSearch by keywords, #hashtags, $cashtags… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/OpenAI_PodcastSentiment_XTwitterScraper_Example.Trump_Iran_XTwitterScraper_Example
🐦 X-Twitter Scraper: Real-Time Search and Data Extraction Tool
Search and scrape X-Twitter (formerly Twitter) for posts by keyword, account, or trending topics. This no-code tool makes it easy to generate real-time, LLM-ready datasets for any AI or content use case.
Get started with real-time scraping and structure tweet data instantly into clean JSON.
🚀 Key Features
⚡ Real-Time Fetch – Stream the latest tweets the moment they’re posted
🎯 Flexible Search –… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/Trump_Iran_XTwitterScraper_Example.on_the_books_example
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/on_the_books_example.korean-current-law-bar-exam-sft-1000
Korean Current-Law Bar Exam SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 스타일 SFT 데이터 1,000문항입니다.
이 데이터셋은 법무부 기출문제를 복제하지 않습니다. 기존 gyung/korean-bar-exam-moj-multiple-choice의 data/questions.csv는 난도와 과목 분포 참고 및 제15회 중복 방지 기준으로만 사용했습니다.
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 과목 분포, 제15회 유사도 QA 결과입니다.
Columns
question_text: 문제와 5개 선택지
answer: 정답 번호, 1부터 5… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-current-law-bar-exam-sft-1000.trove-examples-dataThis repo holds files that are used in Trove examples.
tevatron_msmarco_passage_aug_qrel.jsonl contains the qrels extracted from train.jsonl.gz file in Tevatron/msmarco-passage-aug repo.
examination-of-sars-cov-2-serological-test-results
Examination of SARS-CoV-2 serological test results from multiple commercial and laboratory platforms with an in-house serum panel
Description
Severe acute respiratory syndrome (SARS) coronavirus 2 (SARS-CoV-2) is a novel human coronavirus that was identified in 2019. SARS-CoV-2 infection results in an acute, severe respiratory disease called coronavirus disease 2019 (COVID-19). The emergence and rapid spread of SARS-CoV-2 has led to a global public health crisis, which… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/examination-of-sars-cov-2-serological-test-results.ml-interview-examples-adultIITM_Intro_to_Deep_Learning_Nppe1_exam_dataset
🧠 IITM Intro to Deep Learning & GenAI NPPE1 — Age & Gender Prediction Dataset
This dataset was prepared for the IIT Madras "Intro to Deep Learning & GenAI NPPE1" competition hosted on Kaggle.It contains face images and metadata used for multi-task learning — predicting both age (regression) and gender (classification) from image inputs.
📦 Dataset Structure
Files Included
File
Description
train/
Folder containing training face images.… See the full description on the dataset page: https://huggingface.co/datasets/AyusmanSamasi/IITM_Intro_to_Deep_Learning_Nppe1_exam_dataset.AMIE_PSEAE_Wrenbeck_2017_examplelegal-cross-examination-issue-impeachment-coherence-risk-v0.1What this dataset does
You receive
issues
witness claims
impeachment material
cross plan
objective
You decide
coherent
or
incoherent
Daily use
trial prep
witness attack planning
impeachment detection
example_mlexam_pa_titanicexample_ColameoSee 'Files' tab to download the specific example_Colameo_MS.csv and example_Colameo_RNA.csv
example_quotes
Dataset: Example Quotes
Starting example structure for how to store quotes.
Structure
Quote:
Type: String
Description: The primary content.
Antagonist:
Type: String
Description: The individual responsible for said quote.
Antagonists_id:
Type: Integer
Description: A unique identifier associated with the quoting antagonist. Used for generating URLs or internal referencing.
URL:
Type: List of Strings (separated by |)
Description: A list of URLs offering context or… See the full description on the dataset page: https://huggingface.co/datasets/Mediocreatmybest/example_quotes.
