datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STARK_10k
STARK: Spatial-Temporal reAsoning benchmaRK
STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure.
Dataset Summary
Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity:
State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.taxbench-au
TaxBench-AU
A benchmark for testing whether AI agents can calculate Australian tax.
TaxBench-AU contains 156 Australian tax calculation questions, presented as multiple-choice (4-option) worked tax problems. The benchmark is designed to test whether an AI agent can read the facts, apply the right Australian tax rule for the relevant income year, do the calculation, and choose the correct answer.
The Kaggle mirror is published as Agent Tax Exam for Australian Tax.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Pn101/taxbench-au.Python-Code-Solutions
Python Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
Python Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.10k_rows_cleaned_prompts
10K Rows Cleaned Prompts Dataset
Created by Aipresso LIMITED, London, UK
⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use
You must provide attribution when using this data in publications, research, or commercial products.
Dataset Overview
A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models.
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.plot-palette-100k
Empowering Writers with a Universe of Ideas Plot Palette DataSet HuggingFace » Plot Palette was created to fine-tune large language models for creative writing, generating diverse outputs through iterative loops and seed data. It is designed to be run on a Linux system with systemctl for managing services. Included is the service structure, specific category prompts and ~100k data entries. The dataset is available here or… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/plot-palette-100k.Ko-LAION-Aesthetics-10M
LAION-Aesthetics 10M Dataset Card
Dataset details
Dataset type:
Laion aesthetic is a subset of laion5B that has been estimated by a model trained on top of clip embeddings to be aesthetic. The intended usage of this dataset is image generation
Paper or resources for more information:
https://laion.ai/blog/laion-aesthetics/
Acknowledgements
This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grants funded by the… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/Ko-LAION-Aesthetics-10M.alexa-qa
Alexa Answers from alexaanswers.amazon.com
The Alexa Answers community helps to improve Alexa’s knowledge and answer questions asked by Alexa users. Which contains some very quirky and hard question like
Q: what percent of the population has blackhair
A: The most common hair color in the world is black and its found in wide array of background and ethnicities. About 75 to 85% of the global population has either black hair or the deepest brown shade.
Q: what was the world population… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/alexa-qa.NCERT_Science_10thCPP-Code-Solutions
C++ Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
C++ Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
GeoQuestions1089
A crowdsourced geospatial question-answering dataset that contains 1089 triples of natural language questions, SPARQL/GeoSPARQL queries, and their answers over YAGO2geo.
Overview
GeoQuestions1089 is a crowdsourced geospatial question-answering dataset that targets the Knowledge Graph YAGO2geo. It contains 1089 triples of geospatial questions, their answers, and the respective SPARQL/GeoSPARQL queries.
It has been used to benchmark two state of the art Question… See the full description on the dataset page: https://huggingface.co/datasets/AI-team-UoA/GeoQuestions1089.NCERT_Social_Studies_10thDataset_com_Laws_1039
Dataset Card act of legislation in Thailand.
The dataset contains problem act of legislation in Thailand.
This dataset is taken from (https://www.drthawip.com/), comes from the government website.
JS-Code-Solutions
Python Code Solutions
Features
1000k of JS Code Solutions for Text Generation and Question Answering
JS Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
Kratki-Istorii-Instruct-100kKratki-Istorii-Instruct-100k is a synthetically generated dataset (using INSAIT-Institute/BgGPT-Gemma-2-9B-IT-v1.0) of short stories (3-5) paragraphs, which a young kid should be able to understand. The simplicity of the language used makes it very suitable for training and studying the behaviour of really small Language Models (<500M parameters).
The dataset consists of ~100k texts in Bulgarian. You can use the dataset via the HF interface:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/WT-solutions/Kratki-Istorii-Instruct-100k.birthday_quotes_1_to_100
Birthday Quote 1 to 100 — Full Combination Dataset
The Birthday Quote 1 to 100 dataset is an extensive collection of 3,807 birthday messages generated through complete combinations of tone, theme, and valid recipient types across realistic age groups.This dataset spans ages from 1 to 100 years, providing a highly diverse and customizable resource for generating personalized birthday wishes for any recipient.
Each entry contains a birthday message along with structured metadata — age… See the full description on the dataset page: https://huggingface.co/datasets/tejasashinde/birthday_quotes_1_to_100.korean-current-law-bar-exam-sft-1000
Korean Current-Law Bar Exam SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 스타일 SFT 데이터 1,000문항입니다.
이 데이터셋은 법무부 기출문제를 복제하지 않습니다. 기존 gyung/korean-bar-exam-moj-multiple-choice의 data/questions.csv는 난도와 과목 분포 참고 및 제15회 중복 방지 기준으로만 사용했습니다.
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 과목 분포, 제15회 유사도 QA 결과입니다.
Columns
question_text: 문제와 5개 선택지
answer: 정답 번호, 1부터 5… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-current-law-bar-exam-sft-1000.llm_practiceSystem-Response-100K
System-Response-100K dataset
This dataset contains text and code for machine learning tasks including:
Text Generation
Text Classification
Summarization
Question Answering
The dataset includes text formatted in JSON and is in English.
Dataset Statistics
Number of entries: Not specified in the information you provided.
Modalities
Text
Code
Formats
JSON
Languages
English
Getting Started
This section can include… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/System-Response-100K.NCERT_Science_10th
