datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toxi-text-3MThis is a large multilingual toxicity dataset with 3M rows of text data from 55 natural languages, all of which are written/sent by humans, not machine translation models.
The preprocessed training data alone consists of 2,880,667 rows of comments, tweets, and messages. Among these rows, 416,529 are classified as toxic, while the remaining 2,463,773 are considered neutral. Below is a table to illustrate the data composition:
Toxic
Neutral
Total
multilingual-train-deduplicated.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/toxi-text-3M.gradio-reviewsall-scam-spamThis is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham.
1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT.
Some preprcoessing algorithms
spam_assassin.js, followed by spam_assassin.py
enron_spam.py
Data composition
Description
To make the text… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.FRED_DatasetTextual Time Series Dataset collected from the FRED.gov dataset using the FRED API for finetuning / pretraining in csv format as part of Humanity Unleashed Research.
LuxInstruct
LuxInstruct
Dataset Summary
LuxInstruct is the first large-scale cross-lingual instruction tuning dataset for Luxembourgish, introduced in LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish (Philippy et al., 2025).
It addresses the lack of high-quality instruction–response data for low-resource languages by avoiding direct machine translation into Luxembourgish. Instead, it leverages aligned data from English, French, and German to generate natural… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/LuxInstruct.ocnli3kocnliThis dataset is copied from CLUE with certain modifications.
The paper of CLUE is OCNLI.
The modifications are:
Transform json file to csv file.
Encoding in UTF-8.
Remove data entries whose label value is '-'.
Replace label values, 'neutral' to 1, 'entailment' to 0, and 'contradiction' to 2.
Add one column 'sentence1', whose value is '前提:' + premise value + '结论:' + hypothsis value.
ocnli_train_std.csv comes from train.50k.json.
ocnli_test_std.csv comes from dev.json.
Text2SQL_Workflow_Trace
Text2SQL Workflow Trace
Dataset Description
This dataset contains workflow traces for Text-to-SQL tasks, capturing the intermediate steps of translating natural language queries to executable SQL. It was used as input trace for the research presented in the paper:"HEXGEN-TEXT2SQL: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL Workflow" (arXiv:2505.05286).
The end-to-end Text-to-SQL queries collected in the dataset are from BIRD bench, and the trace… See the full description on the dataset page: https://huggingface.co/datasets/fredpeng/Text2SQL_Workflow_Trace.Material_Selection_EvalA benchmark designed to facilitate evaluation and modify the behavior of a foundation model through different existing techniques in the context of material selection for conceptual design.
The data is collected by conducting a survey of experts in the field of material selection. The same questions mentioned in keyquestions.csv are asked to experts.
This can be used to evaluate a Language model performance and its spread compared to a human evaluation.
To get into a more detailed explanation… See the full description on the dataset page: https://huggingface.co/datasets/Frederick001/Material_Selection_Eval.chatinterface_with_image_csv
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/freddyaboulton/chatinterface_with_image_csv.agi-eval-sat-math-judgmentsdope_data_points_14
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/freddyaboulton/dope_data_points_14.alpaca-csvchess
Dataset
Modified version of lichess_elite_2020-06
Created by: https://lichess.org/@/nikonoel
Source: https://database.nikonoel.fr/
Source: https://database.nikonoel.fr/lichess_elite_2020-06.zip
frederickaffordable-housingdope_data_points
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/freddyaboulton/dope_data_points.chess-400k
Dataset
Modified version of lichess_elite_2020-06
Created by: https://lichess.org/@/nikonoel
Source: https://database.nikonoel.fr/
Source: https://database.nikonoel.fr/lichess_elite_2020-06.zip
all-restaurants-in-frederick-maryland-us-641588
All Restaurants in Frederick, Maryland, US
Free sample dataset from BeamStation
This dataset contains a complete export of all restaurants in Frederick, Maryland, United States. It includes 523 records, each representing a distinct dining establishment, and is refreshed on a weekly basis to keep the information current. The export provides every available column from the source, offering a full profile for each venue—such as name, address, cuisine type, contact details, hours of… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/all-restaurants-in-frederick-maryland-us-641588.dope_data_points_2
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/freddyaboulton/dope_data_points_2.
