CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FredZhang7 /toxi-text-3MThis is a large multilingual toxicity dataset with 3M rows of text data from 55 natural languages, all of which are written/sent by humans, not machine translation models. The preprocessed training data alone consists of 2,880,667 rows of comments, tweets, and messages. Among these rows, 416,529 are classified as toxic, while the remaining 2,463,773 are considered neutral. Below is a table to illustrate the data composition: Toxic Neutral Total multilingual-train-deduplicated.csv… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/toxi-text-3M.texttext-classification1M<n<10M32 likes579 downloads1y agoHugging Face02freddyaboulton /gradio-reviewstabularn<1K2 likes373 downloads2d agoHugging Face03FredZhang7 /all-scam-spamThis is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham. 1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT. Some preprcoessing algorithms spam_assassin.js, followed by spam_assassin.py enron_spam.py Data composition Description To make the text… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.texttext-classification10K<n<100K15 likes239 downloads2y agoHugging Face04KrossKinetic /FRED_DatasetTextual Time Series Dataset collected from the FRED.gov dataset using the FRED API for finetuning / pretraining in csv format as part of Humanity Unleashed Research. text100K<n<1M1 likes112 downloads2y agoHugging Face05fredxlpy /LuxInstruct LuxInstruct Dataset Summary LuxInstruct is the first large-scale cross-lingual instruction tuning dataset for Luxembourgish, introduced in LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish (Philippy et al., 2025). It addresses the lack of high-quality instruction–response data for low-resource languages by avoiding direct machine translation into Luxembourgish. Instead, it leverages aligned data from English, French, and German to generate natural… See the full description on the dataset page: https://huggingface.co/datasets/fredxlpy/LuxInstruct.texttext-generation100K<n<1M2 likes59 downloads1y agoHugging Face06Fred666 /ocnli3ktabular1K<n<10K0 likes37 downloads3y agoHugging Face07Fred666 /ocnliThis dataset is copied from CLUE with certain modifications. The paper of CLUE is OCNLI. The modifications are: Transform json file to csv file. Encoding in UTF-8. Remove data entries whose label value is '-'. Replace label values, 'neutral' to 1, 'entailment' to 0, and 'contradiction' to 2. Add one column 'sentence1', whose value is '前提:' + premise value + '结论:' + hypothsis value. ocnli_train_std.csv comes from train.50k.json. ocnli_test_std.csv comes from dev.json. texttext-classification10K<n<100K1 likes29 downloads3y agoHugging Face08fredpeng /Text2SQL_Workflow_Trace Text2SQL Workflow Trace Dataset Description This dataset contains workflow traces for Text-to-SQL tasks, capturing the intermediate steps of translating natural language queries to executable SQL. It was used as input trace for the research presented in the paper:"HEXGEN-TEXT2SQL: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL Workflow" (arXiv:2505.05286). The end-to-end Text-to-SQL queries collected in the dataset are from BIRD bench, and the trace… See the full description on the dataset page: https://huggingface.co/datasets/fredpeng/Text2SQL_Workflow_Trace.tabular1K<n<10K1 likes25 downloads1y agoHugging Face09Frederick001 /Material_Selection_EvalA benchmark designed to facilitate evaluation and modify the behavior of a foundation model through different existing techniques in the context of material selection for conceptual design. The data is collected by conducting a survey of experts in the field of material selection. The same questions mentioned in keyquestions.csv are asked to experts. This can be used to evaluate a Language model performance and its spread compared to a human evaluation. To get into a more detailed explanation… See the full description on the dataset page: https://huggingface.co/datasets/Frederick001/Material_Selection_Eval.tabulartext-generationn<1K0 likes18 downloads9mo agoHugging Face10freddyaboulton /chatinterface_with_image_csv Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/freddyaboulton/chatinterface_with_image_csv.imagen<1K0 likes17 downloads3y agoHugging Face11freddie /agi-eval-sat-math-judgmentstextn<1K0 likes14 downloads1y agoHugging Face12freddyaboulton /dope_data_points_14 Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/freddyaboulton/dope_data_points_14.tabularn<1K0 likes13 downloads3y agoHugging Face13freddyaboulton /alpaca-csvtext10K<n<100K0 likes11 downloads3y agoHugging Face14Freddiex /chess Dataset Modified version of lichess_elite_2020-06 Created by: https://lichess.org/@/nikonoel Source: https://database.nikonoel.fr/ Source: https://database.nikonoel.fr/lichess_elite_2020-06.zip text100K<n<1M0 likes10 downloads2y agoHugging Face15Kabatubare /fredericktextn<1K0 likes8 downloads3y agoHugging Face16freddyaboulton /affordable-housingtabular1K<n<10K0 likes8 downloads1y agoHugging Face17freddyaboulton /dope_data_points Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/freddyaboulton/dope_data_points.tabularn<1K0 likes7 downloads3y agoHugging Face18Freddiex /chess-400k Dataset Modified version of lichess_elite_2020-06 Created by: https://lichess.org/@/nikonoel Source: https://database.nikonoel.fr/ Source: https://database.nikonoel.fr/lichess_elite_2020-06.zip text100K<n<1M0 likes7 downloads2y agoHugging Face19beamstation /all-restaurants-in-frederick-maryland-us-641588 All Restaurants in Frederick, Maryland, US Free sample dataset from BeamStation This dataset contains a complete export of all restaurants in Frederick, Maryland, United States. It includes 523 records, each representing a distinct dining establishment, and is refreshed on a weekly basis to keep the information current. The export provides every available column from the source, offering a full profile for each venue—such as name, address, cuisine type, contact details, hours of… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/all-restaurants-in-frederick-maryland-us-641588.tabulartabular-classificationn<1K0 likes3 downloads6mo agoHugging Face20freddyaboulton /dope_data_points_2 Dataset Card for Dataset Name Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/freddyaboulton/dope_data_points_2.tabularn<1K0 likes1 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.