CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01scikit-learn /iris Iris Species Dataset The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple Measurements in Taxonomic Problems, and can also be found on the UCI Machine Learning Repository. It includes three iris species with 50 samples each as well as some properties about each flower. One flower species is linearly separable from the other two, but the other two are not linearly separable from each other. The dataset is taken from UCI Machine Learning Repository's… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/iris.tabularn<1K13 likes16k downloads4y agoHugging Face02scikit-learn /churn-predictionCustomer churn prediction dataset of a fictional telecommunication company made by IBM Sample Datasets. Context Predict behavior to retain customers. You can analyze all relevant customer data and develop focused customer retention programs. Content Each row represents a customer, each column contains customer’s attributes described on the column metadata. The data set includes information about: Customers who left within the last month: the column is called Churn Services that each customer… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/churn-prediction.tabular1K<n<10K20 likes5.8k downloads4y agoHugging Face03scikit-learn /adult-census-income Adult Census Income Dataset The following was retrieved from UCI machine learning repository. This data was extracted from the 1994 Census bureau database by Ronny Kohavi and Barry Becker (Data Mining and Visualization, Silicon Graphics). A set of reasonably clean records was extracted using the following conditions: ((AAGE>16) && (AGI>100) && (AFNLWGT>1) && (HRSWK>0)). The prediction task is to determine whether a person makes over $50K a year. Description of fnlwgt (final weight)… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/adult-census-income.tabular10K<n<100K9 likes4.6k downloads4y agoHugging Face04scikit-learn /breast-cancer-wisconsin Breast Cancer Wisconsin Diagnostic Dataset Following description was retrieved from breast cancer dataset on UCI machine learning repository. Features are computed from a digitized image of a fine needle aspirate (FNA) of a breast mass. They describe characteristics of the cell nuclei present in the image. A few of the images can be found at here. Separating plane described above was obtained using Multisurface Method-Tree (MSM-T), a classification method which uses linear… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/breast-cancer-wisconsin.tabularn<1K6 likes844 downloads4y agoHugging Face05scikit-learn /auto-mpg Auto Miles per Gallon (MPG) Dataset Following description was taken from UCI machine learning repository. Source: This dataset was taken from the StatLib library which is maintained at Carnegie Mellon University. The dataset was used in the 1983 American Statistical Association Exposition. Data Set Information: This dataset is a slightly modified version of the dataset provided in the StatLib library. In line with the use by Ross Quinlan (1993) in predicting the attribute… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/auto-mpg.tabulartabular-classificationn<1K3 likes733 downloads3y agoHugging Face06scikit-learn /imdbThis is the sentiment analysis dataset based on IMDB reviews initially released by Stanford University. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well. Raw text and already processed bag of words formats are provided. See the README file contained in the release for more… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/imdb.text10K<n<100K0 likes353 downloads4y agoHugging Face07scikit-learn /Fish Dataset Summary Dataset recording various measurements of 7 different species of fish at a fish market. Predictive models can be used to predict weight, species, etc. Feature Descriptions Species - Species name of fish Weight - Weight of fish in grams Length1 - Vertical length in cm Length2 - Diagonal length in cm Length3 - Cross length in cm Height - Height in cm Width - Width in cm Acknowledgments Dataset created by Aung Pyae, and found on Kaggle tabularn<1K0 likes175 downloads4y agoHugging Face08scikit-learn /student-alcohol-consumption Student Alcohol Consumption Dataset A dataset on social, gender and study data from secondary school students. Following was retrieved from UCI machine learning repository. Context: The data were obtained in a survey of students math and portuguese language courses in secondary school. It contains a lot of interesting social, gender and study information about students. You can use it for some EDA or try to predict students final grade. Content: Attributes for both student-mat.csv… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/student-alcohol-consumption.tabular1K<n<10K3 likes131 downloads4y agoHugging Face09scikit-learn /tips A Waiter's Tips The following description was retrieved from Kaggle page. Food servers’ tips in restaurants may be influenced by many factors, including the nature of the restaurant, size of the party, and table locations in the restaurant. Restaurant managers need to know which factors matter when they assign tables to food servers. For the sake of staff morale, they usually want to avoid either the substance or the appearance of unfair treatment of the servers, for whom tips (at… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/tips.tabularn<1K0 likes105 downloads4y agoHugging Face10relai-ai /scikit-learn-reasoningSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: scikit-learn Data Source Link: https://scikit-learn.org/stable/index.html Data Source License: https://github.com/scikit-learn/scikit-learn/blob/main/COPYING Data Source Authors: scikit-learn contributors AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answering1K<n<10K0 likes26 downloads1y agoHugging Face11relai-ai /scikit-learn-standardSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: scikit-learn Data Source Link: https://scikit-learn.org/stable/index.html Data Source License: https://github.com/scikit-learn/scikit-learn/blob/main/COPYING Data Source Authors: scikit-learn contributors AI Benchmarks by Data Agents. 2025 RELAI.AI. Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answering1K<n<10K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.