CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01scikit-learn /iris Iris Species Dataset The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple Measurements in Taxonomic Problems, and can also be found on the UCI Machine Learning Repository. It includes three iris species with 50 samples each as well as some properties about each flower. One flower species is linearly separable from the other two, but the other two are not linearly separable from each other. The dataset is taken from UCI Machine Learning Repository's… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/iris.tabularn<1K13 likes16k downloads4y agoHugging Face02scikit-learn /churn-predictionCustomer churn prediction dataset of a fictional telecommunication company made by IBM Sample Datasets. Context Predict behavior to retain customers. You can analyze all relevant customer data and develop focused customer retention programs. Content Each row represents a customer, each column contains customer’s attributes described on the column metadata. The data set includes information about: Customers who left within the last month: the column is called Churn Services that each customer… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/churn-prediction.tabular1K<n<10K20 likes5.8k downloads4y agoHugging Face03scikit-learn /adult-census-income Adult Census Income Dataset The following was retrieved from UCI machine learning repository. This data was extracted from the 1994 Census bureau database by Ronny Kohavi and Barry Becker (Data Mining and Visualization, Silicon Graphics). A set of reasonably clean records was extracted using the following conditions: ((AAGE>16) && (AGI>100) && (AFNLWGT>1) && (HRSWK>0)). The prediction task is to determine whether a person makes over $50K a year. Description of fnlwgt (final weight)… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/adult-census-income.tabular10K<n<100K9 likes4.6k downloads4y agoHugging Face04OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K7 likes3.1k downloads29d agoHugging Face05SAIS-Life-Science /Aneumo Aneumo Datasets AneumoDataset is a comprehensive multi-modal dataset containing 3D anatomical structures and simulated hemodynamic parameters for cerebral aneurysms, supporting both computational modeling and AI-based analysis. textn<1K6 likes2.8k downloads6mo agoHugging Face06SciCodePile /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Domain-Code.tabulartext-generation1M<n<10M4 likes2k downloads6mo agoHugging Face07scikit-fingerprints /MoleculeNet_PCBA MoleculeNet PCBA PCBA (PubChem BioAssay) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict biological activity against 128 bioassays, generated by high-throughput screening (HTS). All tasks are binary active/non-active. Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros. Characteristic… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_PCBA.tabulartabular-classification100K<n<1M0 likes1.4k downloads2y agoHugging Face08scikit-fingerprints /MoleculeNet_BACE MoleculeNet BACE BACE dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict binding results for a set of inhibitors of humanβ-secretase 1 (BACE-1). Characteristic Description Tasks 1 Task type classification Total samples 1513 Recommended split scaffold Recommended metricAUROC References [1] Govindan Subramanian et al. "Computational Modeling of β-Secretase 1 (BACE-1)… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_BACE.texttabular-classification1K<n<10K0 likes1.4k downloads2y agoHugging Face09BeIR /scidocs-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/scidocs-qrels.texttext-retrieval10K<n<100K0 likes1.2k downloads4y agoHugging Face10scikit-fingerprints /MoleculeNet_Tox21 MoleculeNet Tox21 Tox21 dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict 12 toxicity targets, including nuclear receptors and stress response pathways. All tasks are binary. Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros. Characteristic Description Tasks 12 Task type multitask… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_Tox21.tabulartabular-classification1K<n<10K1 likes1.1k downloads2y agoHugging Face11scikit-fingerprints /LRGB_Peptides-func LRGB Peptides-func Peptides-func (Peptides functional) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through scikit-fingerprints library. The task is to predict functional properties of peptides. Characteristic Description Tasks 10 Task type classification Total samples 15535 Recommended split stratified random Recommended metric AUPRC References [1] Dwivedi, Vijay Prakash, et al. "Long Range Graph Benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-func.tabulartabular-classification10K<n<100K0 likes1.1k downloads6mo agoHugging Face12scikit-fingerprints /MoleculeNet_ESOL MoleculeNet ESOL ESOL (Estimated SOLubility) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict aqueous solubility. Targets are log-transformed, and the unit is log mols per litre (log Mol/L). Characteristic Description Tasks 1 Task type regression Total samples 1128 Recommended split scaffold Recommended metric RMSE References [1] John S. Delaney "ESOL: Estimating… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ESOL.texttabular-regression1K<n<10K3 likes1.1k downloads2y agoHugging Face13scikit-fingerprints /MoleculeNet_BBBP MoleculeNet BBBP BBBP (Blood-Brain Barrier Penetration) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict blood-brain barrier penetration (barrier permeability) of small drug-like molecules. Characteristic Description Tasks 1 Task type classification Total samples 2039 Recommended split scaffold Recommended metric AUROC References [1] Ines Filipa Martins et al. "A… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_BBBP.texttabular-classification1K<n<10K0 likes1.1k downloads2y agoHugging Face14scikit-fingerprints /MoleculeNet_ClinTox MoleculeNet ClinTox Load and return the ClinTox dataset, part of MoleculeNet [1] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict drug approval viability, by predicting clinical trial toxicity and final FDA approval status. Both tasks are binary. Characteristic Description Tasks 2 Task type multitask classification Total samples 1477 Recommended split scaffold Recommended metric AUROC References [1]… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox.tabulartabular-classification1K<n<10K0 likes988 downloads2y agoHugging Face15scikit-fingerprints /MoleculeNet_SIDER MoleculeNet SIDER Load and return the SIDER (Side Effect Resource) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict adverse drug reactions (ADRs) as drug side effects to 27 system organ classes in MedDRA classification. All tasks are binary. Characteristic Description Tasks 12 Task type multitask classification Total samples 7831 Recommended split scaffold Recommended metric… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_SIDER.tabulartabular-classification1K<n<10K0 likes972 downloads2y agoHugging Face16scikit-fingerprints /MoleculeNet_MUV MoleculeNet MUV Load and return the MUV (Maximum Unbiased Validation) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict 17 targets designed for validation of virtual screening techniques, based on PubChem BioAssays. All tasks are binary. Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros.… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_MUV.tabulartabular-classification10K<n<100K0 likes955 downloads2y agoHugging Face17scikit-fingerprints /MoleculeNet_Lipophilicity MoleculeNet Lipophilicity Lipophilicity dataset, part of MoleculeNet [1] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict octanol/water distribution coefficient (logD) at pH 7.4. Targets are already log transformed, and are a unitless ratio. Characteristic Description Tasks 1 Task type regression Total samples 4200 Recommended split scaffold Recommended metric RMSE References [1] Wu, Zhenqin, et al.… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_Lipophilicity.texttabular-regression1K<n<10K0 likes945 downloads2y agoHugging Face18scikit-fingerprints /MoleculeNet_HIV MoleculeNet HIV HIV dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict ability of molecules to inhibit HIV replication. Characteristic Description Tasks 1 Task type classification Total samples 41127 Recommended split scaffold Recommended metric AUROCWarning: in newer RDKit vesions, 7 molecules from the original dataset are not read correctly due to disallowed hypervalent states of… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_HIV.texttabular-classification10K<n<100K0 likes945 downloads2y agoHugging Face19scikit-fingerprints /MoleculeNet_ToxCast MoleculeNet ToxCast ToxCast dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict 617 toxicity targets from a large library of compounds based on in vitro high-throughput screening. All tasks are binary. Note that targets have missing values. Algorithms should be evaluated only on present labels. For training data, you may want to impute them, e.g. with zeros. Characteristic Description Tasks… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ToxCast.tabulartabular-classification1K<n<10K0 likes935 downloads2y agoHugging Face20scikit-fingerprints /MoleculeNet_FreeSolv MoleculeNet FreeSolv FreeSolv (Free Solvation Database) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through scikit-fingerprints library. The task is to predict hydration free energy of small molecules in water. Targets are in kcal/mol. Characteristic Description Tasks 1 Task type regression Total samples 642 Recommended split scaffold Recommended metric RMSE References [1] Mobley, D.L., Guthrie, J.P. "FreeSolv: a… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_FreeSolv.texttabular-regressionn<1K0 likes905 downloads2y agoHugging Face21scikit-fingerprints /LRGB_Peptides-struct LRGB Peptides-struct Peptides-struct (Peptides structural) dataset, part of Long Range Graph Benchmark (LRGB) [1]. It is intended to be used through scikit-fingerprints library. The task is to predict structural properties of peptides. Note that this is raw data, whereas the original paper [1] specifies that targets should be standardized (mean 0, standard deviation 1) before training and evaluation. scikit-fingerprints does this by default in the loader function, otherwise this… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/LRGB_Peptides-struct.tabulartabular-classification10K<n<100K0 likes862 downloads6mo agoHugging Face22scikit-learn /breast-cancer-wisconsin Breast Cancer Wisconsin Diagnostic Dataset Following description was retrieved from breast cancer dataset on UCI machine learning repository. Features are computed from a digitized image of a fine needle aspirate (FNA) of a breast mass. They describe characteristics of the cell nuclei present in the image. A few of the images can be found at here. Separating plane described above was obtained using Multisurface Method-Tree (MSM-T), a classification method which uses linear… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/breast-cancer-wisconsin.tabularn<1K6 likes844 downloads4y agoHugging Face23scikit-learn /auto-mpg Auto Miles per Gallon (MPG) Dataset Following description was taken from UCI machine learning repository. Source: This dataset was taken from the StatLib library which is maintained at Carnegie Mellon University. The dataset was used in the 1983 American Statistical Association Exposition. Data Set Information: This dataset is a slightly modified version of the dataset provided in the StatLib library. In line with the use by Ross Quinlan (1993) in predicting the attribute… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/auto-mpg.tabulartabular-classificationn<1K3 likes733 downloads3y agoHugging Face24armanc /ScienceQAThis is the ScientificQA dataset by Saikh et al (2022). @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K14 likes723 downloads4y agoHugging Face25stevez80 /Sci-Fi-Books-gutenberg Gutenberg Sci-Fi Book Dataset This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing. Data Format The dataset is provided in CSV format. Each record represents a book and includes the following fields: ID: A unique identifier for the book. Title: The title of the book. Author: The author(s) of the book. Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.texttext-generation1K<n<10K12 likes588 downloads3y agoHugging Face26mariiakoroliuk /generalization-science-datadocumentn<1K0 likes410 downloads3d agoHugging Face27liuhangbiao /SciCode-Domain-Code DATA1: Domain-Specific Code Dataset Dataset Overview DATA1 is a large-scale domain-specific code dataset focusing on code samples from interdisciplinary fields such as biology, chemistry, materials science, and related areas. The dataset is collected and organized from GitHub repositories, covering 178 different domain topics with over 1.1 billion lines of code. Dataset Statistics Total Datasets: 178 CSV files Total Data Size: ~115 GB Total Lines of Code: Over… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/SciCode-Domain-Code.tabulartext-generation1M<n<10M0 likes382 downloads6mo agoHugging Face28scikit-fingerprints /ExpansionRx_OpenADMET_RLM_CLint ExpansionRx-OpenADMET RLM CLint RLM CLint (rat liver microsomal intrinsic clearance) dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through scikit-fingerprints library. The task is to predict the rat liver microsomal intrinsic clearance (RLM CLint) of molecules. Note that this dataset was not part of the original challenge. It was provided by the organizers afterward as an additional endpoint. Characteristic Description Tasks 1… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_RLM_CLint.texttabular-regressionn<1K0 likes368 downloads6mo agoHugging Face29scikit-learn /imdbThis is the sentiment analysis dataset based on IMDB reviews initially released by Stanford University. This is a dataset for binary sentiment classification containing substantially more data than previous benchmark datasets. We provide a set of 25,000 highly polar movie reviews for training, and 25,000 for testing. There is additional unlabeled data for use as well. Raw text and already processed bag of words formats are provided. See the README file contained in the release for more… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/imdb.text10K<n<100K0 likes353 downloads4y agoHugging Face30deep-principle /science_materialstabularn<1K0 likes340 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.