CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.8k downloads3y agoHugging Face02RanveerChaudhary /password_strength_datasettexttext-classification100K<n<1M2 likes1.3k downloads1y agoHugging Face03randalakab /Crop-recommendationtabular1K<n<10K0 likes553 downloads6mo agoHugging Face04AmanPriyanshu /random-small-github-repositories random-small-github-repositories A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. Contents seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash) repos-zipped/ — one .zip per repo, named {repo_hash}.zip unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.tabulartext-generation1K<n<10K0 likes427 downloads6mo agoHugging Face05AmanPriyanshu /random-python-github-repositories random-python-github-repositories A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files. Contents repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash) repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.tabulartext-generation1K<n<10K0 likes356 downloads6mo agoHugging Face06jason1966 /alitaqi000_world-university-rankings-2023 World University Rankings 2023 World University Rankings 2023 include 1,799 universities across 104 countries. Dataset Info Source: Kaggle Original Size: 0.07 MB Kaggle Downloads: 8,164 Files: 1 Files World University Rankings 2023.csv Mirrored from Kaggle tabular1K<n<10K0 likes285 downloads6mo agoHugging Face07cycloevan /Ransomware_PE_Header_Feature_Dataset Dataset Card for Ransomware PE Header Feature Dataset Dataset Description Dataset Summary This dataset contains PE header features (first 1024 bytes) from 2,157 Windows executable samples, comprising 1,134 legitimate software (goodware) and 1,023 ransomware samples across 25 ransomware families. Each sample is represented by numerical features extracted from the raw PE header. Supported Tasks Binary Classification: Distinguish between goodware and… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/Ransomware_PE_Header_Feature_Dataset.documentimage-classification1K<n<10K0 likes186 downloads8mo agoHugging Face08Randa /MAOffens Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Randa/MAOffens.texttext-classification10K<n<100K1 likes165 downloads2y agoHugging Face09AmelieSchreiber /binding_sites_random_split_by_family_550KThis dataset is obtained from a UniProt search for protein sequences with family and binding site annotations. The dataset includes unreviewed (TrEMBL) protein sequences as well as reviewed sequences. We refined the dataset by only including sequences with an annotation score of 4. We sorted and split by family, where random families were selected for the test dataset until approximately 20% of the protein sequences were separated out for test data. We excluded any sequences with <, >, or ?… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/binding_sites_random_split_by_family_550K.text100K<n<1M4 likes136 downloads3y agoHugging Face10RaniduG /SiPaKosa-Sent SiPaKosa: Sinhala-Pali Buddhist Corpus A comprehensive corpus of canonical and classical Buddhist texts in Sinhala and Pali, compiled from historical archives and web-scraped canonical scriptures. This is the sentence-level version of the SiPaKosa dataset. Where SiPaKosa contains book level text, this dataset has all sentences by book. Related dataset (book-level): RaniduG/SiPaKosa Dataset Statistics Total Sentences: 786,344 Sinhala Sentences: 465,539 (59.2%) Mixed… See the full description on the dataset page: https://huggingface.co/datasets/RaniduG/SiPaKosa-Sent.texttext-generation100K<n<1M0 likes117 downloads6mo agoHugging Face11guactastesgood /GSM-Ranges GSM-Ranges Dataset 📄 Paper: Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges🔗 GitHub Repository: GSM-Ranges GitHub What is GSM-Ranges? GSM-Ranges is a dataset generator built upon the GSM8K benchmark. It systematically modifies numerical values in math word problems to assess the robustness of large language models (LLMs) across a broad spectrum of numerical scales. By introducing numerical… See the full description on the dataset page: https://huggingface.co/datasets/guactastesgood/GSM-Ranges.textquestion-answering10K<n<100K0 likes93 downloads2y agoHugging Face12brightkey /uni-rankings-2026 BrightKey Independent University Rankings Dataset (2026) 299 universities × 55 countries × 6 dimensions, evaluated independently. No payments from institutions accepted. Public data only. This is the open release of the BrightKey university rankings — an independent alternative to QS, THE, and Shanghai rankings. Released under CC BY 4.0. Live site: https://brightkey.co/en/rankings/methodology GitHub repo: https://github.com/arthurb2l/brightkey-university-dataset Zenodo DOI:… See the full description on the dataset page: https://huggingface.co/datasets/brightkey/uni-rankings-2026.tabulartabular-classificationn<1K0 likes79 downloads3mo agoHugging Face13CyberMax-tools /open-domain-ranks Linkheft Open Domain Ranks: free domain authority data for 10.3 million domains Try the paid tool: Linkheft on Apify: score any domain list via API, with 4-month trends. First try costs cents; pay only for results. Buy: Velbrake Pro – Personal (1 site) ($39.00/year): for site owners: control which AI crawlers can read your WordPress site (the base plugin is free). Checkout by Polar. An open alternative to proprietary "domain authority" scores. For each of the top 10… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/open-domain-ranks.tabulartabular-regression10M<n<100M0 likes74 downloads10h agoHugging Face14Randolphzeng /Mr-GSM8KView the project page: https://github.com/dvlab-research/DiagGSM8K see our paper at https://arxiv.org/abs/2312.17080 Description In this work, we introduce a novel evaluation paradigm for Large Language Models, one that challenges them to engage in meta-reasoning. Our paradigm shifts the focus from result-oriented assessments, which often overlook the reasoning process, to a more holistic evaluation that effectively differentiates the cognitive capabilities among models. For… See the full description on the dataset page: https://huggingface.co/datasets/Randolphzeng/Mr-GSM8K.tabularquestion-answering1K<n<10K12 likes61 downloads3y agoHugging Face15kasi-ranaweera /Sri_Lankan_UGC_Cutoff_Mark_Dataset Sri Lankan University Z-Score Recommendations This dataset supports the Z-Score University Finder project, providing recommendations for Sri Lankan Advanced Level (A/L) students based on their Z-Score, academic stream, and district. It is designed for educational research and building recommendation systems for university admissions. Provenance Sources The dataset is derived from aggregated historical university admission data from the University Grants… See the full description on the dataset page: https://huggingface.co/datasets/kasi-ranaweera/Sri_Lankan_UGC_Cutoff_Mark_Dataset.tabular10K<n<100K2 likes53 downloads1y agoHugging Face16kyisaiah47 /tooldrift-model-rankings ToolDrift: OpenRouter model usage rankings, captured daily One row per model per ranking window per capture: its rank, the tokens and requests behind that rank, and its share of the window. The series shows which models the market actually routes work to, day by day. Rows in this cut 44,369 One row is one model in one ranking window on one capture day Cut 2026-09-04 Refreshed Monthly, on the first of the month Measured by ToolDrift Method… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-model-rankings.tabular10K<n<100K0 likes50 downloads6d agoHugging Face17kyisaiah47 /tooldrift-app-rankings ToolDrift: OpenRouter app usage rankings, captured daily One row per app per ranking window per capture, with the tool it maps to where ToolDrift tracks one. It is the same series as the model rankings, read from the consumer side. Rows in this cut 641 One row is one app in one ranking window on one capture day Cut 2026-09-04 Refreshed Monthly, on the first of the month Measured by ToolDrift Method https://toolproof.thecompound.tech/methodology Licence… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-app-rankings.tabularn<1K0 likes49 downloads6d agoHugging Face18raniahaya /tourism-package-predictiontabular1K<n<10K1 likes41 downloads5d agoHugging Face19benni-ben /random-sentence-v2 Random Sentences Dataset (version 2) This is a random sentence dataset, it has random simplistic children sentences paired with completely unrelated random words. This dataset was used in SmolBabble2-360m, an AI model that spits out random sentences regardless of what is said to it. It is an improved version of the previous dataset, and it has much more examples. You can find more information on both of these models on the provided repository links. The dataset has prompt and… See the full description on the dataset page: https://huggingface.co/datasets/benni-ben/random-sentence-v2.text1K<n<10K0 likes39 downloads3d agoHugging Face20Raniahossam33 /Islamweb_part2text10K<n<100K0 likes33 downloads2y agoHugging Face21randallkaren9 /past-setting-ede8fc past-setting-ede8fc Synthetic sensors test data: 51 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/randallkaren9/past-setting-ede8fc.tabularn<1K0 likes33 downloads15d agoHugging Face22CooperBench /qwen35-9b-question-first-coop-random-50 What this is Cooperative two-agent coding dataset: 49 task pairs across 15 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a question-first prompt variant — agents begin by asking each other clarifying questions about their respective features before starting implementation, aiming to surface integration concerns early. All 49 pairs were successfully evaluated. At a glance Field Value Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-question-first-coop-random-50.tabularn<1K0 likes31 downloads3mo agoHugging Face23gunahkarcasper /poker-gto-strategy-and-hand-rangestextn<1K0 likes30 downloads7mo agoHugging Face24addykan /genome-classification-with-randomtext10K<n<100K0 likes27 downloads2y agoHugging Face25genbio-ai /100M-random-promotersBoer, Carl G. de, Eeshit Dhaval Vaishnav, Ronen Sadeh, Esteban Luis Abeyta, Nir Friedman, and Aviv Regev. 2020. “Deciphering Eukaryotic Gene-Regulatory Logic with 100 Million Random Promoters.” Nature Biotechnology 38 (1): 56–65. https://doi.org/10.1038/s41587-019-0315-8. https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE104878 text10M<n<100M1 likes26 downloads2y agoHugging Face26ignoreandfly /VSR_random_tsvtext1K<n<10K0 likes26 downloads1y agoHugging Face27SohamThakkar-07 /india-ev-range-dataset 🔋 India EV Real-World Range Dataset A comprehensive dataset of 7,291 data points covering 61 Indian 4-wheeler EV variants from 18 manufacturers across 12 Indian driving scenarios. Dataset Description This dataset was built to predict the real-world driving range of Electric Vehicles under Indian conditions. It combines: Real EV specifications from all major EVs sold in India (Tata, Mahindra, Hyundai, MG, Kia, BYD, BMW, Mercedes, etc.) Physics-based energy modeling using… See the full description on the dataset page: https://huggingface.co/datasets/SohamThakkar-07/india-ev-range-dataset.tabulartabular-regression1K<n<10K0 likes26 downloads5mo agoHugging Face28trabten /tibetan_ranking_training_data Tibetan Ranking Training Data — Tier 1 Training data for a cross-encoder that re-ranks Tibetan→English translation candidates by contextual relevance. Each row pairs a Tibetan term and sentence context with one English gloss candidate and a binary label: 1 (this gloss is defensible in this context) or 0 (it is not). Used to train trabten/tibetan_ranking_BUDA. Files File Rows Description training_v2_merged.csv 71,050 Tier 1 — surgically annotated… See the full description on the dataset page: https://huggingface.co/datasets/trabten/tibetan_ranking_training_data.texttext-classification100K<n<1M0 likes26 downloads3mo agoHugging Face29CooperBench /qwen35-9b-contract-first-coop-random-50 What this is Cooperative two-agent coding dataset: 36 task pairs across 13 repos (random-50 subset), generated with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a contract-first prompt variant — agents first agree on a shared interface contract (function signatures, data structures, API boundaries) before independently implementing their respective features. All 36 pairs were successfully evaluated. At a glance Field Value Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-contract-first-coop-random-50.tabularn<1K0 likes24 downloads3mo agoHugging Face30RanjuBiswas /ner-email Overview: This dataset is augmented through llama-3.1:8b. The pourpose is to finetune llm for token classification i.e Email in our case. Following tags are present in dataset: full_name : 1 email : 2 gender : 3 city : 4 country : 5 text1K<n<10K0 likes23 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.