datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
argument_quality_ranking_30k
Dataset Card for Argument-Quality-Ranking-30k Dataset
Dataset Summary
Argument Quality Ranking
The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets.
The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis.
Argument Topic
This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.password_strength_datasetCrop-recommendationrandom-small-github-repositories
random-small-github-repositories
A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks.
Contents
seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash)
repos-zipped/ — one .zip per repo, named {repo_hash}.zip
unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.random-python-github-repositories
random-python-github-repositories
A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files.
Contents
repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash)
repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.alitaqi000_world-university-rankings-2023
World University Rankings 2023
World University Rankings 2023 include 1,799 universities across 104 countries.
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 8,164
Files: 1
Files
World University Rankings 2023.csv
Mirrored from Kaggle
Ransomware_PE_Header_Feature_Dataset
Dataset Card for Ransomware PE Header Feature Dataset
Dataset Description
Dataset Summary
This dataset contains PE header features (first 1024 bytes) from 2,157 Windows executable samples, comprising 1,134 legitimate software (goodware) and 1,023 ransomware samples across 25 ransomware families. Each sample is represented by numerical features extracted from the raw PE header.
Supported Tasks
Binary Classification: Distinguish between goodware and… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/Ransomware_PE_Header_Feature_Dataset.MAOffens
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Randa/MAOffens.binding_sites_random_split_by_family_550KThis dataset is obtained from a UniProt search
for protein sequences with family and binding site annotations. The dataset includes unreviewed (TrEMBL) protein sequences as well as
reviewed sequences. We refined the dataset by only including sequences with an annotation score of 4. We sorted and split by family, where
random families were selected for the test dataset until approximately 20% of the protein sequences were separated out for test data.
We excluded any sequences with <, >, or ?… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/binding_sites_random_split_by_family_550K.SiPaKosa-Sent
SiPaKosa: Sinhala-Pali Buddhist Corpus
A comprehensive corpus of canonical and classical Buddhist texts in Sinhala and Pali, compiled from historical archives and web-scraped canonical scriptures.
This is the sentence-level version of the SiPaKosa dataset.
Where SiPaKosa contains book level text, this dataset has all sentences by book.
Related dataset (book-level): RaniduG/SiPaKosa
Dataset Statistics
Total Sentences: 786,344
Sinhala Sentences: 465,539 (59.2%)
Mixed… See the full description on the dataset page: https://huggingface.co/datasets/RaniduG/SiPaKosa-Sent.GSM-Ranges
GSM-Ranges Dataset
📄 Paper: Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges🔗 GitHub Repository: GSM-Ranges GitHub
What is GSM-Ranges?
GSM-Ranges is a dataset generator built upon the GSM8K benchmark. It systematically modifies numerical values in math word problems to assess the robustness of large language models (LLMs) across a broad spectrum of numerical scales. By introducing numerical… See the full description on the dataset page: https://huggingface.co/datasets/guactastesgood/GSM-Ranges.uni-rankings-2026
BrightKey Independent University Rankings Dataset (2026)
299 universities × 55 countries × 6 dimensions, evaluated independently. No payments from institutions accepted. Public data only.
This is the open release of the BrightKey university rankings — an independent alternative to QS, THE, and Shanghai rankings. Released under CC BY 4.0.
Live site: https://brightkey.co/en/rankings/methodology
GitHub repo: https://github.com/arthurb2l/brightkey-university-dataset
Zenodo DOI:… See the full description on the dataset page: https://huggingface.co/datasets/brightkey/uni-rankings-2026.open-domain-ranks
Linkheft Open Domain Ranks: free domain authority data for 10.3 million domains
Try the paid tool: Linkheft on Apify: score any domain list via API, with 4-month trends. First try costs cents; pay only for results.
Buy: Velbrake Pro – Personal (1 site) ($39.00/year): for site owners: control which AI crawlers can read your WordPress site (the base plugin is free). Checkout by Polar.
An open alternative to proprietary "domain authority" scores. For each of the top 10… See the full description on the dataset page: https://huggingface.co/datasets/CyberMax-tools/open-domain-ranks.Mr-GSM8KView the project page:
https://github.com/dvlab-research/DiagGSM8K
see our paper at https://arxiv.org/abs/2312.17080
Description
In this work, we introduce a novel evaluation paradigm for Large Language Models,
one that challenges them to engage in meta-reasoning. Our paradigm shifts the focus from result-oriented assessments,
which often overlook the reasoning process, to a more holistic evaluation that effectively differentiates
the cognitive capabilities among models. For… See the full description on the dataset page: https://huggingface.co/datasets/Randolphzeng/Mr-GSM8K.Sri_Lankan_UGC_Cutoff_Mark_Dataset
Sri Lankan University Z-Score Recommendations
This dataset supports the Z-Score University Finder project, providing recommendations for Sri Lankan Advanced Level (A/L) students based on their Z-Score, academic stream, and district. It is designed for educational research and building recommendation systems for university admissions.
Provenance
Sources
The dataset is derived from aggregated historical university admission data from the University Grants… See the full description on the dataset page: https://huggingface.co/datasets/kasi-ranaweera/Sri_Lankan_UGC_Cutoff_Mark_Dataset.tooldrift-model-rankings
ToolDrift: OpenRouter model usage rankings, captured daily
One row per model per ranking window per capture: its rank, the tokens and requests behind that rank, and its share of the window. The series shows which models the market actually routes work to, day by day.
Rows in this cut
44,369
One row is
one model in one ranking window on one capture day
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-model-rankings.tooldrift-app-rankings
ToolDrift: OpenRouter app usage rankings, captured daily
One row per app per ranking window per capture, with the tool it maps to where ToolDrift tracks one. It is the same series as the model rankings, read from the consumer side.
Rows in this cut
641
One row is
one app in one ranking window on one capture day
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
ToolDrift
Method
https://toolproof.thecompound.tech/methodology
Licence… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/tooldrift-app-rankings.tourism-package-predictionrandom-sentence-v2
Random Sentences Dataset (version 2)
This is a random sentence dataset, it has random simplistic children sentences paired with completely unrelated random words. This dataset was used in SmolBabble2-360m, an AI model that spits out random sentences regardless of what is said to it. It is an improved version of the previous dataset, and it has much more examples. You can find more information on both of these models on the provided repository links.
The dataset has prompt and… See the full description on the dataset page: https://huggingface.co/datasets/benni-ben/random-sentence-v2.Islamweb_part2past-setting-ede8fc
past-setting-ede8fc
Synthetic sensors test data: 51 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/randallkaren9/past-setting-ede8fc.qwen35-9b-question-first-coop-random-50
What this is
Cooperative two-agent coding dataset: 49 task pairs across 15 repos (random-50 subset), generated
with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a question-first prompt variant —
agents begin by asking each other clarifying questions about their respective features before
starting implementation, aiming to surface integration concerns early. All 49 pairs were
successfully evaluated.
At a glance
Field
Value
Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-question-first-coop-random-50.poker-gto-strategy-and-hand-rangesgenome-classification-with-random100M-random-promotersBoer, Carl G. de, Eeshit Dhaval Vaishnav, Ronen Sadeh, Esteban Luis Abeyta, Nir Friedman, and Aviv Regev. 2020. “Deciphering Eukaryotic Gene-Regulatory Logic with 100 Million Random Promoters.” Nature Biotechnology 38 (1): 56–65. https://doi.org/10.1038/s41587-019-0315-8.
https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE104878
VSR_random_tsvindia-ev-range-dataset
🔋 India EV Real-World Range Dataset
A comprehensive dataset of 7,291 data points covering 61 Indian 4-wheeler EV variants from 18 manufacturers across 12 Indian driving scenarios.
Dataset Description
This dataset was built to predict the real-world driving range of Electric Vehicles under Indian conditions. It combines:
Real EV specifications from all major EVs sold in India (Tata, Mahindra, Hyundai, MG, Kia, BYD, BMW, Mercedes, etc.)
Physics-based energy modeling using… See the full description on the dataset page: https://huggingface.co/datasets/SohamThakkar-07/india-ev-range-dataset.tibetan_ranking_training_data
Tibetan Ranking Training Data — Tier 1
Training data for a cross-encoder that re-ranks Tibetan→English translation candidates by contextual relevance. Each row pairs a Tibetan term and sentence context with one English gloss candidate and a binary label: 1 (this gloss is defensible in this context) or 0 (it is not).
Used to train trabten/tibetan_ranking_BUDA.
Files
File
Rows
Description
training_v2_merged.csv
71,050
Tier 1 — surgically annotated… See the full description on the dataset page: https://huggingface.co/datasets/trabten/tibetan_ranking_training_data.qwen35-9b-contract-first-coop-random-50
What this is
Cooperative two-agent coding dataset: 36 task pairs across 13 repos (random-50 subset), generated
with mini_swe_agent on Qwen/Qwen3.5-9B in coop setting using a contract-first prompt variant —
agents first agree on a shared interface contract (function signatures, data structures, API
boundaries) before independently implementing their respective features. All 36 pairs were
successfully evaluated.
At a glance
Field
Value
Model… See the full description on the dataset page: https://huggingface.co/datasets/CooperBench/qwen35-9b-contract-first-coop-random-50.ner-email
Overview:
This dataset is augmented through llama-3.1:8b. The pourpose is to finetune llm for token classification i.e Email in our case.
Following tags are present in dataset:
full_name : 1
email : 2
gender : 3
city : 4
country : 5
