datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BAT
BAT Dataset
This dataset provides an alternative way to access the data from the BAT (BAT: Benchmark for Auto-bidding Task) autobidding benchmark.
Related Resources
GitHub Repository: avito-tech/bat-autobidding-benchmark
Paper: BAT: Benchmark for Auto-bidding Task
Dataset Description
This dataset contains auction data for First-Price Auction (FPA) and Vickrey-Clarke-Groves (VCG) mechanisms, used for benchmarking autobidding algorithms.
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/AvitoTech/BAT.TR-News
Citation
If you use the dataset, please cite the paper:
@article{10.1007/s10579-021-09568-y,
year = {2022},
title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}},
author = {Baykara, Batuhan and Güngör, Tunga},
journal = {Language Resources and Evaluation},
issn = {1574-020X},
doi = {10.1007/s10579-021-09568-y},
pages = {1--35}}
BATIS
license: cc-by-nc-4.0
BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models
This repository contains the dataset used in experiments shown in BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models. To download the dataset, you can use the load_dataset function from HuggingFace. For example :
from datasets import load_dataset
# Training Split for Kenya
training_kenya = load_dataset("cathv/BATIS", name="Kenya"… See the full description on the dataset page: https://huggingface.co/datasets/cathv/BATIS.pokemon-battle-outcomesclash-royale-battlesbatman2
What is the purpose of this dataset?
This dataset is a reformatted replica of BATMAN-2.0, a database of traditional Chinese medicine (TCM) ingredients, herbs, and formulas.
Refer to the original peer-reviewed publication for more detail.
What can I do with this dataset?
Check out the demo notebooks in the demo folder.
BATMAN is a powerful source of data that links natural compounds and their compositions to protein targets.
You can use this information to run drug… See the full description on the dataset page: https://huggingface.co/datasets/f-galkin/batman2.africanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
pokemon-battle-statsHU-News
Citation
If you use the dataset, please cite the paper:
@article{10.1007/s10579-021-09568-y,
year = {2022},
title = {{Abstractive text summarization and new large-scale datasets for agglutinative languages Turkish and Hungarian}},
author = {Baykara, Batuhan and Güngör, Tunga},
journal = {Language Resources and Evaluation},
issn = {1574-020X},
doi = {10.1007/s10579-021-09568-y},
pages = {1--35}}
job_resume_fit
Resume-Job Fit Dataset
Description
This dataset contains 2385 resumes matched to 23 different job categories. For each job posting and resume pair, skill matching is evaluated using three different scores: direct AI-based skill matching, string-based skill matching, and fuzzy token matching. The resumes are sourced from the Resume Dataset Source. Each row contains a candidate's resume, the related job posting, its category, and various matching scores.… See the full description on the dataset page: https://huggingface.co/datasets/batuhanmtl/job_resume_fit.sycl
SyCL Data
This is the data used for experiments in Beyond Contrastive Learning: Synthetic Data Enables List-wise Training with Multiple Levels of Relevance paper.
Synthetic Data
The data is created from MS MARCO queries shared in mteb/msmarco repo.
Directories llama33_70b, qwen25_72b, and qwen25_32b contain data generated with Llama 3.3 70B, Qwen2.5 72B, and Qwen2.5 32B, respectively.
The data format in each subdirectory is the same as mteb/msmarco repo (trec20 is the… See the full description on the dataset page: https://huggingface.co/datasets/BatsResearch/sycl.mlb-statcast-battersthe-Pokemon-Trading-Card-Game-Battle-Challenge-DataSetpokemon_tcg_battle_dataset.csv (~6–7k rows from 500 simulated games by default).
the-Pokemon-Trading-Card-Game-Battle-Challenge-DataSet-SampleGenerating a small synthetic dataset
cung-phi-bat-trach
Cung phi và hướng Bát Trạch
Kua number and Bat Trach directions
1. Mô tả · Description
Cung phi theo năm sinh và giới tính cho khoảng 1900 tới 2099, kèm bốn hướng tốt và bốn hướng cần tránh.
Kua number by birth year and sex for 1900 to 2099, with the four favourable and four unfavourable directions.
Số dòng · Rows: 400
Phiên bản · Version: 1.1.0 (2026-09-16)
Mã hoá · Encoding: UTF-8 không BOM
2. Cấu trúc · Structure
Cột · Column
Kiểu · Type… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/cung-phi-bat-trach.metacognitive-monitoring-battery
Metacognitive Monitoring Battery
A cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework.
Paper: The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring
Code: github.com/synthiumjp/metacognitive-monitoring-battery
Author: Jon-Paul Cacioli (Independent Researcher, Melbourne, Australia)
Overview
The battery comprises 524 items across six cognitive domains… See the full description on the dataset page: https://huggingface.co/datasets/synthiumjp/metacognitive-monitoring-battery.paper-abstracts
Battery Abstracts Dataset
This dataset includes 29,472 battery papers and 17,191 non-battery papers, a total of 46,663 papers. These papers are manually labelled in terms of the journals to which they belong. 14 battery journals and 1,044 non battery journals were selected to form this database.
training_data.csv: Battery papers: 20,629, Non-battery papers: 12,034. Total: 32,663.
val_data.csv: Battery papers: 5,895, Non-battery papers: 3,438. Total: 9,333.
test_data.csv: Battery… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/paper-abstracts.ipulse-ai-batch5-advisor-forecast-panel
iPulse AI Batch 5 Advisor Forecast Panel
This dataset exposes a compact, anonymized panel of production forecasts from iPulse AI, Future Edge Group's Open Agentic Investment Research Platform. It is designed for research on forecast combination, disagreement, correlated errors, regime dependence, and the effective number of independent forecasters.
The release contains seven showcase assets, twelve advisor configurations per asset, quarterly forecast paths extending five years… See the full description on the dataset page: https://huggingface.co/datasets/future-edge-group/ipulse-ai-batch5-advisor-forecast-panel.jacobian-research-trajectory-batch-001
Jacobian research trajectory — batch 001
A flattened archival extraction of research attempts concerning polynomial maps in two variables with nonzero constant Jacobian determinant, using Research Trajectory Schema v1.0.0.
Contents
research_trajectory_flat.csv contains 207 records and 145 columns. Each row represents one canonical record: 1 problem, 24 branches, 55 attempts, 56 evaluations, 19 relations, 3 decisions, or 49 artifacts.
Nested object fields use… See the full description on the dataset page: https://huggingface.co/datasets/amphora/jacobian-research-trajectory-batch-001.ADeLe_battery_v1dot0
Dataset Card for ADeLe
Dataset Summary
ADeLe (Annotated-Demand-Levels) battery is a single, unified test set whose every item is labelled with the level (0-5+) it demands on 18 general ability dimensions (e.g. attention and scan, logical reasoning, various knowledge areas) plus an “unguessability” dimension. It is produced by applying the DeLeAn rubrics, via GPT-4o annotators, to AI benchmarks.
Version 1.0 contains 16 108 items drawn from 63 tasks spread across a diverse… See the full description on the dataset page: https://huggingface.co/datasets/CFI-Kinds-of-Intelligence/ADeLe_battery_v1dot0.bulgarian_homographsБългарски Омографи с IPA и Ударения (1016 думи)
(Ако откриете грешки моля пишете в Community. Благодаря! Щом подготвя ще има и нови)
(Има и омографи в неформатиран вид 9800 думи в който са включени и обработените.
Няма цензурирани, но токсичните няма да бъдат описани макар да са част от БАН и БЕРОН)
За повече информация виж - Zakonite_na_Omografa.md
Описание
Този набор от данни съдържа 1016 уникални български омографа – думи, които се изписват еднакво, но се произнасят различно поради… See the full description on the dataset page: https://huggingface.co/datasets/batvanio12/bulgarian_homographs.AI_BATTERY_OPTIMIZER
Dataset Card for AI Battery Optimizer
The AI Battery Optimizer Dataset contains synthetic smartphone battery usage logs created during the development of the AI Battery Optimizer App.It is intended for research and experimentation on battery prediction, app usage forecasting, and adaptive resource management.
Dataset Details
This dataset logs:
Battery percentage over time
Power usage (mW)
Estimated time remaining
Predicted app usage with confidence score
Screen… See the full description on the dataset page: https://huggingface.co/datasets/yu743/AI_BATTERY_OPTIMIZER.diem-bat-dong-giua-cac-phai
Điểm bất đồng giữa các trường phái
Where the schools disagree
1. Mô tả · Description
Từng chỗ mà các nguồn lịch pháp, phong thuỷ và bói toán Việt Nam không thống nhất, kèm cách bộ dữ liệu này chọn, cách khác đang lưu hành, và hệ quả khi đối chiếu hai nguồn. Mọi bảng tra đều phải chọn một cách ở những chỗ ấy; phần lớn chọn rồi im lặng.
Each point where Vietnamese calendrical, feng shui and divination sources disagree, with the reading taken here, the competing… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/diem-bat-dong-giua-cac-phai.uaetaxlawdataBATIS
license: cc-by-nc-4.0
BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models
This repository contains the dataset used in experiments shown in BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models. To download the dataset, you can use the load_dataset function from HuggingFace. For example :
from datasets import load_dataset
# Training Split for Kenya
training_kenya = load_dataset("anonsubmit/BATIS", name="Kenya"… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmit/BATIS.AI_BATTERY_OPTIMIZER
Dataset Card for AI Battery Optimizer
The AI Battery Optimizer Dataset contains synthetic smartphone battery usage logs created during the development of the AI Battery Optimizer App.It is intended for research and experimentation on battery prediction, app usage forecasting, and adaptive resource management.
Dataset Details
This dataset logs:
Battery percentage over time
Power usage (mW)
Estimated time remaining
Predicted app usage with confidence score
Screen… See the full description on the dataset page: https://huggingface.co/datasets/Yonne819/AI_BATTERY_OPTIMIZER.Batch_indexing_machine_tokensTR_synt_medical_multiple_choice_QA_with_reasoningbatik-id-tegalAI_BATTERY_OPTIMIZER
Dataset Card for AI Battery Optimizer
The AI Battery Optimizer Dataset contains synthetic smartphone battery usage logs created during the development of the AI Battery Optimizer App.It is intended for research and experimentation on battery prediction, app usage forecasting, and adaptive resource management.
Dataset Details
This dataset logs:
Battery percentage over time
Power usage (mW)
Estimated time remaining
Predicted app usage with confidence score
Screen… See the full description on the dataset page: https://huggingface.co/datasets/wycwsjdw/AI_BATTERY_OPTIMIZER.
