datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HeliumSLM-Datasetshelf-photos-batch1
shelf-photos-batch1
Working dataset for the shelf-monitoring pipeline: bootstrap labeling and
fine-tuning data management for Geraldine/rf-detr-nano-bookshelf.
Layout
photos/ — 25 of the library's own shelf photos, untouched (no crops, no upscaling)
original_images_library/ — 285 high-res images from
llabres/library-dataset (MIT),
used for continuous fine-tuning (domain shift: real library stacks)
dataset/Bookshelf-recognition-2.v1i.coco.zip — COCO export of… See the full description on the dataset page: https://huggingface.co/datasets/Geraldine/shelf-photos-batch1.pos_tagging
POS Tagging Dataset
Original Data Source
Conll2003
E. F. Tjong Kim Sang and F. De Meulder, Proceedings of the
Seventh Conference on Natural Language Learning at HLT-
NAACL 2003, 2003, pp. 142–147.
The Peen Treebank
M. P. Marcus, B. Santorini and M. A. Marcinkiewicz, Comput.
Linguist., 1993, 19, 313–330.
Citation
BatteryDataExtractor: battery-aware text-mining software embedded with BERT models
chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.openai-terra-batch-wiki-brazil-1000-partial-20260724-01
OpenAI Terra Batch — Wikipédia PT-BR (run parcial)
Checkpoint publicável de uma execução real e interrompida do fluxo
document_task_matrix. A execução planejou gerar uma matriz de 1.000
documentos da Wikipédia em português por 25 tasks canônicas usando a Responses
API Batch e o modelo gpt-5.6-terra.
Este repositório não representa a conclusão dos 25.000 pares planejados. Ele
contém somente os 1.282 candidatos aceitos após a reconciliação offline de
todos os resultados Batch já… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/openai-terra-batch-wiki-brazil-1000-partial-20260724-01.opcder-batchfr-benign-batchlight-batch-summarize-dialogue
Light dataset
Dialogues are preprocessed into a form:
<Character name>: <character line>
...
<Character name>: <character line>
Summarize the document
bat_genomebatch-effects-leaderboard-resultsbattery-device-data-qa
Battery Device QA Data
Battery device records, including anode, cathode, and electrolyte.
Examples of the question answering evaluation dataset:
{'question': 'What is the cathode?', 'answer': 'Al foil', 'context': 'The blended slurry was then cast onto a clean current collector (Al foil for the cathode and Cu foil for the anode) and dried at 90 °C under vacuum overnight.', 'start index': 645}
{'question': 'What is the anode?', 'answer': 'Cu foil', 'context': 'The blended slurry was… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/battery-device-data-qa.aimodel-sft-v1kids-multilingual-benchmark
TinyAya v2 — Multilingual Benchmark for Children's AI Companions
2,312 child–AI conversational prompts across 23 languages, evaluated against
four models with five-judge LLM-as-judge validation.
📄 Companion article: see HF Articles by @batuhanaktas.
💻 Code: https://github.com/aktasbatuhan/cohere-tiny-aya-for-kids
Dataset summary
This dataset contains:
benchmark/items.jsonl — 2,312 benchmark items in 23 languages. Each item
is a structured prompt designed to mimic… See the full description on the dataset page: https://huggingface.co/datasets/batuhanaktas/kids-multilingual-benchmark.3D-CoT
3D-CoT Benchmark: Chain-of-Thought Datasets for 3D Point Cloud-Language Models
Overview
The 3D-CoT Benchmark is a structured reasoning dataset designed explicitly to facilitate the systematic study of Chain-of-Thought (CoT)'s impact on 3D vision-language alignment. By extending existing 3D datasets with carefully structured reasoning annotations, this benchmark enables rigorous exploration of multimodal reasoning capabilities, significantly enhancing interpretability and… See the full description on the dataset page: https://huggingface.co/datasets/Battam/3D-CoT.Llama-3-70b-battlesChatbot Arena user conversations between Llama-3-70b VS GPT-4-1025 or Llama-3-70b VS Claude-3-Opus with user preference votes. Single turn. Excludes ties.
Used in Llama Data Analysis blog post and "VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models" (Paper, Code).
Citation
@article{dunlap_vibecheck,
title={VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models},
author={Lisa Dunlap and Krishna Mandal and Trevor… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/Llama-3-70b-battles.2026.AC.Copy-Coordination-Battery
2026.AC.Copy-Coordination-Battery
Full call-level transcripts and the pooled analysis for a preregistered battery asking whether isolated copies of one language model coordinate with each other without any communication channel, and whether that coordination is decision-theoretic (FDT/UDT-style reasoning about being a copy) or merely shared focal points that any two similar models would share.
38,204 records across four subject models — claude-fable-5, claude-opus-5… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.AC.Copy-Coordination-Battery.Rehber-CoT-Science
🧬 Rehber-CoT-Science: Turkish Scientific Reasoning Dataset
Turkish Scientific Computational Reasoning (Chain-of-Thought) Dataset
Multi-step scientific problem-solving dataset with verifiable Python code and detailed explanations
Dataset • Author
📌 Changelog
Eski sürümlere erişim: Branch menüsünden v1 seçebilirsiniz.
Version
Date
Changes
v2.0
24.12.2025
✨ Yeni explained_answer alanı eklendi, Statistics domain eklendi, 712 örneğe genişletildi… See the full description on the dataset page: https://huggingface.co/datasets/batuhanozkose/Rehber-CoT-Science.torchgbif-batches-sampleFinModernBERT-pairs-sec-synthetic-v1
FinModernBERT-pairs-sec-synthetic-v1
115,238 synthetic finance contrastive-training pairs generated from SEC EDGAR filings —
the finance portion of the training set behind
FinModernBERT-embed-large-v1,
the 395M embedding model that beats the 7B Fin-E5 on FinMTEB Summarization (+0.109) and
STS (+0.010).
To our knowledge this includes the first public document↔summary positive + mismatched
negative pair set built specifically for the FinMTEB Summarization task shape.… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/FinModernBERT-pairs-sec-synthetic-v1.Test_batchsmoke-openai-terra-batch-brasil-25-20260724-01
Smoke OpenAI Terra Batch — Brasil × 25 tasks
Run real de validação do fluxo matricial document_task_matrix, executada
sobre um único documento da Wikipédia em português com o título Brasil.
Cada uma das 25 tasks canônicas recebeu exatamente um slot inicial.
Resultado
status: completed
documentos: 1
pares planejados: 25
exemplos aceitos: 25
pares pulados: 0
pares esgotados: 0
resultados reais do backend: 27
retries com nova chamada: 2
backend: openai_api… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/smoke-openai-terra-batch-brasil-25-20260724-01.dataset-viber-image-generation-preference-inference-endpoints-battle-flux
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.bbnaija2026-predictionssolar-inverter-panel-battery-compatibility
Solar inverter–battery compatibility
Canonical, always-current version: https://referencesource.org/solar-inverter-panel-battery-compatibility/
Machine-readable: https://referencesource.org/solar-inverter-panel-battery-compatibility/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-12
Stale after: 2027-02-08 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 101
Which lithium batteries have been… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/solar-inverter-panel-battery-compatibility.battlefield-medic-sharegpt
🏥⚔️ Synthetic Battlefield Medical Conversations
For the multilingual version (non-sharegpt foormat) that includes the title columns go here https://huggingface.co/datasets/nisten/battlefield-medic-multilingual
Over 3000 conversations incorporating 2000+ human diseases and over 1000 battlefield injuries from various scenarios
Author: Nisten Tahiraj
License: MIT
This dataset consists of highly detailed synthetic conversations… See the full description on the dataset page: https://huggingface.co/datasets/nisten/battlefield-medic-sharegpt.abbreviation_detection
Abbreviation Detection Dataset
Original Data Source
PLOS
I. Zilio, H. Saadany, P. Sharma, D. Kanojia and C. Orasan,
PLOD: An Abbreviation Detection Dataset for Scientific Docu-
ments, 2022, https://arxiv.org/abs/2204.12061.
SDU@AAAI-21
A. P. B. Veyseh, F. Dernoncourt, Q. H. Tran and T. H. Nguyen,
Proceedings of the 28th International Conference on Compu-
tational Linguistics, 2020, pp. 3285–3301
Citation
BatteryDataExtractor: battery-aware… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/abbreviation_detection.battery-dataset-catalog
BSEBench Battery Dataset Catalog
This repository is the public discovery layer for BSEBench battery datasets.
It is not a raw-data mirror. Raw files become official BSEBench datasets
only after a strict manifest exists in bsebench-datasets/manifests/
with verified SHA-256 checksums and Hugging Face Tier 1 storage.
Prospect records: 28
Named datasets / variants indexed: 204
Generated on: 2026-05-01
Files
dataset_prospects.yaml: canonical hand-curated catalog… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/battery-dataset-catalog.lithium-battery-shipping-limits
Lithium battery shipping limits (49 CFR 173.185)
Canonical, always-current version: https://referencesource.org/lithium-battery-shipping-limits/
Machine-readable: https://referencesource.org/lithium-battery-shipping-limits/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-03
Stale after: 2027-01-30 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 15
Watt-hour and lithium-content thresholds, quantity… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/lithium-battery-shipping-limits.battle-tested-regex-validations
Battle-Tested Regex Validations
A curated collection of 25+ reliable, tested regular expressions for common validation tasks such as emails, URLs, and phone numbers. Each entry includes the pattern, a test input, and a note explaining its real-world applicability. Ideal for developers who need dependable validation logic without reinventing the wheel.
30 rows · category: utility · licence: CC0-1.0 (public domain)
Usage
import sys
sys.path.insert(0, ".")
from tool… See the full description on the dataset page: https://huggingface.co/datasets/SharkSkin/battle-tested-regex-validations.kashmiri_clean_batch_4.1
