datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
golden-batch-sentinel-data
Golden Batch Sentinel Data
Benchmark datasets for process monitoring and fault detection in batch manufacturing.
Datasets
IndPenSim (Industrial Penicillin Simulation)
A 100,000L fermentation simulation with 100 batches and rich multivariate signals.
Source: Mendeley Data
Paper: Modern day monitoring and control challenges...
Batches: 100 (90 normal, 10 faulty)
Variables: 37 process variables (Raman spectra excluded for efficiency)
Time resolution: 0.2 hours… See the full description on the dataset page: https://huggingface.co/datasets/foundation-models/golden-batch-sentinel-data.ai_arena_udtraek
Danish data from AI-Arenaen
This data is an extraction of conversations and reactions from the website AI-Arenaen.dk, taken from datasets published by ComparIA.
Source
The dataset is the Danish subset of:
'ministere-culture/comparia-conversations'
'ministere-culture/comparia-reactions'
Both comparia datasets are released under the cc-by compatible license: Etalab Open License 2.0 (etalab-2.0)
We extracted the data on 28/01/2026. In the future ai-arena dat, AI-arena… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai_arena_udtraek.time-series-foundation-models-papers
Time Series Foundation Models Papers — FineSet
A research-paper dataset on Time Series Foundation Models Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Time Series Foundation Models Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/time-series-foundation-models-papers.training-output
Training Output
Tabular run summaries exported from AlphaOPT evaluation (all_test_results.json).
Data
default split: one row per benchmark run with metrics (pass rates, tokens, duration, config flags, paths).
Format on the Hub
Rows are stored as Apache Parquet (Hugging Face datasets default), which is efficient for analytics and the Dataset viewer.
Source
Generated locally under AlphaOPT/testing/; re-export if you re-run evaluations.
croco-munin-apertus-8b-da-simpo-fullcroco-munin-apertus-8b-da-50kcroco-munin-apertus-8b-da-simpo-full-50kcroco-munin-apertus-8b-da-generatedai-arenaen-conversations
AI Arenaen Conversations
A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform.
Origin of the data: what is AI-Arenaen?
The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.croco-munin-apertus-8b-da-goldkaenguruen
Kænguruen Danish Math Competition
A dataset of multiple-choice math problems from the Danish Kangaroo math competition
(Matematikkens Kænguru), a popular international mathematics contest held annually in
Denmark for students in grades 4–9.
Dataset description
Kænguruen originates from France (1991) and is now held in 100+ countries with around
6 million participants per year. In Denmark it is organized by
Danmarks Matematiklærerforening and takes place on the third… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/kaenguruen.croco-munin-apertus-8b-da-simpo-tunedmedical-ai-clinical-foundation-models-2026
🏥 Medical AI & Clinical Foundation Models Dataset (2026 Edition)
Sample dataset of 30 audit-verified research papers covering Clinical LLMs, Medical Foundation Models, Radiology/Pathology Vision, and EHR Intelligence with 384d PyTorch embeddings.
🛒 Full 1,000 Paper B2B Dataset Available on Gumroad
Get the complete 1,000 paper dataset (1,000 papers + OpenAlex Citations + IP Safety Score + SQLite/CSV/Parquet + Quickstart Script) on Gumroad:
👉 Get Full 1,000… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/medical-ai-clinical-foundation-models-2026.croco-munin-apertus-8b-dacroco-munin-apertus-8b-da-simpocroco-munin-apertus-8b-da-lsai-arenaen-votes
