datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linguistic-similaritysimverse2026
SimVerse
⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review.
A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.sf-index-history
SimpleFunctions Index History
Time series of the SF Index: a four-number summary of prediction-market consensus — disagreement (0-100), geo-risk (0-100), breadth (-1..+1), and activity (0-100) — computed every 15 minutes from ~50K markets. Flat JSONL for easy charting / analysis.
License and Use
This dataset is released under Creative Commons Attribution 4.0 International
(CC-BY-4.0; https://creativecommons.org/licenses/by/4.0/). You may use it
freely for personal… See the full description on the dataset page: https://huggingface.co/datasets/SimpleFunctions/sf-index-history.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.lm-similarity
Great Models Think Alike and this Undermines AI Oversight
This is the data collection for the publication "Great Models Think Alike and this Undermines AI Oversight."
judge_scores_mmlu_pro_free_filtered: Judge scores of nine judges without access to the reference answers on the filtered, open-style MMLU-Pro dataset.
judge_w_gt_mmlu_pro_free_filtered: Ensemble judge scores of five judges with access to the reference options and ground-truth information on the filtered OSQ MMLU-Pro.… See the full description on the dataset page: https://huggingface.co/datasets/bethgelab/lm-similarity.chinese-ai-and-robotics-open-intelligence
🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.age-specific-text-simplification
Age-Specific Text Simplification Dataset
Dataset Description
This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group.
Dataset Summary
Total Examples: 17,177
Training Split: 15,459 examples
Validation Split: 1,718 examples
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.SimpleS2
SimpleS2
To load the data:
import json
import pickle
# Read cube
with open('cubo1_pickle', 'rb') as file:
data = pickle.load(file).to_dataset(dim='band')
# Read metadata
with open('cubo1.json') as f:
meta = json.load(f)
Citation
This dataset is related to the paper: arXiv:2506.196560.8k-data-SimpleDeepSearchernvidia_openmathinstruct-2-simple-processed元データ
https://huggingface.co/datasets/nvidia/OpenMathInstruct-2
VidChain-Datageo_wikipedia_geonamesSimulBenchTTTXXX01__Mistral-7B-Base-SimPO2-5e-7-details
Dataset Card for Evaluation run of TTTXXX01/Mistral-7B-Base-SimPO2-5e-7
Dataset automatically created during the evaluation run of model TTTXXX01/Mistral-7B-Base-SimPO2-5e-7
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TTTXXX01__Mistral-7B-Base-SimPO2-5e-7-details.video_training_v1
Video Training Dataset (v1)
Private dataset. Mixed-provenance research data. Not cleared for public
redistribution — see Provenance & License before changing visibility.
Generated by scripts/build_video_training_dataset.py.
Contents
Videos: 648 (≈0.94 h, ~1.9 GiB)
NPZ training pairs: 79 (causal_forcing: 33, pframe_npz: 46; ~12 GiB)
Total payload: ~13.9 GiB
Splits — train: 577, val: 36, test: 35
Provenance & License
This dataset mixes several… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/video_training_v1.SimulBench-resultsblockchain-simspanish_simplification_20k
Dataset Card: spanish_simplification_20k
This dataset provides 20K samples for Spanish simplification across a diverse range of topics, designed for people with cognitive challenges. The samples vary in length, from short to long. This dataset can supplement a larger training corpus for fine-tuning small language models, such as Gemma 4B, to simplify complex Spanish into simpler Spanish or translate English directly into simplified Spanish.
Dataset Contributors… See the full description on the dataset page: https://huggingface.co/datasets/khaledmahmoud/spanish_simplification_20k.vitl
Model Card for DINOv3
DINOv3 is a family of versatile vision foundation models that outperforms the specialized state of the art across a broad range of settings, without fine-tuning. DINOv3 produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models.
Model Details
These are Vision Transformer and ConvNeXt models trained following the method described in… See the full description on the dataset page: https://huggingface.co/datasets/simon123905/vitl.nerel_simpleultimate_epic_battle_simulator_recordings_01
史诗战争模拟器 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_2423771c06c259fd62c5d065481aca95
Collection: general (泛数据)
Recordings: 22
Layout: recordings/<recording_id>/<raw component>
Simple-agent-traces
📱 Simple Agent Traces – Tiny Tool‑Calling Conversations for Small Models
Simple Agent Traces is a compact, hand‑picked dataset of 605 real‑world tool‑calling conversations, each carefully truncated to ≤8,192 tokens (using the SmolLM2‑360M tokenizer).It is purpose‑built for training and fine‑tuning tiny language models (≤500M) that must run on‑device – smartphones, edge devices, or any environment with strict memory and latency constraints.
🧹 No chain‑of‑thought, no fluff.Every… See the full description on the dataset page: https://huggingface.co/datasets/LiteMind/Simple-agent-traces.qwen3-1.7b-features-similar-k100000countdown-qwen3-0.6b
Countdown Qwen3-0.6B Pass@10 Buckets
Countdown arithmetic problems filtered by observed local Qwen/Qwen3-0.6B success rate over 10 rollouts per problem.
Each problem asks for an arithmetic expression that reaches a target using each listed source number at most once. The final answer should be inside \boxed{...}. Canonical solutions are provided, but any verifier-valid expression is accepted.
Subsets
subset
source bucket
count
observed successes out of 10… See the full description on the dataset page: https://huggingface.co/datasets/simpissa/countdown-qwen3-0.6b.security_dataAALF__gemma-2-27b-it-SimPO-37K-details
Dataset Card for Evaluation run of AALF/gemma-2-27b-it-SimPO-37K
Dataset automatically created during the evaluation run of model AALF/gemma-2-27b-it-SimPO-37K
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AALF__gemma-2-27b-it-SimPO-37K-details.AALF__gemma-2-27b-it-SimPO-37K-100steps-details
Dataset Card for Evaluation run of AALF/gemma-2-27b-it-SimPO-37K-100steps
Dataset automatically created during the evaluation run of model AALF/gemma-2-27b-it-SimPO-37K-100steps
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AALF__gemma-2-27b-it-SimPO-37K-100steps-details.0.5k-data-SimpleDeepSearcher
