datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jax-gcm-data
jax-gcm boundary conditions and emissions
Input data for jax-gcm
(jcm), a fully differentiable atmospheric GCM in JAX. Two tiers:
products/ — grid-independent source products
store
contents
source
ceds_anthro.zarr
anthropogenic SO2/BC/OC/NH3 flux, sector-summed, 0.5°, monthly 1850–2023 + PI (1850–59) / PD (2005–14) climatologies
CEDS-CMIP-2025-04-18 (input4MIPs CMIP7)
bb4cmip7.zarr
open-burning SO2/BC/OC/NH3 flux, 0.25°, monthly 1850–2023 + PI/PD… See the full description on the dataset page: https://huggingface.co/datasets/climate-analytics-lab/jax-gcm-data.jax-fli-experiments
jax-fli experiments
Data, samples, and reference catalogs for the jax-fli
forward-modelling experiments. Each experiment is exposed as one or more
HuggingFace dataset configs; load a config with datasets.load_dataset.
Experiment 00 — CosmoGrid reference
A single CosmoGrid simulation (cosmo_000001) packaged as jax-fli Catalog
parquet files, used as reference truth for the lensing / masked-shear experiments.
Config
Field
Shape
NSIDE
Description… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/jax-fli-experiments.jaxgmg-sus-vis-data
jaxgmg-sus-vis-data
HTML plot storage for the
Susceptibility Visualizations
Space.
Structure
<model_name>/
comparison_*.html
<run_name>_nbeta:<N>_perturbation_type:<type>/
conv_vs_fc_direction.html
conv_blocks_3d_direction.html
fc_layers_3d_direction.html
pca_direction.html
Adding plots
cp -r /path/to/figs/<model_name> .
git add . && git commit -m "Add <model_name> plots" && git push
HundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例
HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.val-teammatessmoltalk-gemma3-1024eval-teammateszhwiki-latestThis repository demonstrates access to the latest Chinese Wikipedia corpora.
Download
You can download the latest Chinese Wikipedia dump from the following link:
Chinese Wikipedia Dump
English Wikipedia Dump (For reference)
Extraction
After you download the dump, you can extract the data using the following commands:
# install wikiextractor
pip install wikiextractor
# extract the data
wikiextractor --json -o <output_dir> zhwiki-latest-pages-articles.xml.bz2
Then, you… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/zhwiki-latest.jaxaht-hanabi
JaxAHT Partner Checkpoints
Trained partner and ego agent checkpoints for LARG/jax-aht. Includes Hanabi (full 5c/5r/25, mini 3c/3r/9) and LBF (Level-Based Foraging) BC-LSTM human proxies.
All checkpoints are Orbax PyTreeCheckpointer format.
Self-Play Baselines
Sampling mode, 256 episodes, mean over listed seeds.
Other-Play mini-Hanabi scores use the corrected color-hint action permutation (commit 1852d9e). Full Hanabi OP retraining in progress; current full Hanabi… See the full description on the dataset page: https://huggingface.co/datasets/lainwired/jaxaht-hanabi.DialogES
DialogES: An Large Dataset for Generating Dialogue Events and Summaries
简介
本项目提出一个对话事件抽取和摘要生成数据集——DialogES,数据集包括共 44,672 组多轮对话,每组对话采用自动的方式标注出对话事件和对话摘要。
该数据集主要用于训练对话摘要模型,研究者可采用单任务和多任务学习的方式利用本数据集。
收集过程
对话采集:本数据集中的对话数据收集自两个已有的开源数据集,即 NaturalConv 和 HundredCV-Chat;
事件标注:采用 few-shot in-context learning 的方式标注,人工标注出 5 个对话-事件样本,作为演示样例嵌入大模型的提示词中,然后引导模型对输入对话进行标注;
摘要标注:采用 zero-shot learning 的方式标注,编写提示词"请总结对话的主要内容:",要求大模型生成对话摘要;
自动标注:按照以上准则,本项目利用 Deepseek-V3… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/DialogES.Spotify_Million_Playlist_Dataset_ChallengeMultitask_Preplay_JaxMaze_models
Multitask Preplay — jaxmaze model data
Data for the paper "Multitask Preplay" (PNAS).
Analysis code: https://github.com/wcarvalho/multitask_preplay (branch pnas).
Splits: qlearning, usfa, dyna, preplay, her, bfs, dfs, greedy_euclidean, memory_based_euclidean.
Each split is also available as a top-level parquet file.
HundredCVs
百人简历数据集
HundredCVs: A Curriculum Vitae Dataset of 100 Young Chinese People
简介
本项目提出一个全新的中文简历数据集(HundredCVs),包含了 100 位青年的个人简历。HundredCVs 具有以下特点:
年轻化、多样性:数据集中的人物年龄分布在 15~30 岁之间,广泛涵盖了不同性别、不同职业、不同学历(高中至博士不等)。
结构完整:每份简历中的信息包括人物的个人名片、性格特征、主要事迹,以及详细经历/个人自述等。
安全性:我们使用化名替代了人物的真实姓名,此外,人物经历也使用大语言模型的改写和提炼,表现出标准化和一致性的语言风格。
设计意图
HundredCVs 的提出主要是为了方便研究者开展基于简历的自然语言处理任务。数据集中提供两个文件:
profile.json:每条记录只包含人物的个人名片、性格特征和主要事迹。可用于开展角色扮演、人物画像构建等任务。… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCVs.example-dataset
ORB Transformation Applied on diffusiondb Dataset
This dataset consists of images, captions and images that are transformed to extract features using ORB transform.
You can find the original dataset here.
An example sample is below:
Caption: "spider - man, cinematic, photography "
Image:
Transformation:
eval-teammates-brdatasetseval-teammates-br-sampleThe sample was created from jaxaht/eval-teammates-br.
One best-response policy was sampled from each environment configuration.
For different variations of the same environment (e.g. lbf_7x7_nolevels and lbf_12x12), the best-response to the same type of agent was selected.
genai-ml-2025-hw7-evalLite-Thinking
Lite-Thinking: A Large-Scale Math Dataset with Moderate Reasoning Steps
Motivation
With the rapid popularization of large reasoning models, like GPT-4o, Deepseek-R1, and Qwen3, there are increasing researchers seeking to build their own reasoning models.
Typically, small foundation models are chosen; following the mature technology of Deepseek-R1, mathematical datasets are mainly adopted to build training corpora.
Despite existing available datasets, represented by… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/Lite-Thinking.IFEval-gemma3-chat
Dataset Card for Dataset Name
This dataset is a subset of google/IFEval, selected by the token length of applying chat template of google/gemma-3-4b-it.
Dataset Details
Dataset Description
Curated by: jaxon3062
Language(s) (NLP): en
License: Apache 2.0 Licence
Dataset Sources [optional]
Repository: google/IFEval
Paper [optional]: Instruction-Following Evaluation for Large Language Models
Uses
Direct Use
This can… See the full description on the dataset page: https://huggingface.co/datasets/jaxon3062/IFEval-gemma3-chat.jaxaht-benchmark-leaderboard
JaxAHT benchmark — leaderboard + submitted checkpoints
Backing store for the live JaxAHT benchmark API at https://lainwired-jaxaht-benchmark.hf.space.
Layout:
leaderboard/<env>_<version>.json — leaderboard entries per env+version
checkpoints/<entry_id>/checkpoint.safetensors — submitted ego checkpoints (one per submission)
checkpoints/<entry_id>/meta.json — per-submission metadata
Updated automatically by the API on submission.
HumTrans2
HumTrans Dataset
Dataset Name: HumTrans
Dataset Type: Humming audio in .wav format and corresponding label MIDI file
Primary Use: Humming melody transcription and as a foundation for downstream tasks such as humming melody based music generation
Summary: 500 musical compositions of different genres and languages, 1000 music segments in total; sampled at a frequency of 44,100 Hz; approximately 56.22 hours of audio; 14,614 files in total.
File Description: all_wav.zip includes… See the full description on the dataset page: https://huggingface.co/datasets/jax321/HumTrans2.canny_diffusiondb
Canny DiffusionDB
This dataset is the DiffusionDB dataset that is transformed using Canny transformation.
You can see samples below 👇
Sample:
Original Image:
Transformed Image:
Caption:
"a small wheat field beside a forest, studio lighting, golden ratio, details, masterpiece, fine art, intricate, decadent, ornate, highly detailed, digital painting, octane render, ray tracing reflections, 8 k, featured, by claude monet and vincent van gogh "Below you can find a small script used… See the full description on the dataset page: https://huggingface.co/datasets/jax-diffusers-event/canny_diffusiondb.wams2026-am-process-jaxontologies_phenotype_to_genes_JAXwhisper-jax-test-files
Dataset Card for "whisper-jax-test-files"
More Information needed
whisper-jax-examples
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Samhita/whisper-jax-examples.jax-tasks-v1
jax-tasks-v1
Task dataset for a JAX RL / eval environment, in the shape used by the
Prime Intellect Environments Hub.
38 JAX tasks across 5 categories. Each task gives the model one or more input arrays and an
instruction; the answer is the array left in result, graded with numpy.allclose against a
reference. Grading is deterministic — no LLM judge, no external API, CPU only.
Category
Tasks
Covers
array_ops
14
reshape, transpose, axis reductions, clip, sort/argsort… See the full description on the dataset page: https://huggingface.co/datasets/eltociear/jax-tasks-v1.github-issues
Dataset Card for "github-issues"
More Information needed
HumanPhenotype_JAX
