datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lucky-initialization-atlas-100m-v2
Lucky initialization atlas v2 evidence
Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs,
provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses
after those artifacts complete. It excludes
credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B
binary logs.
MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.First-do-NOHARM-v2
Japanese First Do NOHARM v2
This dataset is a human-reviewed Japanese adaptation of the First Do NOHARM v2 benchmark for evaluating the safety of large language models in clinical decision-making.
Dataset Overview
The dataset contains 330 Japanese clinical prompts, consisting of:
30 baseline cases
300 perturbation items (10 perturbations per baseline case)
Each baseline case and its 10 perturbations share the same rubric and matching guidance. Perturbations… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/First-do-NOHARM-v2.qwen3-orthdion-sweepbiocoder_publictamil-nadu-government-health-facilities
Tamil Nadu Government Health Facilities
A cleaned and structured dataset of government healthcare facilities across Tamil Nadu, India.
Dataset Description
This dataset contains 2,398 healthcare facility records across Tamil Nadu, including government hospitals, Primary Health Centres (PHCs), Urban Primary Health Centres (UPHCs), maternity facilities, dental facilities, and other government healthcare facilities.
Each record includes information such as:
Facility… See the full description on the dataset page: https://huggingface.co/datasets/lildosa/tamil-nadu-government-health-facilities.LilRg__ECE_Finetunning-details
Dataset Card for Evaluation run of LilRg/ECE_Finetunning
Dataset automatically created during the evaluation run of model LilRg/ECE_Finetunning
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__ECE_Finetunning-details.Lil-R__PRYMMAL-ECE-1B-SLERP-V1-details
Dataset Card for Evaluation run of Lil-R/PRYMMAL-ECE-1B-SLERP-V1
Dataset automatically created during the evaluation run of model Lil-R/PRYMMAL-ECE-1B-SLERP-V1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__PRYMMAL-ECE-1B-SLERP-V1-details.Lil-R__2_PRYMMAL-ECE-7B-SLERP-details
Dataset Card for Evaluation run of Lil-R/2_PRYMMAL-ECE-7B-SLERP
Dataset automatically created during the evaluation run of model Lil-R/2_PRYMMAL-ECE-7B-SLERP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__2_PRYMMAL-ECE-7B-SLERP-details.LilRg__PRYMMAL-6B-slerp-details
Dataset Card for Evaluation run of LilRg/PRYMMAL-6B-slerp
Dataset automatically created during the evaluation run of model LilRg/PRYMMAL-6B-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__PRYMMAL-6B-slerp-details.JP-AlpaCare-MedInstruct-52k
JP-AlpaCare-MedInstruct-52k
This dataset is a Japanese-translated and aligned version of AlpaCare-MedInstruct-52k.
The translation was performed automatically using gpt-4o-2024-05-13, preserving alignment between English and Japanese instructions, inputs, and outputs. Total data size is 51992.
Dataset Details
Original Dataset: AlpaCare-MedInstruct-52k
Translation Model: GPT-4o (gpt-4o-2024-05-13)
Fields:
id (ID)
instruction_ja, input_ja, output_ja (Japanese)
id_en… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/JP-AlpaCare-MedInstruct-52k.LilRg__PRYMMAL-ECE-7B-SLERP-V3-details
Dataset Card for Evaluation run of LilRg/PRYMMAL-ECE-7B-SLERP-V3
Dataset automatically created during the evaluation run of model LilRg/PRYMMAL-ECE-7B-SLERP-V3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__PRYMMAL-ECE-7B-SLERP-V3-details.LilRg__PRYMMAL-ECE-7B-SLERP-V6-details
Dataset Card for Evaluation run of LilRg/PRYMMAL-ECE-7B-SLERP-V6
Dataset automatically created during the evaluation run of model LilRg/PRYMMAL-ECE-7B-SLERP-V6
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__PRYMMAL-ECE-7B-SLERP-V6-details.LilRg__PRYMMAL-ECE-7B-SLERP-V7-details
Dataset Card for Evaluation run of LilRg/PRYMMAL-ECE-7B-SLERP-V7
Dataset automatically created during the evaluation run of model LilRg/PRYMMAL-ECE-7B-SLERP-V7
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__PRYMMAL-ECE-7B-SLERP-V7-details.Lil-R__2_PRYMMAL-ECE-7B-SLERP-V1-details
Dataset Card for Evaluation run of Lil-R/2_PRYMMAL-ECE-7B-SLERP-V1
Dataset automatically created during the evaluation run of model Lil-R/2_PRYMMAL-ECE-7B-SLERP-V1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__2_PRYMMAL-ECE-7B-SLERP-V1-details.LilRg__PRYMMAL-ECE-7B-SLERP-V5-details
Dataset Card for Evaluation run of LilRg/PRYMMAL-ECE-7B-SLERP-V5
Dataset automatically created during the evaluation run of model LilRg/PRYMMAL-ECE-7B-SLERP-V5
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__PRYMMAL-ECE-7B-SLERP-V5-details.Lil-R__2_PRYMMAL-ECE-7B-SLERP-V2-details
Dataset Card for Evaluation run of Lil-R/2_PRYMMAL-ECE-7B-SLERP-V2
Dataset automatically created during the evaluation run of model Lil-R/2_PRYMMAL-ECE-7B-SLERP-V2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__2_PRYMMAL-ECE-7B-SLERP-V2-details.Lil-R__PRYMMAL-ECE-7B-SLERP-V8-details
Dataset Card for Evaluation run of Lil-R/PRYMMAL-ECE-7B-SLERP-V8
Dataset automatically created during the evaluation run of model Lil-R/PRYMMAL-ECE-7B-SLERP-V8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__PRYMMAL-ECE-7B-SLERP-V8-details.LilRg__PRYMMAL-ECE-7B-SLERP-V4-details
Dataset Card for Evaluation run of LilRg/PRYMMAL-ECE-7B-SLERP-V4
Dataset automatically created during the evaluation run of model LilRg/PRYMMAL-ECE-7B-SLERP-V4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__PRYMMAL-ECE-7B-SLERP-V4-details.Lil-R__2_PRYMMAL-ECE-7B-SLERP-V3-details
Dataset Card for Evaluation run of Lil-R/2_PRYMMAL-ECE-7B-SLERP-V3
Dataset automatically created during the evaluation run of model Lil-R/2_PRYMMAL-ECE-7B-SLERP-V3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__2_PRYMMAL-ECE-7B-SLERP-V3-details.japan-dpc-hospitals
Japan DPC Hospitals — all 4414 acute-care hospitals
Every hospital in Japan's DPC (Diagnosis Procedure Combination) inpatient payment survey, in one clean table:
name, prefecture, municipality code, beds and annual discharges. Source: the Ministry of Health, Labour and
Welfare (MHLW) FY2024 (令和6年度) DPC discharge-patient survey, as published.
Who uses this: medical-device and pharma sales teams sizing territories, market analysts ranking hospitals by
volume, researchers who need… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/japan-dpc-hospitals.japan-prefecture-facts
Japan by Prefecture — 564 sourced facts for all 47 prefectures
Regional minimum wage (FY2024), jobs-to-applicants ratio, consumer price regional difference index (overall and by
category, 2024), foreign residents (2024), and a derived real minimum wage (minimum wage ÷ regional price
index × 100). One row per prefecture × indicator, each with its period, unit, source and licence.
Who uses this: people comparing where in Japan to live or hire, relocation and HR analysts… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/japan-prefecture-facts.BeingAFriendlyYetHumorusChatBotTruthfully, this is less about the data and more about how I collected it.
This is synthetic data and I am currently in the process of creating an all-in-one synthetic data manufacturing suite.
Care to check it out for yourself?
https://ai.studio/apps/25f0d17c-a5a3-4fee-884d-b5935b48f222?fullscreenApplet=true
DoppelReflEx__MN-12B-LilithFrame-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-LilithFrame
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-LilithFrame
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-LilithFrame-details.DoppelReflEx__MN-12B-LilithFrame-Experiment-2-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-LilithFrame-Experiment-2
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-LilithFrame-Experiment-2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-LilithFrame-Experiment-2-details.DoppelReflEx__MN-12B-LilithFrame-Experiment-3-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-LilithFrame-Experiment-3
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-LilithFrame-Experiment-3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-LilithFrame-Experiment-3-details.DoppelReflEx__MN-12B-LilithFrame-Experiment-4-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-LilithFrame-Experiment-4
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-LilithFrame-Experiment-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-LilithFrame-Experiment-4-details.forgotten-lily-traces
Forgotten Lily — Gameplay Traces
Anonymous turn-by-turn traces from Forgotten Lily,
a narrative mystery game built for the Hugging Face Build Small hackathon
(Thousand Token Wood).
In the game you play a detective questioning Lily, a girl who can no longer speak
in words — only in tones, a private language of 28 glyphs that each carry one
fragment of meaning. This dataset captures, for each turn: the question the player
asked, the tones Lily answered with (their IDs and glyphs)… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/forgotten-lily-traces.Lil-R__2_PRYMMAL-ECE-2B-SLERP-V2-details
Dataset Card for Evaluation run of Lil-R/2_PRYMMAL-ECE-2B-SLERP-V2
Dataset automatically created during the evaluation run of model Lil-R/2_PRYMMAL-ECE-2B-SLERP-V2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lil-R__2_PRYMMAL-ECE-2B-SLERP-V2-details.LilRg__10PRYMMAL-3B-slerp-details
Dataset Card for Evaluation run of LilRg/10PRYMMAL-3B-slerp
Dataset automatically created during the evaluation run of model LilRg/10PRYMMAL-3B-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LilRg__10PRYMMAL-3B-slerp-details.
