datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bio-overrefusal-v0.1
Bio Over-Refusal Dataset v0.1.0
Dataset Summary
The Bio Over-Refusal Dataset is a domain-expert-authored and tier-annotated benchmark of 201 legitimate biology research queries stratified by sensitivity tier. It is designed to measure the false-positive refusal rate (FPR) of large language models — specifically, the rate at which models refuse or hedge on questions that credentialed biology researchers would consider appropriate to answer.
The dataset does not… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/bio-overrefusal-v0.1.lm-eval-results-jan-hq-stealth-v2-private
Dataset Card for Evaluation run of jan-hq/stealth-v2
Dataset automatically created during the evaluation run of model jan-hq/stealth-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-jan-hq-stealth-v2-private.sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.ambiguity-casebook
Dual-Use Ambiguity Casebook
A 35-row, single-annotator research corpus for studying context-sensitive
adjudication in AI-mediated biology. Records follow one exact 21-field schema
and cover six descriptive categories.
This is a dataset release, not a model leaderboard, compliance tool, or
laboratory guide. Raw model responses, model-comparison results, live-provider
evaluation code, and adversarial failure-mode material are intentionally
excluded.
Dataset structure… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/ambiguity-casebook.LabCraft-Eval
LabCraft-Eval
LabCraft-Eval is an Inspect AI evaluation environment for measuring how well AI
agents execute benign molecular-microbiology protocols inside a seeded
laboratory simulator with task-dependent stochasticity. It pairs task prompts
and tool-accessible lab operations with deterministic, multi-axis trajectory
scoring.
This Hugging Face dataset export is generated from the GitHub repository:
https://github.com/jang1563/LabCraft-Eval.git
Release
Release… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/LabCraft-Eval.jan_2026_financial_advice
Crypto Trading Signals, Risk Screening & OHLCV Candles (Jan 2026)
This dataset is a worked example of a pump-and-dump
The headline is +52.2% over 22 trades. Almost all of it is two low-cap pumps.
52% ≈ PUMP (+22.58) + FARTCOIN (+38.33) — ~61 п.п. on three trades. Everything else is in the red.
Sharpe 0.30, a real mark-to-market drawdown of ~21%, Recovery Factor 2.75 — on 22 trades over
27 days this is not statistics, it is the description of one lucky episode.
On… See the full description on the dataset page: https://huggingface.co/datasets/tripolskypetr/jan_2026_financial_advice.kaist-ai__janus-rm-7b-details
Dataset Card for Evaluation run of kaist-ai/janus-rm-7b
Dataset automatically created during the evaluation run of model kaist-ai/janus-rm-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-rm-7b-details.protein-structure-trust-benchmark
Protein-Structure Trust-Routing Benchmark (Boltz-2)
Leakage-controlled benchmarks for confidence-calibrated trust routing over a protein-structure
predictor: given a specialist model's confidence (Boltz-2 ipTM / pLDDT) for a target, decide whether to
trust the prediction or pay to verify it — and score that decision against experimentally-measured
correctness. Evaluation substrate for the report "When does an LLM trust a specialist model? A cost-aware
trust-routing audit"… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/protein-structure-trust-benchmark.janli_synthetic_rationale
JaNLI synthetic rationale
JaNLI: 日本語の言語現象に基づく 敵対的推論データセットの回答の判断根拠を、実験的に、言語モデルによって付与したデータセットです。
Assessing the Generalization Capacity of Pre-trained Language Models through Japanese Adversarial Natural Language Inference
判断根拠の付与にはmicrosoft/Phi-3-medium-4k-instructを用いました。
特徴
1件のサンプルにつき、4件の回答の判断根拠の候補文を付与しています。
Greedy Search(do_sample=False)では判断根拠を述べない事例が多く確認されたため、生成パラメータを変動させて4件の判断根拠の候補文を出力させています。
どの判断根拠の候補文を採用すべきかは作成者もまだ回答を持っていません。
引用… See the full description on the dataset page: https://huggingface.co/datasets/ryota39/janli_synthetic_rationale.Long_Covid_word_frequency_TFIDF_21_Nov_22_Janwhole_text_TF_21_Nov_22_Janfarmer-price-data-2026-jan-aprPJMixers-Dev__LLaMa-3.2-Instruct-JankMixBread-v0.1-3B-details
Dataset Card for Evaluation run of PJMixers-Dev/LLaMa-3.2-Instruct-JankMixBread-v0.1-3B
Dataset automatically created during the evaluation run of model PJMixers-Dev/LLaMa-3.2-Instruct-JankMixBread-v0.1-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers-Dev__LLaMa-3.2-Instruct-JankMixBread-v0.1-3B-details.PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.2-SFT-3B-details
Dataset Card for Evaluation run of PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.2-SFT-3B
Dataset automatically created during the evaluation run of model PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.2-SFT-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.2-SFT-3B-details.PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.2-SFT-HailMary-v0.1-KTO-3B-details
Dataset Card for Evaluation run of PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.2-SFT-HailMary-v0.1-KTO-3B
Dataset automatically created during the evaluation run of model PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.2-SFT-HailMary-v0.1-KTO-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.2-SFT-HailMary-v0.1-KTO-3B-details.astra-final-datasetPJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.1-SFT-3B-details
Dataset Card for Evaluation run of PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.1-SFT-3B
Dataset automatically created during the evaluation run of model PJMixers-Dev/LLaMa-3.2-Instruct-JankMix-v0.1-SFT-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PJMixers-Dev__LLaMa-3.2-Instruct-JankMix-v0.1-SFT-3B-details.kaist-ai__janus-dpo-7b-details
Dataset Card for Evaluation run of kaist-ai/janus-dpo-7b
Dataset automatically created during the evaluation run of model kaist-ai/janus-dpo-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-dpo-7b-details.kaist-ai__janus-7b-details
Dataset Card for Evaluation run of kaist-ai/janus-7b
Dataset automatically created during the evaluation run of model kaist-ai/janus-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-7b-details.file_test_1
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Jandodev/file_test_1.
