datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
control-arena-persistent-state-glm-openweightdetails_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted-sync.automationbench600-open-weight-runs
AutomationBench 600 — raw run artifacts (open-weight models)
Complete raw exports from running the public 600-task
AutomationBench (Zapier, v1.0.5, --toolset api)
on open-weight models served locally with SGLang 0.5.12 on 8x H200.
Each .json.gz is the untouched --export-json output: run metadata, summary, and one record per
task including every message of the trajectory, the final environment state, and per-assertion
results.
File
Model
Pass rate
Partial credit… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/automationbench600-open-weight-runs.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-nonscale-weighted.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval-weighted.evalap-comparing-openweight-with-patronusaiglider-114
Comparing openweight with PatronusAI/glider (ID: 114)
Comparing openweight Albert-API with specific judge PatronusAI/glider
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt
Scores… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-patronusaiglider-114.eu-open-weight-models
EU-readiness of open-weight LLMs
Curated by LLM Radar — updated 2026-05-03 — 55 models.
A manually-reviewed dataset assessing open-weight Large Language Models (LLMs)
on their suitability for EU deployment and commercial use. Each model is
evaluated on licence, commercial use, training data, and
origin, with quality / speed / price metrics from
Artificial Analysis where available.
Primary use cases:
Selecting open-weight models for self-hosted EU deployment
Licensing and… See the full description on the dataset page: https://huggingface.co/datasets/llmradar/eu-open-weight-models.evalap-compare-open-weight-models-31th-83
Compare Open Weight Models 31th (ID: 83)
Comparing open weight models
Overview
This dataset contains 44 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: Groq/Llama-3-Groq-8B-Tool-Use, Qwen/Qwen3-VL-4B-Instruct, Qwen/Qwen3-VL-8B-Thinking, meta-llama/Llama-3.1-8B-Instruct, mistral-medium-2508, mistralai/Magistral-Small-2509, mistralai/Mistral-Small-3.2-24B-Instruct-2506… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-compare-open-weight-models-31th-83.evalap-comparing-openweight-with-openaigpt-oss-120b-113
Comparing openweight with openai/gpt-oss-120b (ID: 113)
Comparing openweight Albert-API with specific judge openai/gpt-oss-120b
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-openaigpt-oss-120b-113.open-weights-archiving-playbook
The Open-Weights Archiving Playbook
A practical, hard-won guide to keeping open-weight models offline — for independence, reproducibility, and peace of mind.
Part 1 of The Open-Weights Lifecycle — see roadmap.md for the full series arc (and article.md for the short story-version).
Why this exists. Open weights can disappear: providers deprecate repos, licenses change, geopolitics tightens export rules, or an API you depend on simply gets turned off. If a model matters to you… See the full description on the dataset page: https://huggingface.co/datasets/AXOlotlvaNschmozelot/open-weights-archiving-playbook.open-weight-judges-eval
Open-Weight Judges Evaluation Dataset
Evaluation outputs from a study on cross-task transferability of family-diverse open-weight judge panels for synthetic text evaluation.
Overview
This dataset contains ~6,180 judge evaluations across three synthetic-text tasks, produced by a three-member open-weight panel, proprietary baselines (GPT-o4-mini, Gemini 2.5 Flash), a specialized open evaluator (Atla Selene Mini), and four bias audits.
Tasks
Task
Items… See the full description on the dataset page: https://huggingface.co/datasets/klusai/open-weight-judges-eval.evalap-comparing-openweight-with-atlaaiselene-1-mini-llama-31-8b-115
Comparing openweight with AtlaAI/Selene-1-Mini-Llama-3.1-8B (ID: 115)
Comparing openweight Albert-API with specific judge AtlaAI/Selene-1-Mini-Llama-3.1-8B
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion… See the full description on the dataset page: https://huggingface.co/datasets/kaaloo/evalap-comparing-openweight-with-atlaaiselene-1-mini-llama-31-8b-115.retire-open-weights
Let it go — retiring an open-weights model with dignity
Part 5 of The Open-Weights Lifecycle — acquire → verify → run → maintain → retire.
Every guide about archiving tells you how to keep things. Almost none tell you how to stop keeping something — and that's the habit that separates a curated library from a hoard. A healthy archive isn't the one that never deletes; it's the one that deletes on purpose.
This is the part of the lifecycle nobody writes about. So here it is: how… See the full description on the dataset page: https://huggingface.co/datasets/AXOlotlvaNschmozelot/retire-open-weights.evalap-comparing-openweight-with-patronusaiglider-114
Comparing openweight with PatronusAI/glider (ID: 114)
Comparing openweight Albert-API with specific judge PatronusAI/glider
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt
Scores… See the full description on the dataset page: https://huggingface.co/datasets/kaaloo/evalap-comparing-openweight-with-patronusaiglider-114.evalap-comparing-openweight-with-atlaaiselene-1-mini-llama-31-8b-115
Comparing openweight with AtlaAI/Selene-1-Mini-Llama-3.1-8B (ID: 115)
Comparing openweight Albert-API with specific judge AtlaAI/Selene-1-Mini-Llama-3.1-8B
Overview
This dataset contains 24 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: openweight-large, openweight-medium, openweight-small
Metrics: energy_consumption, generation_time, gwp_consumption, judge_notator, judge_precision, nb_tokens_completion… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-openweight-with-atlaaiselene-1-mini-llama-31-8b-115.run-open-weights-locally
Run your first local open-weights model in 5 minutes
Part 3 of The Open-Weights Lifecycle — acquire → verify → run → maintain → retire.
You've acquired a model (Part 1) and worked out which quant fits your machine (Part 2). Now the fun part: actually running it — on your own hardware, offline, no API key, no monthly bill.
The surprise most people don't expect: you don't need a big GPU, or any GPU at all. A mainstream laptop runs small models comfortably on the CPU alone. Here's… See the full description on the dataset page: https://huggingface.co/datasets/AXOlotlvaNschmozelot/run-open-weights-locally.keep-open-weights-alive
Keep your local model library alive
Part 4 of The Open-Weights Lifecycle — acquire → verify → run → maintain → retire.
You downloaded the models (Part 1), sized them (Part 2), and ran them (Part 3). But an archive is not a "set it and forget it" thing. Files rot, disks die, and future-you forgets what half the folders even are. Keeping a library alive is a little bit of light maintenance, done on a rhythm — not a big project.
Here are the five habits that keep a local model… See the full description on the dataset page: https://huggingface.co/datasets/AXOlotlvaNschmozelot/keep-open-weights-alive.costos-openweight-modelsopen-weight-model-registry
Open-Weight Model Registry
Interactive registry and direct comparison for open-weight AI models
Open-Weight Model Registry is a data-driven Hugging Face Space for exploring and directly comparing downloadable AI models across:
license
access status
architecture
parameter count
context length
modalities
self-hosting
commercial-use status
modification and redistribution status
documented runtimes
verification date
The Space reads from:… See the full description on the dataset page: https://huggingface.co/datasets/open-weights/open-weight-model-registry.
