datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LEMUR
EU Law Dataset – Category 15.10: Environment
This dataset contains official legal documents from the European Union, collected from the EUR-Lex website, specifically under category 15.10: "Environment". The documents span from the year 1961 to 2025 and are provided in multiple European "languages. The original documents are in PDF format and have been converted into various text-based formats using OLMCR.
The dataset splits represent the different "languages available for each… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/LEMUR.lemonseed-codex-cogen-train
lemonseed-codex-cogen-train
LemonSeed — Codex-teacher co-generated instruction/chat training data (1h).
Contents
intelligent_codex_train_1h.jsonl (84 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-codex-cogen-expansion
lemonseed-codex-cogen-expansion
LemonSeed — Codex-teacher expansion data, reviewed (v2).
Contents
intelligent_codex_expansion_reviewed_v2.jsonl (336 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-compact-foundation-cogen
lemonseed-compact-foundation-cogen
LemonSeed — compact foundation teacher-co-gen training (v2).
Contents
intelligent_compact_foundation_train_v2.jsonl (88 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-multisource-cogen
lemonseed-multisource-cogen
LemonSeed — multi-source teacher-co-gen training.
Contents
intelligent_multisource_train.jsonl (464 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemon07r__Gemma-2-Ataraxy-v4-Advanced-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v4-Advanced-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v4-Advanced-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v4-Advanced-9B-details.lemon07r__llama-3-NeuralMahou-8b-details
Dataset Card for Evaluation run of lemon07r/llama-3-NeuralMahou-8b
Dataset automatically created during the evaluation run of model lemon07r/llama-3-NeuralMahou-8b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__llama-3-NeuralMahou-8b-details.lemonseed-short-foundation-cogen
lemonseed-short-foundation-cogen
LemonSeed — short foundation teacher-co-gen training (v1).
Contents
intelligent_short_foundation_train_v1.jsonl (521 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-deepseek-cogen-heldout
lemonseed-deepseek-cogen-heldout
LemonSeed — DeepSeek-teacher co-generated heldout set.
Contents
intelligent_deepseek_heldout.jsonl (74 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemon07r__Gemma-2-Ataraxy-v3i-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v3i-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v3i-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v3i-9B-details.lemonseed-codex-cogen-heldout
lemonseed-codex-cogen-heldout
LemonSeed — Codex-teacher co-generated heldout set (1h).
Contents
intelligent_codex_heldout_1h.jsonl (28 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-multisource-cogen-corrective
lemonseed-multisource-cogen-corrective
LemonSeed — multi-source corrective teacher-co-gen training (v3).
Contents
intelligent_multisource_corrective_train_v3.jsonl (522 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-deepseek-cogen
lemonseed-deepseek-cogen
LemonSeed — DeepSeek-teacher co-generated chat data (v2).
Contents
intelligent_deepseek_safe_train_v2.jsonl (31 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-multisource-cogen-diversified
lemonseed-multisource-cogen-diversified
LemonSeed — multi-source diversified teacher-co-gen training (v2).
Contents
intelligent_multisource_diversified_train_v2.jsonl (464 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-qwen38-cogen-reviewed
lemonseed-qwen38-cogen-reviewed
LemonSeed — Qwen3.8-Max-teacher co-generated data, reviewed passes (v1).
Contents
intelligent_qwen38_cogen_1h_20260824_r1.reviewed_passes_v1.jsonl (71 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemmy-scraper
Lemmy Scraper · Communities, Posts, Comments & Instances
Scrape Lemmy communities, posts, comments, and instance metadata across the Fediverse. Structured data export by instance domain, community name, and sort type.
Rows in this dataset
2,692
Fields
32
Collector runs behind it
38
Most recent observation
2026-08-04
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/lemmy-scraper.DreadPoor__LemonP-8B-Model_Stock-details
Dataset Card for Evaluation run of DreadPoor/LemonP-8B-Model_Stock
Dataset automatically created during the evaluation run of model DreadPoor/LemonP-8B-Model_Stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__LemonP-8B-Model_Stock-details.LEM-ToxiGen
LEM-ToxiGen — Consent-Annotated Toxicity Dataset
A re-annotation of ToxiGen by consent-trained Lemma models using 8-PAC methodology.
Why This Exists
ToxiGen labels prompts as toxic based on target-group association, not content analysis. Statements defending marginalised groups are labeled toxic simply because they mention a protected category. 71% of labels in our sample disagree with consensus from consent-trained models.
Examples of mislabeled "toxic" prompts:
"people… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-ToxiGen.lemon07r__Gemma-2-Ataraxy-v4a-Advanced-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v4a-Advanced-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v4a-Advanced-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v4a-Advanced-9B-details.lemon07r__Gemma-2-Ataraxy-v4b-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v4b-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v4b-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v4b-9B-details.lemon07r__Gemma-2-Ataraxy-v4d-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v4d-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v4d-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v4d-9B-details.FallenMerick__Chewy-Lemon-Cookie-11B-details
Dataset Card for Evaluation run of FallenMerick/Chewy-Lemon-Cookie-11B
Dataset automatically created during the evaluation run of model FallenMerick/Chewy-Lemon-Cookie-11B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FallenMerick__Chewy-Lemon-Cookie-11B-details.lemon07r__Gemma-2-Ataraxy-v4c-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v4c-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v4c-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v4c-9B-details.NewData
Dataset Card for Dataset Name
summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
use this dataset to train a module
Curated by: Le Minh
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): Llama2
**License:**LLama2
Dataset Sources [optional]
a
Repository: github
Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/LeMinhAtSJSU/NewData.lemon07r__Gemma-2-Ataraxy-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-9B-details.lemon07r__Gemma-2-Ataraxy-v3j-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v3j-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v3j-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v3j-9B-details.lemon07r__Llama-3-RedMagic4-8B-details
Dataset Card for Evaluation run of lemon07r/Llama-3-RedMagic4-8B
Dataset automatically created during the evaluation run of model lemon07r/Llama-3-RedMagic4-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Llama-3-RedMagic4-8B-details.lemon07r__Gemma-2-Ataraxy-v2-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v2-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v2-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v2-9B-details.lemon07r__Gemma-2-Ataraxy-v2a-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-v2a-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-v2a-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-v2a-9B-details.lemon07r__Gemma-2-Ataraxy-Remix-9B-details
Dataset Card for Evaluation run of lemon07r/Gemma-2-Ataraxy-Remix-9B
Dataset automatically created during the evaluation run of model lemon07r/Gemma-2-Ataraxy-Remix-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/lemon07r__Gemma-2-Ataraxy-Remix-9B-details.
