datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LEM-benchmarks
LEM-benchmarks
Canonical 8-PAC benchmark results for the Lemma model family.
This dataset is an aggregated store of per-round evaluation data produced by
lthn/LEM-Eval. Every row
represents one model's answer to one question in one round of a paired A/B
run against its unmodified base, and the dataset grows monotonically as more
workers contribute — different machines, different sampling states, different
hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.LEMONADE
🍋 EPFL-Smart-Kitchen: Lemonade benchmark
Paper | GitHub
📚 Introduction
we introduce Lemonade: Language models Evaluation of MOtion aNd Action-Driven Enquiries.
Lemonade consists of 36,521 closed-ended QA pairs linked to egocentric video clips, categorized in three groups and six subcategories. 18,857 QAs focus on behavior understanding, leveraging the rich ground truth behavior annotations of the EPFL-Smart Kitchen to interrogate models about perceived actions… See the full description on the dataset page: https://huggingface.co/datasets/amathislab/LEMONADE.smol-koreantalkSmolLM2의 인스트럭션 훈련 데이터 HuggingFaceTB/smol-smoltalk를 한국어로 번역했어요.
wildchatnew
Dataset Card for WildChat
Dataset Description
Paper: https://arxiv.org/abs/2405.01470
Interactive Search Tool: https://wildvisualizer.com (paper)
License: ODC-BY
Language(s) (NLP): multi-lingual
Point of Contact: Yuntian Deng
Dataset Summary
WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat… See the full description on the dataset page: https://huggingface.co/datasets/Lemonnn123/wildchatnew.lemonseed-rl-tasks-v2
lemonseed-rl-tasks-v2
LemonSeed — GRPO RL tasks v2 (vocab MC / antonym MC / cloze MC / grammar / dialogue / arithmetic).
Contents
rl_tasks_v2.jsonl (12492 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-cumulative-r2
lemonseed-cumulative-r2
LemonSeed — cumulative round 2 (multi-skill with arithmetic replay).
Contents
cum_r2.jsonl (20000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-mixed-skills
lemonseed-mixed-skills
LemonSeed — mixed Go + Sudoku + addition + prose stream.
Contents
mixed.jsonl (11428 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-games-r2
lemonseed-games-r2
LemonSeed — games round 2 (Go/Sudoku atari + constraint reasoning).
Contents
games_r2.jsonl (12000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-cumulative-r5
lemonseed-cumulative-r5
LemonSeed — cumulative round 5 (first fully successful multi-skill revival: prose + arithmetic + scaffold removal).
Contents
cum_r5.jsonl (20000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemone-docs-embedded
Lemone-embedded, pre-built embeddings dataset for French taxation.
This database presents the embeddings generated by the Lemone-embed-pro model and aims at a large-scale distribution of the model even for the GPU-poor.
This sentence transformers model, specifically designed for French taxation, has been fine-tuned on a dataset comprising 43 million tokens, integrating a blend of semi-synthetic and fully synthetic data generated by GPT-4 Turbo and Llama 3.1 70B, which have… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/lemone-docs-embedded.lemonseed-sudoku
lemonseed-sudoku
LemonSeed — 4x4 Latin-square / Sudoku completion with stepwise solving.
Contents
sudoku.jsonl (5000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-arithmetic
lemonseed-arithmetic
LemonSeed — multi-digit arithmetic scratchpad (decomposition method).
Contents
arithmetic.jsonl (16000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-rl-tasks-cogen-snapshot
lemonseed-rl-tasks-cogen-snapshot
LemonSeed — RL co-gen task snapshot.
Contents
rl_tasks_cogen_snapshot.jsonl (1809 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-sudoku-mixed
lemonseed-sudoku-mixed
LemonSeed — Sudoku reasoning interleaved with arithmetic replay.
Contents
sudoku_mixed.jsonl (5714 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-rl-tasks-cogen
lemonseed-rl-tasks-cogen
LemonSeed — RL co-generated chat tasks (teacher-authored).
Contents
rl_tasks_cogen.jsonl (2500 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-rl-chat-tasks
lemonseed-rl-chat-tasks
LemonSeed — chat-alignment RL tasks (prompt/gold single-turn).
Contents
rl_chat_tasks.jsonl (9852 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-cot-math-clean
lemonseed-cot-math-clean
LemonSeed — cleaned chain-of-thought math word problems.
Contents
cot_math_clean.jsonl (3530 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-vocabulary
lemonseed-vocabulary
LemonSeed — WordNet vocabulary Q&A with chain-of-thought definitions.
Contents
vocab.jsonl (4000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-teacher-sft
lemonseed-teacher-sft
LemonSeed — teacher-model SFT traces (multiple-choice vocabulary).
Contents
teacher_sft.jsonl (3333 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-addition-carry-method
lemonseed-addition-carry-method
LemonSeed — carry-method addition scratchpad (LSB-first, single-digit facts). Teaches digit-level addition with explicit written steps.
Contents
addition.jsonl (10000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-word-problems
lemonseed-word-problems
LemonSeed — 13-category arithmetic word problems with plan + scratchpad.
Contents
wordproblems.jsonl (2500 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-go-mixed
lemonseed-go-mixed
LemonSeed — Go reasoning interleaved with arithmetic replay.
Contents
go_mixed.jsonl (5714 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
algorithmic-reasoning-seed
Dataset Card for Algorithmic Reasoning (seed)
Note: This dataset is WIP and most question's answer section is empty or incomplete! See also "Other Known Limitations" section
Warning: If you somehow do use this dataset, remember to NOT do any eval after training on the questions in this dataset!
Dataset Summary
Dataset to help LLM learn how to reason about code, especially on algorithmic tasks, by seeing human demostration.
Supported Tasks and Leaderboards
[More… See the full description on the dataset page: https://huggingface.co/datasets/lemonteaa/algorithmic-reasoning-seed.lemonseed-cot-math
lemonseed-cot-math
LemonSeed — chain-of-thought math word problems (11 generators, computed solutions).
Contents
cot_math.jsonl (4000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-go
lemonseed-go
LemonSeed — 5x5 Go liberty/atari/capture reasoning with verification steps.
Contents
go.jsonl (5000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-cumulative-r1
lemonseed-cumulative-r1
LemonSeed — cumulative round 1 (Go 30% / Sudoku 30% / addition 28% / prose 12%).
Contents
cum_r1.jsonl (15000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
qwen_qa_pairs_cli_training.jsonl
Data sources
Multiple datasets from Hugging Face related to natural language to CLI pairs were gathered.
Human reviewed synthetic data from Claude Opus4.6 and ChatGPT4.5 were added.
A handful of grounding rows related to the organisation "Spicy Lemonade" were added (see details below)
Data processing
As part of the processing, data was converted to the Alpaca format with instruction (natural language), input (typically blank) and output (the CLI command) columns.
The… See the full description on the dataset page: https://huggingface.co/datasets/spicy-lemonade/qwen_qa_pairs_cli_training.jsonl.lemonseed-cumulative-r4
lemonseed-cumulative-r4
LemonSeed — cumulative round 4 (multi-skill, prose recovering).
Contents
cum_r4.jsonl (20000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-rl-tasks
lemonseed-rl-tasks
LemonSeed — GRPO RL tasks (vocab/cloze/antonym/dialogue-QA, machine-checkable).
Contents
rl_tasks.jsonl (9583 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
lemonseed-rl-tasks-cogen-pilot
lemonseed-rl-tasks-cogen-pilot
LemonSeed — RL co-gen pilot tasks.
Contents
rl_tasks_cogen_pilot.jsonl (61 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
