datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lemmanaid-afp-reruns
Lemmanaid AFP-pool Reproducibility Reruns
Reproducibility study for claude-opus-4-5 on the yalhessi/lemexp-commerical-llm-experiment benchmark, using an AFP demo pool (honest eval — no train/test theory leakage).
Companion to ggranberry/lemmanaid-commercial-results, which holds the earlier shot-count + retrieval sweeps under test-LOO.
Configs
Two configs, one per benchmark domain:
Config
Source HF config
Test rows
octonions
template_octonions_2026… See the full description on the dataset page: https://huggingface.co/datasets/ggranberry/lemmanaid-afp-reruns.renderer_test_datasetIMO-Lemmas
Tencent-IMO: Towards Solving More Challenging IMO Problems via Decoupled Reasoning and Proving
This dataset contains strategic subgoals (lemmas) and their formal proofs for a challenging set of post-2000 International Mathematical Olympiad (IMO) problems. All statements and proofs are formalized in the Lean 4 theorem proving language.
The data was generated using the Decoupled Reasoning and Proving framework, introduced in our paper: Towards Solving More Challenging IMO Problems… See the full description on the dataset page: https://huggingface.co/datasets/Tencent-IMO/IMO-Lemmas.lemma-long-horizon-rewrite-benchmark
LEMMA Long-Horizon Symbolic Rewriting Benchmark
Rewrite one expression into another, one verified step at a time — for up to 128 steps.
Every problem gives you a start expression, an exact target, and a reference derivation in which
every single step was accepted by a symbolic verifier. The hard part is not any individual
rewrite. It is picking the right rewrite at each of up to 128 consecutive states, where a
plausible-looking legal move can quietly take you away from the… See the full description on the dataset page: https://huggingface.co/datasets/BlackdromeAILabs/lemma-long-horizon-rewrite-benchmark.tcs_find_lemmaCloud_Computing_Preprocessed
Data Description:
Preprocessed system metrics and log data from Cloud Computing Platform.
Constructed the metric time series (as npy format) from the original metrics data (Json format).
Extracted the log messages from the original log data (Json format). Parsed the log messages into log event templates.
Note: 20240207 data does not contain EKS log data; it solely comprises CloudTrail log data in CSV format. Consequently, this dataset does not require preprocessing with a log… See the full description on the dataset page: https://huggingface.co/datasets/Lemma-RCA-NEC/Cloud_Computing_Preprocessed.word_net_synset_lemma
Dataset Card for "word_net_synset_lemma"
More Information needed
Cloud_Computing_Original
Data Description:
Both system metrics (json format) and log data (json format) were collected from Cloud Computing Platform with hundreds of system entities invovled. Six different types of real faults (including cryptojacking, silent pod degradation fault, malware attack, mistakes made by GitOps, configuration change failure, and bug infection) were simulated.
Citation:
Lecheng Zheng, Zhengzhang Chen, Dongjie Wang, Chengyuan Deng, Reon Matsuoka, and Haifeng… See the full description on the dataset page: https://huggingface.co/datasets/Lemma-RCA-NEC/Cloud_Computing_Original.kavram-yanilgisi-envanteri
Kavram Yanılgısı Envanteri — Türkiye ortaokul matematiği (6-8. sınıf)
505 kayıt · 31 kaynak · 24 sütun · Türkçe
Bu veri seti, Türkiye'de tam metnine erişilen 31 akademik çalışmanın (11 yüksek lisans
tezi + 20 dergi makalesi) okunarak kodlanmış kavram yanılgısı kayıtlarından oluşur.
Her satır bir yanılgıyı tanımlar; kaynağın künyesini ve sayfa numarasını taşır, bu
sayede her kayıt tek tek doğrulanabilir (505 kaydın 505'inde sayfa referansı vardır).
Kapsam
6, 7 ve… See the full description on the dataset page: https://huggingface.co/datasets/lemmaakademi/kavram-yanilgisi-envanteri.Product_Review_Original
Data Description:
Both metrics and log data were collected from Product Review Microservice Platform with hundreds of system entities invovled. Four different types of real faults (including DDoS attack, external storage failure, node resource contention stress test, and noisy neighbor issue) were simulated.
Citation:
Lecheng Zheng, Zhengzhang Chen, Dongjie Wang, Chengyuan Deng, Reon Matsuoka, and Haifeng Chen: LEMMA-RCA: A Large Multi-modal Multi-domain Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Lemma-RCA-NEC/Product_Review_Original.Product_Review_Preprocessed
Data Description:
Preprocessed metrics and log data from Product Review Microservice Platform.
Constructed the metric time series (as npy format) from the original metrics data (Json format).
Extracted the log messages from the original log data (Json format). Parsed the log messages into log event templates.
Timezone Note (Important)
PPTX scenario slides use Japan Standard Time (JST) for measurement periods and failure timestamps, because the real testbed is… See the full description on the dataset page: https://huggingface.co/datasets/Lemma-RCA-NEC/Product_Review_Preprocessed.UD_v2_17_POS_LEMMA
Dataset Information
This dataset is designed for the Part-of-Speech (POS) tagging and Lemmatization tasks and is used to train the Airudit multitask model.
Dataset Description
The dataset combines Universal Dependencies v2.17 French corpora available atcommul/universal_dependencies.
Included corpora:
commul/universal_dependencies/fr_gsd
commul/universal_dependencies/fr_sequoia
commul/universal_dependencies/fr_partut
commul/universal_dependencies/fr_parisstories… See the full description on the dataset page: https://huggingface.co/datasets/airudit/UD_v2_17_POS_LEMMA.hellenistic-greek-lemmas
Dataset Card for "hellenistic-greek-lemmas"
More Information needed
LEMMA
Dataset Card for Dataset Name
The LEMMA is collected from MATH and GSM8K. The training set of MATH and GSM8K is used to generate error-corrective reasoning trajectories. For each question in these datasets, the student model (LLaMA3-8B) generates self-generated errors, and the teacher model (GPT-4o) deliberately introduces errors based on the error type distribution of the student model. Then, both "Fix & Continue" and "Fresh & Restart" correction strategies are applied to these… See the full description on the dataset page: https://huggingface.co/datasets/panzs19/LEMMA.bbc-news_lemma_trainsanskrit-lemmasanskrit-lemma-paragraphag_news_lemma_train
Dataset Card for Dataset Name
Dataset Summary
This is lemmatized version of Ag News Data.
Languages
English
Citation Information
@inproceedings{xu-etal-2023-vontss,
title = "v{ONTSS}: v{MF} based semi-supervised neural topic modeling with optimal transport",
author = "Xu, Weijie and
Jiang, Xiaoyu and
Sengamedu Hanumantha Rao, Srinivasan and
Iannacci, Francis and
Zhao, Jinjin",
booktitle = "Findings of the… See the full description on the dataset page: https://huggingface.co/datasets/xwjzds/ag_news_lemma_train.tajik_lemmasnepali-lemma-gold
Nepali Word-Lemma Gold Data
5,000+ manually annotated word-lemma pairs for Nepali.
Source: https://github.com/dpakpdl/NepaliLemmatizer
Usage
from datasets import load_dataset
ds = load_dataset("Titung/nepali-lemma-gold")
bangla_lemmasanskrit-unsandhi-lemma-morphosyntax-tagging-paragraphsanskrit-lemma-without-reclemma-pos-depThis contains tokenizations, lemmatizations, part-of-speech tags, and dependency information for various sentences in seven languages.
In general, they should be high quality and consistent. The dependency information is probably a little worse than the rest. (I haven't put as much effort into it.)
There are some uniquities to the tokenization format I chose. Multiword proper nouns are considered one token, so for example "Grand Canyon" will be a single token. Hyphens are also their own tokens… See the full description on the dataset page: https://huggingface.co/datasets/anchpop/lemma-pos-dep.lemmanaid-rawfiltered_lemma41kV0.0.04
Dataset Card for "filtered_lemma41kV0.0.04"
More Information needed
semantic-domains-greek-lemmatized
Dataset Card for semantic-domains-greek-lemmatized
Dataset Summary
Semantic domains aligned to tokens, broken down by sentences. Tokens have been lemmatized according to data in Clear-Bible/macula-greek.
Domains are based on Louw and Nida's semantic domains for the Greek New Testament.
Languages
Greek, Hellenistic Greek, Koine Greek, Greek of the New Testament
Dataset Structure
Data Instances
DatasetDict({
train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/ryderwishart/semantic-domains-greek-lemmatized.20_newsgroups_lemma_testbabylm1_no_preemption_laugh-lemmatizedbbc-news_lemma_test
