CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ggranberry /lemmanaid-afp-reruns Lemmanaid AFP-pool Reproducibility Reruns Reproducibility study for claude-opus-4-5 on the yalhessi/lemexp-commerical-llm-experiment benchmark, using an AFP demo pool (honest eval — no train/test theory leakage). Companion to ggranberry/lemmanaid-commercial-results, which holds the earlier shot-count + retrieval sweeps under test-LOO. Configs Two configs, one per benchmark domain: Config Source HF config Test rows octonions template_octonions_2026… See the full description on the dataset page: https://huggingface.co/datasets/ggranberry/lemmanaid-afp-reruns.texttext-generation1K<n<10K0 likes618 downloads4mo agoHugging Face02Lemma456 /renderer_test_dataset0 likes312 downloads7mo agoHugging Face03Tencent-IMO /IMO-Lemmas Tencent-IMO: Towards Solving More Challenging IMO Problems via Decoupled Reasoning and Proving This dataset contains strategic subgoals (lemmas) and their formal proofs for a challenging set of post-2000 International Mathematical Olympiad (IMO) problems. All statements and proofs are formalized in the Lean 4 theorem proving language. The data was generated using the Decoupled Reasoning and Proving framework, introduced in our paper: Towards Solving More Challenging IMO Problems… See the full description on the dataset page: https://huggingface.co/datasets/Tencent-IMO/IMO-Lemmas.textn<1K3 likes125 downloads1y agoHugging Face04BlackdromeAILabs /lemma-long-horizon-rewrite-benchmark LEMMA Long-Horizon Symbolic Rewriting Benchmark Rewrite one expression into another, one verified step at a time — for up to 128 steps. Every problem gives you a start expression, an exact target, and a reference derivation in which every single step was accepted by a symbolic verifier. The hard part is not any individual rewrite. It is picking the right rewrite at each of up to 128 consecutive states, where a plausible-looking legal move can quietly take you away from the… See the full description on the dataset page: https://huggingface.co/datasets/BlackdromeAILabs/lemma-long-horizon-rewrite-benchmark.tabularothern<1K0 likes94 downloads14d agoHugging Face05AI-Math-TCS /tcs_find_lemmatext10K<n<100K0 likes86 downloads4mo agoHugging Face06Lemma-RCA-NEC /Cloud_Computing_Preprocessed Data Description: Preprocessed system metrics and log data from Cloud Computing Platform. Constructed the metric time series (as npy format) from the original metrics data (Json format). Extracted the log messages from the original log data (Json format). Parsed the log messages into log event templates. Note: 20240207 data does not contain EKS log data; it solely comprises CloudTrail log data in CSV format. Consequently, this dataset does not require preprocessing with a log… See the full description on the dataset page: https://huggingface.co/datasets/Lemma-RCA-NEC/Cloud_Computing_Preprocessed.time-series-forecasting100M<n<1B7 likes85 downloads9mo agoHugging Face07whu9 /word_net_synset_lemma Dataset Card for "word_net_synset_lemma" More Information needed text100K<n<1M1 likes76 downloads3y agoHugging Face08Lemma-RCA-NEC /Cloud_Computing_Original Data Description: Both system metrics (json format) and log data (json format) were collected from Cloud Computing Platform with hundreds of system entities invovled. Six different types of real faults (including cryptojacking, silent pod degradation fault, malware attack, mistakes made by GitOps, configuration change failure, and bug infection) were simulated. Citation: Lecheng Zheng, Zhengzhang Chen, Dongjie Wang, Chengyuan Deng, Reon Matsuoka, and Haifeng… See the full description on the dataset page: https://huggingface.co/datasets/Lemma-RCA-NEC/Cloud_Computing_Original.time-series-forecasting100M<n<1B4 likes64 downloads1y agoHugging Face09lemmaakademi /kavram-yanilgisi-envanteri Kavram Yanılgısı Envanteri — Türkiye ortaokul matematiği (6-8. sınıf) 505 kayıt · 31 kaynak · 24 sütun · Türkçe Bu veri seti, Türkiye'de tam metnine erişilen 31 akademik çalışmanın (11 yüksek lisans tezi + 20 dergi makalesi) okunarak kodlanmış kavram yanılgısı kayıtlarından oluşur. Her satır bir yanılgıyı tanımlar; kaynağın künyesini ve sayfa numarasını taşır, bu sayede her kayıt tek tek doğrulanabilir (505 kaydın 505'inde sayfa referansı vardır). Kapsam 6, 7 ve… See the full description on the dataset page: https://huggingface.co/datasets/lemmaakademi/kavram-yanilgisi-envanteri.tabularn<1K0 likes58 downloads14d agoHugging Face10Lemma-RCA-NEC /Product_Review_Original Data Description: Both metrics and log data were collected from Product Review Microservice Platform with hundreds of system entities invovled. Four different types of real faults (including DDoS attack, external storage failure, node resource contention stress test, and noisy neighbor issue) were simulated. Citation: Lecheng Zheng, Zhengzhang Chen, Dongjie Wang, Chengyuan Deng, Reon Matsuoka, and Haifeng Chen: LEMMA-RCA: A Large Multi-modal Multi-domain Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Lemma-RCA-NEC/Product_Review_Original.time-series-forecasting100M<n<1B2 likes57 downloads1y agoHugging Face11Lemma-RCA-NEC /Product_Review_Preprocessed Data Description: Preprocessed metrics and log data from Product Review Microservice Platform. Constructed the metric time series (as npy format) from the original metrics data (Json format). Extracted the log messages from the original log data (Json format). Parsed the log messages into log event templates. Timezone Note (Important) PPTX scenario slides use Japan Standard Time (JST) for measurement periods and failure timestamps, because the real testbed is… See the full description on the dataset page: https://huggingface.co/datasets/Lemma-RCA-NEC/Product_Review_Preprocessed.time-series-forecasting100M<n<1B4 likes56 downloads9mo agoHugging Face12airudit /UD_v2_17_POS_LEMMA Dataset Information This dataset is designed for the Part-of-Speech (POS) tagging and Lemmatization tasks and is used to train the Airudit multitask model. Dataset Description The dataset combines Universal Dependencies v2.17 French corpora available atcommul/universal_dependencies. Included corpora: commul/universal_dependencies/fr_gsd commul/universal_dependencies/fr_sequoia commul/universal_dependencies/fr_partut commul/universal_dependencies/fr_parisstories… See the full description on the dataset page: https://huggingface.co/datasets/airudit/UD_v2_17_POS_LEMMA.text10K<n<100K0 likes46 downloads4mo agoHugging Face13ryderwishart /hellenistic-greek-lemmas Dataset Card for "hellenistic-greek-lemmas" More Information needed text1M<n<10M1 likes42 downloads4y agoHugging Face14panzs19 /LEMMA Dataset Card for Dataset Name The LEMMA is collected from MATH and GSM8K. The training set of MATH and GSM8K is used to generate error-corrective reasoning trajectories. For each question in these datasets, the student model (LLaMA3-8B) generates self-generated errors, and the teacher model (GPT-4o) deliberately introduces errors based on the error type distribution of the student model. Then, both "Fix & Continue" and "Fresh & Restart" correction strategies are applied to these… See the full description on the dataset page: https://huggingface.co/datasets/panzs19/LEMMA.text10K<n<100K0 likes42 downloads1y agoHugging Face15xwjzds /bbc-news_lemma_traintext1K<n<10K0 likes35 downloads3y agoHugging Face16chronbmm /sanskrit-lemmatext100K<n<1M0 likes34 downloads2y agoHugging Face17chronbmm /sanskrit-lemma-paragraphtext100K<n<1M0 likes31 downloads2y agoHugging Face18xwjzds /ag_news_lemma_train Dataset Card for Dataset Name Dataset Summary This is lemmatized version of Ag News Data. Languages English Citation Information @inproceedings{xu-etal-2023-vontss, title = "v{ONTSS}: v{MF} based semi-supervised neural topic modeling with optimal transport", author = "Xu, Weijie and Jiang, Xiaoyu and Sengamedu Hanumantha Rao, Srinivasan and Iannacci, Francis and Zhao, Jinjin", booktitle = "Findings of the… See the full description on the dataset page: https://huggingface.co/datasets/xwjzds/ag_news_lemma_train.text100K<n<1M0 likes27 downloads3y agoHugging Face19f1rdavs /tajik_lemmastext10K<n<100K0 likes27 downloads1y agoHugging Face20Titung /nepali-lemma-gold Nepali Word-Lemma Gold Data 5,000+ manually annotated word-lemma pairs for Nepali. Source: https://github.com/dpakpdl/NepaliLemmatizer Usage from datasets import load_dataset ds = load_dataset("Titung/nepali-lemma-gold") texttoken-classification10K<n<100K0 likes27 downloads6mo agoHugging Face21arbitropy /bangla_lemmatext10K<n<100K0 likes25 downloads2y agoHugging Face22chronbmm /sanskrit-unsandhi-lemma-morphosyntax-tagging-paragraphtext100K<n<1M0 likes25 downloads2y agoHugging Face23chronbmm /sanskrit-lemma-without-rectext100K<n<1M0 likes24 downloads2y agoHugging Face24anchpop /lemma-pos-depThis contains tokenizations, lemmatizations, part-of-speech tags, and dependency information for various sentences in seven languages. In general, they should be high quality and consistent. The dependency information is probably a little worse than the rest. (I haven't put as much effort into it.) There are some uniquities to the tokenization format I chose. Multiword proper nouns are considered one token, so for example "Grand Canyon" will be a single token. Hyphens are also their own tokens… See the full description on the dataset page: https://huggingface.co/datasets/anchpop/lemma-pos-dep.text100K<n<1M1 likes23 downloads2mo agoHugging Face25yalhessi /lemmanaid-rawtext100K<n<1M0 likes22 downloads6mo agoHugging Face26fia24 /filtered_lemma41kV0.0.04 Dataset Card for "filtered_lemma41kV0.0.04" More Information needed text10K<n<100K0 likes21 downloads3y agoHugging Face27ryderwishart /semantic-domains-greek-lemmatized Dataset Card for semantic-domains-greek-lemmatized Dataset Summary Semantic domains aligned to tokens, broken down by sentences. Tokens have been lemmatized according to data in Clear-Bible/macula-greek. Domains are based on Louw and Nida's semantic domains for the Greek New Testament. Languages Greek, Hellenistic Greek, Koine Greek, Greek of the New Testament Dataset Structure Data Instances DatasetDict({ train: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/ryderwishart/semantic-domains-greek-lemmatized.texttoken-classification1K<n<10K0 likes20 downloads4y agoHugging Face28xwjzds /20_newsgroups_lemma_testtext1K<n<10K0 likes20 downloads3y agoHugging Face29kanishka /babylm1_no_preemption_laugh-lemmatizedtext10M<n<100M0 likes20 downloads5mo agoHugging Face30xwjzds /bbc-news_lemma_testtext1K<n<10K0 likes19 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.