CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /lambada_openai Dataset Summary This dataset is comprised of the LAMBADA test split as pre-processed by OpenAI (see relevant discussions here and here). It also contains machine translated versions of the split in German, Spanish, French, and Italian. LAMBADA is used to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative texts sharing the characteristic that human subjects are able to guess their last word… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lambada_openai.text10K<n<100K49 likes110k downloads1y agoHugging Face02cimec /lambada Dataset Card for LAMBADA Dataset Summary The LAMBADA evaluates the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative passages sharing the characteristic that human subjects are able to guess their last word if they are exposed to the whole passage, but not if they only see the last sentence preceding the target word. To succeed on LAMBADA, computational models cannot simply rely on local… See the full description on the dataset page: https://huggingface.co/datasets/cimec/lambada.text10K<n<100K67 likes23k downloads3y agoHugging Face03lambda /pokemon-blip-captionsgated Notice of DMCA Takedown Action We have received a DMCA takedown notice from The Pokémon Company International, Inc. In response to this action, we have taken down the dataset. We appreciate your understanding. imagetext-to-imagen<1K314 likes4.9k downloads3y agoHugging Face04Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4 total-300-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4140625 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads14d agoHugging Face05Stage-jh-monitor /total-300-lambda00-s_signal_type6-jh-epoch4 total-300-lambda00-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3875 Action score: 0.43125 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads14d agoHugging Face06Stage-jh-monitor /total-300-lambda05-s_signal_type6-jh-epoch4 total-300-lambda05-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.35703125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads14d agoHugging Face07Stage-jh-monitor /total-300-lambda08-s_signal_type6-jh-epoch4 total-300-lambda08-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38046875 Action score: 0.4078125 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads14d agoHugging Face08Stage-jh-monitor /total-300-lambda10-s_signal_type6-jh-epoch4 total-300-lambda10-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.41875 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads14d agoHugging Face09Stage-jh-monitor /total-300noapp-lambda02-s_signal_type6-jh-epoch4 total-300noapp-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.409375 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads14d agoHugging Face10Stage-jh-monitor /total-300app-lambda02-s_signal_type6-jh-epoch4 total-300app-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3625 Action score: 0.4015625 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads14d agoHugging Face11Stage-jh-monitor /total-131-lambda02-residual-s_signal_type6-jh-epoch4 total-131-lambda02-residual-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3765625 Action score: 0.4171875 Valid samples: 320/320 tabularn<1K0 likes4.1k downloads14d agoHugging Face12Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-retry-epoch4 total-300-lambda02-s_signal_type6-jh-retry-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36953125 Action score: 0.3984375 Valid samples: 320/320 tabularn<1K0 likes4k downloads14d agoHugging Face13Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4-reeval2 total-300-lambda02-s_signal_type6-jh-epoch4-reeval2 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4125 Action score: 0.4265625 Valid samples: 320/320 tabularn<1K0 likes4k downloads14d agoHugging Face14Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4-reeval1 total-300-lambda02-s_signal_type6-jh-epoch4-reeval1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38828125 Action score: 0.4234375 Valid samples: 320/320 tabularn<1K0 likes4k downloads14d agoHugging Face15LAMDA-NeSy /ChinaTravel ChinaTravel Query Dataset This dataset is licensed under Creative Commons Attribution 4.0 International (CC BY 4.0). ChinaTravel is an open-ended travel-planning benchmark with compositional constraint validation for language agents. See the paper, Hugging Face paper page, code, and bilingual sandbox database (ModelScope mirror) for the complete benchmark resources. Introduction For a given query, a language agent uses the sandbox tools to collect information and… See the full description on the dataset page: https://huggingface.co/datasets/LAMDA-NeSy/ChinaTravel.tabulartext-generation1K<n<10K13 likes1.3k downloads8d agoHugging Face16lambda /hermes-agent-reasoning-traces Hermes Agent Reasoning Traces Multi-turn tool-calling trajectories for training AI agents using the Hermes Agent harness. Each sample is a real agent conversation with step-by-step reasoning (<think> blocks) and actual tool execution results. This dataset has two configs, one per source model: Config Model Samples kimi Moonshot AI Kimi-K2.5 7,646 glm-5.1 ZhipuAI GLM-5.1-FP8 7,055 Loading from datasets import load_dataset # Kimi-K2.5 traces ds =… See the full description on the dataset page: https://huggingface.co/datasets/lambda/hermes-agent-reasoning-traces.texttext-generation10K<n<100K382 likes1.1k downloads5mo agoHugging Face17IQSeC-Lab /LAMDA LAMDA: A Longitudinal Android Malware Dataset for Drift Analysis This dataset contains a longitudinal benchmark for Android malware detection designed to analyze and evaluate concept drift in machine learning models. It includes labeled and feature-engineered Android APK data from 2013 to 2025 (excluding 2015), with over 1 million samples collected from real-world sources. Dataset Details LAMDA is the largest and most temporally diverse Android malware dataset to date. It… See the full description on the dataset page: https://huggingface.co/datasets/IQSeC-Lab/LAMDA.tabular1M<n<10M6 likes1.1k downloads7mo agoHugging Face18atutej /m_lamaExtension/Modification of the original m_lama dataset text100K<n<1M0 likes991 downloads3y agoHugging Face19PtaSack /LAMDA LAMDA: A Longitudinal Android Malware Dataset for Drift Analysis This dataset contains a longitudinal benchmark for Android malware detection designed to analyze and evaluate concept drift in machine learning models. It includes labeled and feature-engineered Android APK data from 2013 to 2025 (excluding 2015), with over 1 million samples collected from real-world sources. Dataset Details LAMDA is the largest and most temporally diverse Android malware dataset to date. It… See the full description on the dataset page: https://huggingface.co/datasets/PtaSack/LAMDA.tabular1M<n<10M0 likes830 downloads7mo agoHugging Face20lamini /lamini_docs Dataset Card for "lamini_docs" More Information needed text1K<n<10K23 likes733 downloads3y agoHugging Face21lamhieu /medical_advice_dialogue_en Description The dataset is from medalpaca/medical_meadow_health_advice, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it. Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth. Structure View online through viewer. Note We advise you to reconsider before use, thank you. If you find it useful, please like and… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_advice_dialogue_en.texttext-generation1K<n<10K1 likes659 downloads2y agoHugging Face22lamooon /vocalaudio10K<n<100K1 likes632 downloads5mo agoHugging Face23PestoRosso /lamoda-fashion-product-images High-Resolution Fashion Product Images This dataset is a highly optimized, high-resolution subset of the popular Fashion Product Images Dataset originally hosted on Kaggle. It contains thousands of unique e-commerce fashion products, combining high-resolution product images with multiple descriptive label attributes. All low-resolution thumbnails and anomalies have been aggressively filtered out. Every image in this dataset has a minimum resolution of 640px on its shortest… See the full description on the dataset page: https://huggingface.co/datasets/PestoRosso/lamoda-fashion-product-images.imageimage-classification10K<n<100K1 likes577 downloads3mo agoHugging Face24craffel /openai_lambadaLAMBADA dataset variant used by OpenAI to evaluate GPT-2 and GPT-3.text1K<n<10K2 likes482 downloads5y agoHugging Face25lamini /earnings-calls-qa Lamini Earning Calls QA Dataset Description This dataset contains transcripts of earning calls for various companies, along with questions and answers related to the companies' financial performance and other relevant topics. Format The transcripts, questions, and answers are in the form of jsonlines files, with each json object in the file containing the transcript of an earning call for a single company. Data Pipeline Code The entire data pipeline… See the full description on the dataset page: https://huggingface.co/datasets/lamini/earnings-calls-qa.texttext-classification100K<n<1M54 likes479 downloads3y agoHugging Face26Boldt /lambada_openai_de LAMBADA (DE) — Boldt German Evaluation Suite A modernized German translation of the LAMBADA benchmark (Paperno et al., 2016), part of the Boldt German Evaluation Suite. LAMBADA tests a model's ability to track discourse-level context. Each instance consists of a passage where the final word can only be predicted correctly if the model has understood the broader narrative — it cannot be inferred from the final sentence alone. The target word is always the last token of the passage.… See the full description on the dataset page: https://huggingface.co/datasets/Boldt/lambada_openai_de.text1K<n<10K0 likes471 downloads5mo agoHugging Face27janck /bigscience-lama Dataset Card for LAMA: LAnguage Model Analysis - a dataset for probing and analyzing the factual and commonsense knowledge contained in pretrained language models. @inproceedings{petroni2020how, title={How Context Affects Language Models' Factual Predictions}, author={Fabio Petroni and Patrick Lewis and Aleksandra Piktus and Tim Rockt{"a}schel and Yuxiang Wu and Alexander H. Miller and Sebastian Riedel}, booktitle={Automated Knowledge Base Construction}, year={2020}… See the full description on the dataset page: https://huggingface.co/datasets/janck/bigscience-lama.texttext-retrieval10K<n<100K1 likes464 downloads4y agoHugging Face28Yanyi10086 /LAMDA LAMDA: A Longitudinal Android Malware Dataset for Drift Analysis This dataset contains a longitudinal benchmark for Android malware detection designed to analyze and evaluate concept drift in machine learning models. It includes labeled and feature-engineered Android APK data from 2013 to 2025 (excluding 2015), with over 1 million samples collected from real-world sources. Dataset Details LAMDA is the largest and most temporally diverse Android malware dataset to date. It… See the full description on the dataset page: https://huggingface.co/datasets/Yanyi10086/LAMDA.tabular1M<n<10M0 likes440 downloads9mo agoHugging Face29lamini /bird_text_to_sql Dataset Card for "bird_text_to_sql" More Information needed text10K<n<100K7 likes433 downloads3y agoHugging Face30MBZUAI /LaMini-instruction Dataset Card for "LaMini-Instruction" Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, Alham Fikri Aji, Dataset Description We distill the knowledge from large language models by performing sentence/offline distillation (Kim and Rush, 2016). We generate a total of 2.58M pairs of instructions and responses using gpt-3.5-turbo based on several existing resources of prompts, including self-instruct (Wang et al., 2022), P3 (Sanh et al., 2022), FLAN… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/LaMini-instruction.text1M<n<10M147 likes425 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.