CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Green-Sky /mmlu-redux-2.0-for-llama.cppMMLU-redux-v2.0 converted for the llama.cpp perplexity multiple choice tool. Only valid entries where kept, there is no error based prompting included. Dataset Card for MMLU-Redux-2.0 MMLU-Redux is a subset of 5,700 manually re-annotated questions across 57 MMLU subjects. Citation BibTeX: @misc{gema2024mmlu, title={Are We Done with MMLU?}, author={Aryo Pradipta Gema and Joshua Ong Jun Leang and Giwon Hong and Alessio Devoto and Alberto Carlo Maria… See the full description on the dataset page: https://huggingface.co/datasets/Green-Sky/mmlu-redux-2.0-for-llama.cpp.question-answering1K<n<10K0 likes441 downloads6mo agoHugging Face02Green-Sky /mmlu-redux-for-llama.cppMMLU-redux converted for the llama.cpp perplexity multiple choice tool. Only valid entries where kept, there is no error based prompting included. Dataset Card for MMLU-Redux [!TIP] Please consider using MMLU-Redux-2.0 which contains all 57 MMLU subjects. MMLU-Redux is a subset of 3,000 manually re-annotated questions across 30 MMLU subjects. Citation BibTeX: @misc{gema2024mmlu, title={Are We Done with MMLU?}, author={Aryo Pradipta Gema and Joshua Ong… See the full description on the dataset page: https://huggingface.co/datasets/Green-Sky/mmlu-redux-for-llama.cpp.question-answering1K<n<10K0 likes200 downloads6mo agoHugging Face03GreenNode /nano-hotpotqa-vn NanoHotpotQA-VN An MTEB dataset Massive Text Embedding Benchmark A translated dataset from HotpotQA is a question answering dataset featuring natural, multi-hop questions, with strong supervision for supporting facts to enable more explainable question answering systems. The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-hotpotqa-vn.texttext-retrieval100K<n<1M0 likes70 downloads9mo agoHugging Face04greencalculus /emission-factor-benchmark Emission-Factor Accuracy Benchmark 3,299 rows. Five frontier models answering identical factual questions, with ground truth traced to a named document and an exact cell — plus the same questions re-run with a lookup tool, and a second study on which data vendors those models recommend unprompted. Collected 10 September 2026. Models: claude-opus-5, gpt-5.5, gemini-3.1-pro-preview, gemini-3.6-flash, grok-4.6. All answers were produced through each provider's API with no tools and… See the full description on the dataset page: https://huggingface.co/datasets/greencalculus/emission-factor-benchmark.tabularquestion-answering1K<n<10K1 likes62 downloads15d agoHugging Face05Green-Sky /LongBench-v2-for-llama.cppLongBench v2 converted for the llama.cpp perplexity multiple choice tool. [!WARNING] !! Currently does not work, will fix it in the near future. Probably. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems… See the full description on the dataset page: https://huggingface.co/datasets/Green-Sky/LongBench-v2-for-llama.cpp.textmultiple-choicen<1K0 likes47 downloads6mo agoHugging Face06GreenNode /nano-msmarco-vn NanoMSMARCO-VN An MTEB dataset Massive Text Embedding Benchmark A translated dataset from MS MARCO is a collection of datasets focused on deep learning in search The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding models to filter the translations. - Use LLM-as-a-judge to… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-msmarco-vn.texttext-retrieval100K<n<1M0 likes43 downloads9mo agoHugging Face07greenwich157 /telco-5G-data-faultsSynthetic test dataset for 5G data service faults in the core and RAN network domains. It is used to train the telcoLLM to simulate assistance model for network operations. textquestion-answeringn<1K4 likes40 downloads1y agoHugging Face08GreenNode /nano-nq-vn NanoNQ-VN An MTEB dataset Massive Text Embedding Benchmark A translated dataset from NFCorpus: A Full-Text Learning to Rank Dataset for Medical Information Retrieval The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding models to filter the translations. - Use LLM-as-a-judge… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-nq-vn.texttext-retrieval100K<n<1M0 likes33 downloads9mo agoHugging Face09greenfit-ai /synthetic-sport-products-sustainability Dataset Card for synthetic-sport-products-sustainability This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/as-cle-bert/synthetic-sport-products-sustainability/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/greenfit-ai/synthetic-sport-products-sustainability.texttext-generationn<1K1 likes27 downloads2y agoHugging Face10GreenRalph /rootmodel-knf-philippines-v1 rootmodel-knf-philippines-v1 Adaptive agricultural instruction dataset for regenerative tropical farming informed by Korean Natural Farming (KNF), built from a working farm in Nabua, Camarines Sur, Bicol, Philippines. Released for the AutoScientist Challenge — Agriculture (Part 2, 2026). To the maintainer's knowledge, no equivalent KNF-specific instruction dataset currently exists in the public domain. "Modern AI was trained on the internet. ROOTMODEL is trained on living… See the full description on the dataset page: https://huggingface.co/datasets/GreenRalph/rootmodel-knf-philippines-v1.texttext-generationn<1K0 likes12 downloads2mo agoHugging Face11Nurlykhan /GreenBond-Spillover-Instruct GreenBond-Spillover-Instruct Specialized instruction-tuning dataset for sovereign green bond analysis, spillover-effect detection, and narrative risk assessment. Dataset Info Property Value Total examples 3,000 Format JSONL Fields instruction, input, output Language English Category Distribution Category Count Sentiment classification 752 Greenwashing detection 593 Narrative tagging 467 Spillover Q&A 331 Event extraction… See the full description on the dataset page: https://huggingface.co/datasets/Nurlykhan/GreenBond-Spillover-Instruct.texttext-classification1K<n<10K0 likes10 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.