CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gregb /product-taxonomy-bench Dataset Summary product-taxonomy-bench is an anonymised benchmark dataset for predicting Shopify Product Taxonomy categories from Shopify product tags. This dataset does not include raw product titles, raw tags, or product URLs. Tags are anonymised as tagNNNNNN. Start Here Read this dataset card for the snapshot layout and field definitions. Open the benchmark notebook at notebooks/product_taxonomy_bench.ipynb. It defaults to the fixed paper snapshot at revision… See the full description on the dataset page: https://huggingface.co/datasets/gregb/product-taxonomy-bench.tabulartext-classification10K<n<100K0 likes209 downloads14d agoHugging Face02greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes143 downloads2mo agoHugging Face03reapxdev /greenhouse-jobs-scraper Greenhouse Jobs Scraper Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL. Rows in this dataset 14,091 Fields 36 Collector runs behind it 61 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.tabular10K<n<100K0 likes101 downloads2mo agoHugging Face04gregH /OccuBench OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models Dataset Description OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation. Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.tabulartext-generationn<1K5 likes63 downloads5mo agoHugging Face05greencalculus /emission-factor-benchmark Emission-Factor Accuracy Benchmark 3,299 rows. Five frontier models answering identical factual questions, with ground truth traced to a named document and an exact cell — plus the same questions re-run with a lookup tool, and a second study on which data vendors those models recommend unprompted. Collected 10 September 2026. Models: claude-opus-5, gpt-5.5, gemini-3.1-pro-preview, gemini-3.6-flash, grok-4.6. All answers were produced through each provider's API with no tools and… See the full description on the dataset page: https://huggingface.co/datasets/greencalculus/emission-factor-benchmark.tabularquestion-answering1K<n<10K1 likes59 downloads13d agoHugging Face06open-llm-leaderboard /DoppelReflEx__MN-12B-Mimicore-GreenSnake-detailsgated Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-GreenSnake Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-GreenSnake The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-GreenSnake-details.tabular10K<n<100K0 likes45 downloads2y agoHugging Face07open-llm-leaderboard /OnlyCheeini__greesychat-turbo-detailsgated Dataset Card for Evaluation run of OnlyCheeini/greesychat-turbo Dataset automatically created during the evaluation run of model OnlyCheeini/greesychat-turbo The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OnlyCheeini__greesychat-turbo-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face08G-reen /cc-2021-raw cc-2021-raw English web documents extracted from the Common Crawl CC-MAIN-2021-49 snapshot, intended as a pre-2022 human-authored text corpus (i.e. crawled before generative-model output became widespread on the web). Pipeline Stream — Common Crawl WET records, prefiltered on length, replacement-character ratio, and printable/alphabetic character ratios. Language filter — langdetect at p >= 0.95 on three sampled spans of each document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2021-raw.tabulartext-generation1M<n<10M0 likes36 downloads2mo agoHugging Face0913point5 /swe-grep-rlm-reputable-recent-5plus swe-grep-rlm-reputable-recent-5plus This dataset is a GitHub-mined collection of issue- or PR-linked retrieval examples for repository-level code search and localization. Each row is built from a merged pull request in a reputable, actively maintained open-source repository. The target labels are the PR's changed files, with a focus on non-test files. Summary Rows: 799 Repositories: 46 Query source: 519 rows use linked issue title/body when GitHub exposed it 280 rows… See the full description on the dataset page: https://huggingface.co/datasets/13point5/swe-grep-rlm-reputable-recent-5plus.tabulartext-retrievaln<1K0 likes19 downloads6mo agoHugging Face10open-llm-leaderboard /aixonlab__Grey-12b-detailsgated Dataset Card for Evaluation run of aixonlab/Grey-12b Dataset automatically created during the evaluation run of model aixonlab/Grey-12b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/aixonlab__Grey-12b-details.tabular10K<n<100K0 likes18 downloads2y agoHugging Face11GreenBed4725 /arxiv-cs2021-embeddings-bge-m3tabularn<1K0 likes18 downloads2d agoHugging Face12open-llm-leaderboard /GreenNode__GreenNode-small-9B-it-detailsgated Dataset Card for Evaluation run of GreenNode/GreenNode-small-9B-it Dataset automatically created during the evaluation run of model GreenNode/GreenNode-small-9B-it The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/GreenNode__GreenNode-small-9B-it-details.tabular10K<n<100K0 likes14 downloads2y agoHugging Face13G-reen /sumthink_responses_summarizedSee: https://huggingface.co/datasets/G-reen/sumthink This contains the summarized GLM 4.7 nvfp4 thinking traces as well as the source data. tabular10K<n<100K0 likes9 downloads9mo agoHugging Face14G-reen /cc-2020-raw cc-2020-raw English web documents extracted from the Common Crawl CC-MAIN-2020-50 snapshot, intended as a pre-2022 human-authored text corpus (i.e. crawled before generative-model output became widespread on the web). Pipeline Stream — Common Crawl WET records, prefiltered on length, replacement-character ratio, and printable/alphabetic character ratios. Language filter — langdetect at p >= 0.95 on three sampled spans of each document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2020-raw.tabulartext-generation1M<n<10M0 likes7 downloads2mo agoHugging Face15Edols /robomme-vus-green-blue-varied-layout-20260804tabularn<1K0 likes5 downloads2mo agoHugging Face16Benxiaogu /dp_greencuubetabularn<1K0 likes3 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.