CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01meta-math /MetaMathQAView the project page: https://meta-math.github.io/ see our paper at https://arxiv.org/abs/2309.12284 Note All MetaMathQA data are augmented from the training sets of GSM8K and MATH. None of the augmented data is from the testing set. You can check the original_question in meta-math/MetaMathQA, each item is from the GSM8K or MATH train set. Model Details MetaMath-Mistral-7B is fully fine-tuned on the MetaMathQA datasets and based on the powerful Mistral-7B model. It is… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/MetaMathQA.text100K<n<1M476 likes113k downloads3y agoHugging Face02hf-internal-testing /imagefolder_with_metadataimagen<1K0 likes57k downloads2y agoHugging Face03meta-agents-research-environments /gaia2 Gaia2 Paper | Code | Project Page Dataset Summary Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically. The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.textreinforcement-learningn<1K46 likes36k downloads1y agoHugging Face04meta-agents-research-environments /gaia2_filesystem GAIA2 Filesystem This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset. Dataset Link https://huggingface.co/datasets/meta-agents-research-environments/gaia2 Contact Details Publishing POC: Meta AI Research Team Affiliation: Meta Platforms, Inc. Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.imagen<1K1 likes20k downloads1y agoHugging Face05model-metadata /code_execution_filestextn<1K0 likes12k downloads7mo agoHugging Face06suriyagunasekar /stackoverflow-with-meta-data Dataset Card for "stackoverflow-with-meta-data" More Information needed text10M<n<100M13 likes5.5k downloads4y agoHugging Face07opendatalab /SlimPajama-Meta-rater Annotated SlimPajama Dataset Dataset Description This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions. Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.tabulartext-generation10M<n<100M7 likes5.3k downloads1y agoHugging Face08facebook /meta-active-readingtext1B<n<10B37 likes5k downloads1y agoHugging Face09bigcode /the-stack-metadata Dataset Card for The Stack Metadata Changelog Release Description v1.1 This is the first release of the metadata. It is for The Stack v1.1 v1.2 Metadata dataset matching The Stack v1.2 Dataset Summary This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories. Supported Tasks and Leaderboards The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.tabulartext-generation10B<n<100B10 likes4.9k downloads4y agoHugging Face10librarian-bots /arxiv-metadata-snapshot Dataset Card for "arxiv-metadata-oai-snapshot" More Information needed This is a mirror of the metadata portion of the arXiv dataset. The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset. Metadata This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing: id: ArXiv ID (can be used to access the paper, see below) submitter:… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/arxiv-metadata-snapshot.texttext-generation1M<n<10M22 likes4.7k downloads5d agoHugging Face11dsixteen /Niji_1_Man-metaimage10K<n<100K0 likes4.6k downloads4d agoHugging Face12SaylorTwift /details_meta-llama__Llama-3.1-8B-Instruct_private Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct. The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.textn<1K0 likes4.5k downloads1y agoHugging Face13model-metadata /custom_code_execution_filestext1K<n<10K1 likes4.1k downloads11mo agoHugging Face14allenai /metaicl-dataThis is the downloaded and processed data from Meta's MetaICL. We follow their "How to Download and Preprocess" instructions to obtain their modified versions of CrossFit and UnifiedQA. Citation information @inproceedings{ min2022metaicl, title={ Meta{ICL}: Learning to Learn In Context }, author={ Min, Sewon and Lewis, Mike and Zettlemoyer, Luke and Hajishirzi, Hannaneh }, booktitle={ NAACL-HLT }, year={ 2022 } } @inproceedings{ ye2021crossfit, title={ {C}ross{F}it:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/metaicl-data.text100K<n<1M5 likes3.9k downloads4y agoHugging Face15meta-math /MetaMathQA-40Karxiv.org/abs/2309.12284 View the project page: https://meta-math.github.io/ text10K<n<100K27 likes3.3k downloads3y agoHugging Face16huggingface /transformers-metadata Transformers metadata text1K<n<10K44 likes3.1k downloads2d agoHugging Face17daisuke9999 /Japanese_NicoNico_Douga_Movie_Meta_Data_2016tabular10M<n<100M0 likes3k downloads2y agoHugging Face18bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes2.7k downloads4y agoHugging Face19friendshipkim /metaiclThis is the downloaded and processed data from Meta's MetaICL. We follow their "How to Download and Preprocess" instructions to obtain their modified versions of CrossFit and UnifiedQA. Citation information @inproceedings{ min2022metaicl, title={ Meta{ICL}: Learning to Learn In Context }, author={ Min, Sewon and Lewis, Mike and Zettlemoyer, Luke and Hajishirzi, Hannaneh }, booktitle={ NAACL-HLT }, year={ 2022 } } @inproceedings{ ye2021crossfit, title={ {C}ross{F}it:… See the full description on the dataset page: https://huggingface.co/datasets/friendshipkim/metaicl.text1M<n<10M0 likes2.2k downloads2y agoHugging Face20MedOtter /Colorectal-Liver_Metastases Colorectal-Liver-Metastases (CRLM) Preoperative contrast-enhanced CT scans + manual radiologist segmentations from 197 patients who underwent hepatic resection for colorectal liver metastases at Memorial Sloan Kettering Cancer Center (MSKCC). Released under TCIA in 2023 alongside the Scientific Data descriptor by Simpson et al. (2024). Dataset Details Field Value Modality CT (preoperative, portal-venous phase, contrast-enhanced MDCT) Body part Liver… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Colorectal-Liver_Metastases.imageimage-segmentationn<1K1 likes2k downloads3mo agoHugging Face21huggingface /diffusers-metadatatextn<1K35 likes2k downloads15h agoHugging Face22labofsahil /pypi-packages-metadata-datasettext10M<n<100M0 likes1.7k downloads8mo agoHugging Face23xcpan /MetaQuery_Instruct_2.4M_512resThe data is licensed CC-by-NC. Third party content pulled from other locations are subject to their own licenses and you may have other legal obligations or restrictions that govern your use of that content. The MetaQuery dataset is also released under ODC-BY and Common Crawl terms of use, because it is sourced from mmc4. imagetext-to-image1M<n<10M10 likes1.6k downloads1y agoHugging Face24trojblue /danbooru2025-metadata 🎨 Danbooru 2025 Metadata Latest Post ID: 9,158,800 (as of Apr 16, 2025) 📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork. Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes: More consistent tag history tracking Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.imagetext-to-image1M<n<10M38 likes1.6k downloads1y agoHugging Face25TAUR-Lab /Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructtext10K<n<100K0 likes1.6k downloads2y agoHugging Face26mmarone /fineweb-edu-full-metadata[WIP] FineWeb-Edu with Metadata This repo contains 3 versions of the FineWeb-Edu v1 dataset: fwedu1-metaonly/ fwedu1-text-content-zstd/ fineweb-edu-1.0.0-meta-and-text/ These are all joinable via the hash column, which is xxhash64 in pyspark, calculated on the text column. This hash is unique for all instances in the dataset. For convenience, this join is done for you in the third table fwedu1-metaonly is just the metadata of the data exactly as it comes from the FineWeb-Edu v1… See the full description on the dataset page: https://huggingface.co/datasets/mmarone/fineweb-edu-full-metadata.tabular100M<n<1B0 likes1.5k downloads1y agoHugging Face27Metanova /Proteinstext1M<n<10M2 likes1.4k downloads1y agoHugging Face28smartcat /Amazon_Sample_Metadata_2023 Dataset Card for Dataset Name Original datasets can be found on: https://amazon-reviews-2023.github.io/ Dataset Details This dataset was made as sample of several datasets from the link above. Dataset Description This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors, Health and Personal Care, Amazon Clothing Shoes and Jewlery, Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.tabular1M<n<10M1 likes1.3k downloads2y agoHugging Face29ReadyAi /5000-podcast-conversations-with-metadata-and-embedding-dataset 🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications. This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network. AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.text10K<n<100K8 likes1.3k downloads1y agoHugging Face30meta-math /GSM8K_zh Dataset GSM8K_zh is a dataset for mathematical reasoning in Chinese, question-answer pairs are translated from GSM8K (https://github.com/openai/grade-school-math/tree/master) by GPT-3.5-Turbo with few-shot prompting. The dataset consists of 7473 training samples and 1319 testing samples. The former is for supervised fine-tuning, while the latter is for evaluation. for training samples, question_zh and answer_zh are question and answer keys, respectively; for testing samples, only… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/GSM8K_zh.textquestion-answering1K<n<10K30 likes1.3k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.