CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LadyMia /x_dataset_44311 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_44311.texttext-classification100M<n<1B0 likes983 downloads1y agoHugging Face02LadyMia /x_dataset_7480 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_7480.texttext-classification100M<n<1B0 likes684 downloads1y agoHugging Face03LadyMia /x_dataset_14253 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_14253.texttext-classification100M<n<1B0 likes450 downloads1y agoHugging Face04LadyMia /x_dataset_40563 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_40563.texttext-classification100M<n<1B0 likes354 downloads1y agoHugging Face05LadyMia /x_dataset_21893 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_21893.texttext-classification100M<n<1B0 likes339 downloads1y agoHugging Face06LadyMia /x_dataset_51674 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_51674.texttext-classification100M<n<1B0 likes261 downloads1y agoHugging Face07LadyMia /x_dataset_41147 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_41147.texttext-classification100M<n<1B0 likes239 downloads1y agoHugging Face08LadyMia /x_dataset_21318 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_21318.texttext-classification100M<n<1B0 likes233 downloads1y agoHugging Face09LadyMia /x_dataset_2447 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_2447.texttext-classification100M<n<1B0 likes220 downloads1y agoHugging Face10ladybugdb /scg-columnar Dataset Card for the Scientific Contribution Graph (SCG) Compressed Parquet release: this Hub dataset is a compact Parquet encoding of the Scientific Contribution Graph.It reduces the on-disk footprint from ~80 GB uncompressed JSON/tree files → ~6 GB Parquet, while preserving the same contributions, prerequisite edges, and metadata as SCG v1.1. The Scientific Contribution Graph maps how science is built “on the shoulders of giants”: it extracts fine-grained scientific… See the full description on the dataset page: https://huggingface.co/datasets/ladybugdb/scg-columnar.graph-ml1M<n<10M0 likes209 downloads5h agoHugging Face11LadyMia /x_dataset_55395 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_55395.texttext-classification100M<n<1B0 likes193 downloads1y agoHugging Face12LadyMia /x_dataset_36129 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_36129.texttext-classification100M<n<1B0 likes172 downloads1y agoHugging Face13LadyMia /x_dataset_41414 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_41414.texttext-classification100M<n<1B0 likes130 downloads1y agoHugging Face14LadyMia /x_dataset_17682 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_17682.texttext-classification100M<n<1B0 likes91 downloads1y agoHugging Face15LadyMia /x_dataset_12949 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_12949.texttext-classification100M<n<1B0 likes83 downloads1y agoHugging Face16anonymous-insightladder-2026 /insight-ladder-imo2024 Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission). Overview A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with: 4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.tabulartext-generationn<1K0 likes83 downloads5mo agoHugging Face17LadyMia /x_dataset_24095 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_24095.texttext-classification100M<n<1B0 likes64 downloads1y agoHugging Face18LadyMia /x_dataset_47139 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_47139.texttext-classification100M<n<1B0 likes53 downloads1y agoHugging Face19LadyMia /x_dataset_63681 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_63681.texttext-classification100M<n<1B0 likes50 downloads1y agoHugging Face20LadiesMan69 /FastAPI_dataset FastAPI Documentation Assistant Dataset Dataset Summary Dataset Structure Data Fields Data Splits Dataset Creation Usage Limitations Licensing Citation Dataset Summary This dataset was used to fine-tune Qwen2.5-Coder-7B-FastAPI-LoRA, a LoRA adapter built on top of unsloth/Qwen2.5-Coder-7B-Instruct-bnb-4bit. It consists of instruction-style examples that turn the model into a FastAPI documentation assistant: given a snippet of FastAPI documentation as context and… See the full description on the dataset page: https://huggingface.co/datasets/LadiesMan69/FastAPI_dataset.text-generationn<1K0 likes48 downloads2mo agoHugging Face21LadyMia /x_dataset_63648 Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/LadyMia/x_dataset_63648.texttext-classification100M<n<1B0 likes44 downloads1y agoHugging Face22laddermedia /srs-prompts SRS Prompts srs-prompts is a collection of highlight-anchored memory prompts, each paired with its highlight, source context, and a tier label (T0–T3). See the Memory Machines appendix for details on the dataset, including the taxonomy, collection methodology, and schema. A companion dataset, laddermedia/srs-highlights, aggregates these prompts back to the highlight level. tabulartext-classification1K<n<10K8 likes40 downloads5mo agoHugging Face23laddermedia /srs-highlights SRS Highlights srs-highlights is a collection of reader highlights, each paired with its source context and the full set of rated memory prompts written for it. It is the highlight-level view of laddermedia/srs-prompts: one row per highlight, with every tier-labeled prompt attached as a reference. See the Memory Machines appendix for details on the dataset, including the taxonomy, collection methodology, and schema. texttext-generationn<1K0 likes35 downloads5mo agoHugging Face24collectivat /salom-ladino-articles Şalom Ladino articles text corpus Text corpus compiled from 397 articles from the Judeo-Espanyol section of Şalom newspaper. Original sentences and articles belong to Şalom. Size: 176,843 words Citation If you use this dataset, please cite: Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish This repository is developed as part of project "Judeo-Spanish: Connecting the two ends of the Mediterranean" carried out by Col·lectivaT and… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/salom-ladino-articles.texttext-generation10K<n<100K0 likes29 downloads11mo agoHugging Face25ladyrhanes /alpaca Dataset Card for Alpaca Dataset Summary Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. The authors built on the data generation pipeline from Self-Instruct framework and made the following modifications: The text-davinci-003 engine to generate the instruction data instead… See the full description on the dataset page: https://huggingface.co/datasets/ladyrhanes/alpaca.texttext-generation10K<n<100K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.