CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01semeru /text-code-galeras-code-generation-from-docstring-3k-dedupedtabular1K<n<10K0 likes598 downloads3y agoHugging Face02vwxyzjn /summarize_from_feedback_tldr_3_filteredThis is the query dataset taken directly from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset textsummarization100K<n<1M1 likes456 downloads3y agoHugging Face03silk-road /ChatHaruhi-from-RoleLLMAdapt English Role in RoleBench into ChatHaruhi format only using profiles part in ZenMoore/RoleBench Great thanks to on authors of RoleLLM! usage: # if you pip installed chatharuhi it should be # from chatharuhi import ChatHaruhi from ChatHaruhi import ChatHaruhi chatbot = ChatHaruhi( role_from_hf = 'silk-road/ChatHaruhi-from-RoleLLM/Sherlock Holmes', \ llm = 'openai', embedding = 'bge_en') response = chatbot.chat(role='Police Chief', text = 'Oh… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/ChatHaruhi-from-RoleLLM.text10K<n<100K3 likes399 downloads3y agoHugging Face04PegasusJEsus /DATASET_FROM_ALL_DOMAINStext100M<n<1B1 likes274 downloads7d agoHugging Face05DeSTA-ntu /DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct📑 Paper | 👩‍💻 Github | 🤗 Model | 🤗 Dataset DeSTA-AQA5M comprises 50 speech, environmental sound, and music datasets, totaling over 7,000 hours of audio. Our training framework centers on self-generated response for efficient cross-modal alignment. (see our paper!). In DeSTA, each audio clip is first transformed into a textual description using its metadata. A Large Language Model (LLM) is then prompted with this description to self-generate a response. Ultimately, we construct a… See the full description on the dataset page: https://huggingface.co/datasets/DeSTA-ntu/DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct.text1M<n<10M5 likes231 downloads1y agoHugging Face06fromziro /EOT-2004-Raw End of Term 2004 Original dump: https://eotarchive.org/data/data-2004/ The End of Term Web Archive is a crawl of U.S. government websites conducted at the end of each presidential administration. This is a filtered version of the 2004 crawl. Notice This dataset is still a work in progress. Data Curation We download the 2004 EOT WARC files and parse the HTML using Trafilatura. We then filter the extracted text by length (minimum of 550 characters)… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/EOT-2004-Raw.tabular100K<n<1M1 likes208 downloads2mo agoHugging Face07fromziro /jetoncount_corpus JetonCount's Corpus This is the corpus used to train JetonCount. JSONL Format { "source_file": "token_stats\\HuggingFaceFW_fineweb-edu00000.jsonl", "dataset_dir": "HuggingFaceFW/fineweb-edu", "index": 4020, "chars": 2178, "words": 335, "avg_chars_per_word": 5.504478, "longest_word_chars": 33, "punctuation_ratio": 0.037649, "symbol_ratio": 0.00551, "tokens": 664, "vocab_size": 2560, "tokenizer_dir": "fromziro/Er-Tiny-1.3M" }… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/jetoncount_corpus.tabularfeature-extraction10M<n<100M1 likes198 downloads3mo agoHugging Face08tickleliu /all-skills-from-skills-sh Dataset Overview This dataset is collected from skills.sh, a website that aggregates various command-line skills. Each skill typically includes a brief description, detailed documentation, and an installation command. We have crawled all publicly available skill pages, resulting in approximately 40,000 to 50,000 skill records. The data is stored in a structured format, making it suitable for analysis, retrieval, or further development. Field Descriptions Each skill… See the full description on the dataset page: https://huggingface.co/datasets/tickleliu/all-skills-from-skills-sh.text10K<n<100K7 likes117 downloads7mo agoHugging Face09mondk /Claude-classified_from-Manusagentsreal: Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset but categorized, retaining only the complete sections from Claude. text10K<n<100K4 likes95 downloads27d agoHugging Face10fpan /text-to-ocl-from-ecore Introduction This is a small size dataset containing 52 meta-models (EMF files and PlantUML descriptions), 369 OCL constraints and 369 constraint specification in natural language. The meta-models and OCL constraints are collected from open source github projects and are (syntactically) processable by Eclipse. The constraint specifications of OCL constraints are generated via GPT-4-Turbo. The meta-models can be found in models\ Usage Generation of OCL constraints based on… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.texttranslationn<1K0 likes93 downloads2y agoHugging Face11fpan /text-to-xmi-from-ecoreThis is a small test set for XMI instance model generation task. It containing 26 pairs of meta-models (Ecore), specifications (natural language) and instance models (XMI). In each pair, the meta-model and instance model share the same name. To proper open the instance model in Eclipse EMF, the instance model and meta-model should be placed in the same folder. The meta-models are selected from https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore. The specifications are generated via… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-xmi-from-ecore.texttext-generationn<1K0 likes71 downloads1y agoHugging Face12fromziro /py-docs-2004 Python Docs 2004 Original dump: https://www.python.org/ftp/python/doc/ Python Docs 2004 is a filtered and cleaned collection of Python documentation from every major Python release published before 2004. Stats Version Size Lines 2.3 2.2MB 1215 2.2 1.7MB 1142 2.1 1.3MB 891 2.0 1.2MB 895 1.6 1MB 720 1.5 837KB 449 1.4 744KB 397 1.3 569KB 408 1.2 513KB 384 Total 10.1MB 6501 Notice This dataset is a filtered and cleaned… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/py-docs-2004.texttext-generation10K<n<100K0 likes64 downloads2mo agoHugging Face13Learning-from-Peers /DeepSeek-R1-Distill-Qwen-32B-LeaPPaper: Learning from Peers in Reasoning Models Project Page: https://learning-from-peers.github.io/ Code: https://github.com/tongxuluo/LeaP textquestion-answering1K<n<10K1 likes61 downloads1y agoHugging Face14fromhope /acubench AcuBench Indication-based acupoint-set recommendation. Given an indication / symptom string (e.g. "Headache"), predict the set of WHO-standard acupoints indicated for it, grounded in AcuKG's Indication table, over a fixed 361-point label space. AcuBench is a small (446-sample) benchmark with a dedicated conformal-prediction calibration split. Not a clinical prescription benchmark. A row's acupoint set is "acupoints indicated for this symptom in AcuKG", i.e. a candidate pool… See the full description on the dataset page: https://huggingface.co/datasets/fromhope/acubench.texttext-classificationn<1K0 likes60 downloads28d agoHugging Face15philosopher-from-god /HuggingChat-AI-Assistants-Deleted-System-Promptstextn<1K1 likes57 downloads6mo agoHugging Face16ewhk9887 /korean_code_reviews_from_githubtext10K<n<100K1 likes56 downloads2y agoHugging Face17CL-From-Nothing /rose_code_samples rose_code samples (pass@8 rollouts) vLLM pass@8 samples on the CL-From-Nothing/rose_code train split (23,688 codeforces stdin/stdout problems), scored by the deepcoder verifier (reward=1.0 iff all test cases pass). Qwen3-1.7B/ — student model rollouts. 23,688 questions × 8 samples = 189,504 lines. Qwen3-4B-Thinking-2507/ — teacher model rollouts. Sampling: temperature 0.7, top_p 0.9, max_tokens 16384, 8 samples/question (pass@8). Each cluster file holds a contiguous… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/rose_code_samples.tabulartext-generation100K<n<1M0 likes55 downloads4mo agoHugging Face18SeanWang0027 /polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507 Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's actual continuation for each, and the token accounting behind it. The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is decoded to text, the teacher is shown it under its own chat template, and the teacher's reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.tabulartext-generation10K<n<100K0 likes55 downloads27d agoHugging Face19ErrareHumanumEst /internal-state-from-gibberish Injection-method A/B data (introspection-leakage robustness check) Data behind the "Robustness: does the leak depend on the injection method?" section of reports/experiment.md. Each .pt is one full collection run for one Qwen2.5 model under one injection method, at the original per-model dose (auto-tuner off, so the dose is identical across arms). Builds on open-introspection (Otto Stegmaier) — concept set, difference-vector extraction, and layer/strength calibration. See the… See the full description on the dataset page: https://huggingface.co/datasets/ErrareHumanumEst/internal-state-from-gibberish.tabularn<1K0 likes49 downloads2mo agoHugging Face20CL-From-Nothing /RLVE_envs65_data1000text10K<n<100K0 likes48 downloads6mo agoHugging Face21visv-Bro /repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes48 downloads2mo agoHugging Face22fromziro /arxiv-abstracts-2004 ArXiv Abstracts 2004 Original Dataset: common-pile/arxiv_abstracts ArXiv-Abstracts-2004 is a filtered collection of abstracts from the Common-Pile ArXiv dataset containing works created on or before 2004. Stats Size (MB) Lines 351MB 303,761 Note: The lines, in the .jsonl file, are ordered from oldest to newest. Notice We do not claim ownership of or credit for any prior work done by the Common-Pile team. This dataset is only a… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/arxiv-abstracts-2004.texttext-generation100K<n<1M1 likes47 downloads2mo agoHugging Face23semeru /code-code-galeras-code-completion-from-docstring-3k-dedupedtabular1K<n<10K0 likes44 downloads3y agoHugging Face24CL-From-Nothing /ROSE-polaris-popetabular10K<n<100K0 likes44 downloads4mo agoHugging Face25martimfasantos /summarize-from-feedback-protext100K<n<1M0 likes43 downloads2y agoHugging Face26divaspoudel /repro-a-theory-of-learning-data-statistics-in-diffusion-models-from-easy-to-hard-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes41 downloads2mo agoHugging Face27riteshhf /repro-from-prior-to-pro-efficient-skill-mastery-via-distribution-contractive-rl-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes41 downloads2mo agoHugging Face28databounty-io /extract-metrics-from-log-fixtures-cmskdlip Extract Metrics from Log Fixtures A dataset of extract metrics from log fixtures examples for training and evaluation. Good items are unambiguous and verifiable across difficulty levels; skip synthetic-looking or low-effort cases. About This dataset was produced by the DataBounty community and published here as part of an open, karma-only program. Contributor items exported: 1000 Language: Python Framework: Community License: CC-BY-4.0 Contributors… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/extract-metrics-from-log-fixtures-cmskdlip.text1K<n<10K0 likes40 downloads18d agoHugging Face29GeoGPT-Research-Project /GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl Description This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset. This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl: id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.texttext-generation10M<n<100M1 likes37 downloads1y agoHugging Face30CL-From-Nothing /RLVE-Qwen3-1.7B-Pass1-Rollouts RLVE teacher rollouts — Qwen3-1.7B (pass@1) Teacher rollouts for on-policy distillation on the RLVE environment suite. Teacher / sampler: Qwen3-1.7B Source prompts: RLVE train split — 9000 questions across RLVE-Eval Gym environments (counting / combinatorics / optimization tasks) Sampling: 1 sample/question (pass@1) = 9000 records, temperature 0.7, max 4096 new tokens Rewards: inline RLVE-Eval Gym verifier score (continuous, in [-1, 1]). Teacher accuracy (reward>0): 20 / 9000 =… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/RLVE-Qwen3-1.7B-Pass1-Rollouts.tabulartext-generation1K<n<10K0 likes35 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.