CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01epfl-llm /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.texttext-generation10K<n<100K158 likes3k downloads3y agoHugging Face02guildlm /go-swe-bench-v0 go_swe_bench v0 — real Go bug fixes, verified by the Go toolchain 246 tasks from 79 real Go repositories. Each task is a bug-fix commit whose co-committed test is red on the parent and green on the fix. No LLM anywhere in the build. Mined on 2026-09-19 from the GuildLM Go mining pipeline by inverting the filter that had thrown the tests away (the pipeline was built for SFT data; a benchmark needs the opposite). Every task was verified twice with go test: green at the commit (≥ 1… See the full description on the dataset page: https://huggingface.co/datasets/guildlm/go-swe-bench-v0.texttext-generationn<1K0 likes376 downloads16h agoHugging Face03KMK040412 /guiowl-curated-corpus GUI-Owl Curated Corpus This dataset publishes the full curated mobile GUI-agent supervised fine-tuning corpus in a unified norm1000 mobile_use action format. Each row pairs a mobile UI screenshot with an instruction and a normalized target tool call for training GUI agents. The published files are the curated parquet shards as produced by the source canonicalizers. No parquet shards are merged, re-sharded, or sampled during upload. Sources Source Episodes… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/guiowl-curated-corpus.tabularimage-text-to-textn<1K0 likes334 downloads4mo agoHugging Face04hkust-nlp /GUIMid Breaking the Data Barrier – Building GUI Agents Through Task Generalization 🐙 GitHub | 📝 Paper | 🤗 Mid-training Data | 🤗 Post-Training Data TODO List Report and release the GUIMid with larger size and more domains (10th May expecetd) 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/GUIMid.texttext-generation1M<n<10M7 likes223 downloads1y agoHugging Face05AiMijie /EC-Guide This repo is only used for dataset viewer. Please download from here. Amazon KDDCup 2024 Team ZJU-AI4H’s Solution and Dataset (Track 2 Top 2; Track 5 Top 5) The Amazon KDD Cup’24 competition presents a unique challenge by focusing on the application of LLMs in E-commerce across multiple tasks. Our solution for addressing Tracks 2 and 5 involves a comprehensive pipeline encompassing dataset construction, instruction tuning, post-training quantization, and inference… See the full description on the dataset page: https://huggingface.co/datasets/AiMijie/EC-Guide.textquestion-answering10K<n<100K2 likes220 downloads2y agoHugging Face06aisc-team-a1 /guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/guidelines.texttext-generation10K<n<100K0 likes149 downloads3y agoHugging Face07aisc-team-b1 /guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-b1/guidelines.texttext-generation10K<n<100K0 likes145 downloads3y agoHugging Face08jojo-ai-mst /Myanmar-Tuberculosis-Guidelines-Instructions Myanmar Tuberculosis Guidelines Instructions A bilingual instructional dataset built to support Myanmar's ongoing fight against tuberculosis — turning life-saving guidelines into a usable resource for healthcare workers, educators, and AI researchers working with low-resource languages. Authors: Min Si Thu, Khin Myat Noe Abstract Tuberculosis is still one of Myanmar's biggest public health problems. Part of the difficulty is that good, standardized TB education… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Myanmar-Tuberculosis-Guidelines-Instructions.imagequestion-answering1K<n<10K1 likes143 downloads5mo agoHugging Face09guicybercode /japan-math-philosophy-prompts Japan Math Philosophy Prompts Microdataset autoral com problemas que combinam matemática e reflexão filosófica. Há 24 registros: oito instâncias editoriais, cada uma localizada em pt-BR, en e ja e mantida integralmente no split train. Todo o conteúdo foi gerado por modelo e permanece sem revisão humana. As respostas matemáticas funcionam como gabaritos curtos; os critérios filosóficos indicam qualidades esperadas de uma justificativa, não uma opinião obrigatória.… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/japan-math-philosophy-prompts.textquestion-answeringn<1K0 likes119 downloads27d agoHugging Face10juancopi81 /mutopia_guitar_dataset Mutopia Guitar Dataset Dataset Summary Mutopia guitar dataset consists of the soloist guitar pieces of the Mutopia Project. I encoded the MIDI files into text tokens using the excellent implementation of Dr. Tristan Beheren of the paper: MMM: Exploring Conditional Multi-Track Music Generation with the Transformer. The dataset mainly contains guitar music from western classical composers, such as Sor, Aguado, Carcassi, and Giuliani. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/juancopi81/mutopia_guitar_dataset.texttext-generation1K<n<10K6 likes111 downloads4y agoHugging Face11Lots-of-LoRAs /task879_schema_guided_dstc8_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task879_schema_guided_dstc8_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task879_schema_guided_dstc8_classification.texttext-generation1K<n<10K0 likes90 downloads2y agoHugging Face12vldsavelyev /guitar_tabDataset of music tablature, in alphaTex (https://alphatab.net/docs/alphatex) format, converted from Guitar Pro files (gp3, gp4, gp5, which are downloaded from https://rutracker.org/forum/viewtopic.php?t=2888130texttext-generation10K<n<100K11 likes84 downloads3y agoHugging Face13ghemdd /gui_actor_webdataset GUI-Actor WebDataset A WebDataset format version of the GUI-Actor dataset for training vision-language models on GUI interaction tasks. Usage import webdataset as wds # Load the dataset dataset = wds.WebDataset("path/to/shards-*.tar") dataset = dataset.decode("pilrgb").to_tuple("jpg", "json") for image, metadata in dataset: # Process image and metadata pass Citation Please cite the original GUI-Actor paper if you use this dataset in your research. imagetext-generation1M<n<10M1 likes80 downloads1y agoHugging Face14Guilherme34 /anthropic-Awareness-interview anthropic-Awareness-interview This dataset contains full transcripts of user research interviews where an AI assistant (Claude) interviews people about how they use AI in their work and how they feel about that collaboration.[web:1] Each example includes a long meta-cognitive system prompt plus a complete back-and-forth conversation. Dataset overview Domain: Human–AI interaction in professional and day-to-day work. Format: Multi-turn chat logs with explicit roles. Scale:… See the full description on the dataset page: https://huggingface.co/datasets/Guilherme34/anthropic-Awareness-interview.texttext-generation1K<n<10K2 likes70 downloads10mo agoHugging Face15HeinKoZin /Sora-Ecommerce-Guide Sora Ecommerce Guide Dataset This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks. Splits train: 9 samples test: 2 samples Features instruction: System/task instruction context. input: The prompt, question, or user query. output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.textquestion-answeringn<1K0 likes63 downloads4d agoHugging Face16georgeqiao12138 /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K… See the full description on the dataset page: https://huggingface.co/datasets/georgeqiao12138/guidelines.texttext-generation10K<n<100K0 likes58 downloads6d agoHugging Face17guicybercode /iceland-tech-christian-ethics-prompts Fictional Icelandic Landscapes, Technology and Christian Ethics Prompts This microdataset contains 24 original discussion prompts arranged as 12 parallel pt-BR/English pairs. Each explicitly fictional scenario combines a landscape motif inspired by Iceland, a technology-governance dilemma, and concepts that may be explored through Christian ethics. The records do not describe real Icelandic institutions, policies, communities, or practices, and they do not claim that Christians… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/iceland-tech-christian-ethics-prompts.texttext-generationn<1K0 likes57 downloads27d agoHugging Face18ArunKr /gui_grounding_dataset-1k Supported Tasks Natural Language → GUI Action Grounding Convert user instructions into JSON action objects. Instruction Following Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”). Multi-step UI Automation Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot). Languages English (en) Generated with simple variations (synonyms, phrasings). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-1k.texttext-generation1K<n<10K0 likes56 downloads1y agoHugging Face19PJMixers /epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining. texttext-generation10K<n<100K0 likes48 downloads3y agoHugging Face20guicybercode /br-sovereign-llm-corpus BR Sovereign LLM Corpus Status This public repository is an audited corpus protocol and initial validation snapshot. Public release 0.1.1 contains only a small, explicitly identified corpus sample for pipeline verification. It is not a target-scale pretraining corpus and does not support model-quality claims. The scientific corpus remains under construction. Aggregate counts, source shares, and token counts are not reported until a content-addressed snapshot… See the full description on the dataset page: https://huggingface.co/datasets/guicybercode/br-sovereign-llm-corpus.texttext-generationn<1K0 likes40 downloads28d agoHugging Face21ArunKr /gui_grounding_dataset-100 Supported Tasks Natural Language → GUI Action Grounding Convert user instructions into JSON action objects. Instruction Following Models learn to interpret varying natural language formulations (e.g., “press submit” vs “click the submit button”). Multi-step UI Automation Some samples involve sequences of actions (e.g., open site → type → press Enter → screenshot). Languages English (en) Generated with simple variations (synonyms, phrasings). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ArunKr/gui_grounding_dataset-100.texttext-generationn<1K0 likes38 downloads1y agoHugging Face22WassimLab /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/WassimLab/guidelines.texttext-generation10K<n<100K0 likes35 downloads5mo agoHugging Face23guinansu /paragen-security-sft-alpaca paragen-security-sft-alpaca Alpaca-format instruction-tuning data used to train the Vanilla security baseline (and as the source for the tokenized multi-stream cache used to train the Stream(Ours) security checkpoint) in the paragen_llm security/prompt-injection-robustness experiments (Table 3: TensorTrust, Gandalf, Purple, RuLES, StruQ, NESSiE, IFEval). Format: JSONL, one object per line, fields instruction / input / output (standard Alpaca schema). Size: 48,538 examples.… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/paragen-security-sft-alpaca.texttext-generation10K<n<100K0 likes35 downloads10d agoHugging Face24Billy6310 /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/Billy6310/guidelines.texttext-generation10K<n<100K0 likes34 downloads6mo agoHugging Face25smolify /smolified-bengali-local-food-guide 🤏 smolified-bengali-local-food-guide Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-bengali-local-food-guide. 📦 Asset Details Origin: Smolify Foundry (Job ID: 638d3b25) Records: 1050 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes33 downloads6mo agoHugging Face26Remixonwin /prepware_study_guide-dataset Prepware_Study_Guide Dataset Generated by DocParserEngine. Field Value Documents 1 Records 1 Schema full Usage from datasets import load_dataset ds = load_dataset("Remixonwin/prepware_study_guide-dataset") imagetext-generationn<1K0 likes30 downloads7mo agoHugging Face27mmrech /guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/guidelines.texttext-generation10K<n<100K0 likes30 downloads6mo agoHugging Face28guigux /hulk_dataset_0.1This dataset is AFAIK (12 january 2024) the biggest ready to use open source dataset to finetune LLMs. It contains more than 3.8 million chat samples. Its a collection of multiple different datasets. Some of them have been built using GPT4 or using scraped data. Here is the list: gathnex/Gath_baize teknium/openhermes nomic-ai/gpt4all-j-prompt-generations teknium/dataforge-economics Anthropic/hh-rlhf: we kept only the selected prompts teknium1_GPTeacher_codegen… See the full description on the dataset page: https://huggingface.co/datasets/guigux/hulk_dataset_0.1.texttext-generation1M<n<10M3 likes29 downloads3y agoHugging Face29minsu /epfl-llm_guidelines 🎉 NEW DROP 🎉 PubMed Guidelines We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas! Clinical Guidelines The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/minsu/epfl-llm_guidelines.texttext-generation10K<n<100K0 likes27 downloads7mo agoHugging Face30rebeccazzzz /gui-vs-cli GUI-vs-CLI: A Unified Benchmark This dataset contains task descriptions and verification specifications for 440 desktop software tasks from the GUI-vs-CLI benchmark. The Hugging Face dataset is intended for browsing and lightweight programmatic access to task descriptions. Full runnable task assets, environment files, and execution code are maintained in the GitHub repository. Files data/tasks.jsonl: one task per line. Fields Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/rebeccazzzz/gui-vs-cli.texttext-generationn<1K0 likes26 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.