CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01baridhi /SEC-EDGARDatamule, Teraflop AI, and Eventual collaborated to release the SEC-EDGAR dataset. The dataset contains 590 gbs of data, spanning 8 million samples and 43 billion tokens from all major filings in the SEC EDGAR database. The bulk data was collected using datamule-python library and the official datamule api created by John Friedman. The datamule Python library is a package for collecting, manipulating, and processing the SEC Edgar data at scale. Datamule provides a simple open-source api… See the full description on the dataset page: https://huggingface.co/datasets/baridhi/SEC-EDGAR.texttext-generation1M<n<10M0 likes1.9k downloads5mo agoHugging Face02barc0 /200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds. We generate the dataset with the following steps and two approaches: Generate ~110k descriptions by GPT4o. Approach 1: Generate ~110k codes follow each description by GPT4o-mini. Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions. Run the ~220k codes and do auto-filtering. Get the final ~200k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M11 likes652 downloads2y agoHugging Face03barissozudogru /swe-bench-mini SWE-bench-mini 34 self-contained bug-fix tasks in the SWE-bench format — a small repository snapshot carrying a defect, a test that fails because of it, and a gold patch that fixes it (difficulty mix: 12 easy / 19 medium / 3 hard, author estimate). Built for the swe_bench_mini agent and the make demo-swe-mini evaluator in adk-agent-playground, to demonstrate the framework's range on code-modification and to exercise the CaMeL filesystem-capability gate. A second harder config… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/swe-bench-mini.texttext-generationn<1K1 likes437 downloads2mo agoHugging Face04voidful /barbet-sft Barbet SFT 以臺灣繁體中文為中心的 1,000,000 筆結構化決策 SFT。 train 980,000、validation 10,000、test 10,000。 from datasets import load_dataset dataset = load_dataset("voidful/barbet-sft", split="train", streaming=True) messages = next(iter(dataset))["messages"] messages 可直接用作 system / user / assistant 訓練資料。 預設只讀固定 schema Parquet;canonical 和品質 JSON 不參與資料欄位推斷。 每筆 request 明示輸出格式,JSON、FUNCTION_CALL、COMPACT_TAGS、KEY_VALUE 共用相同 canonical 目標,沒有額外自由形式思考過程。 品質與臺灣用語 全部 1,000,000… See the full description on the dataset page: https://huggingface.co/datasets/voidful/barbet-sft.texttext-generation1M<n<10M0 likes325 downloads17d agoHugging Face05AUEB-NLP /greek-bar-bench Dataset Card for GreekBarBench 🇬🇷🏛️⚖️ GreekBarBench is a benchmark designed to evaluate LLMs on challenging legal reasoning questions across five different legal areas from the Greek Bar exams, requiring citations to statutory articles and case facts. This repository hosts two related benchmarks: Benchmark Subsets Task GreekBarBench (GBB) greekbarbench, gbb-jme Free-text legal reasoning with citations, and LLM-judge meta-evaluation GreekBarRetrieval (GBR)… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/greek-bar-bench.tabularquestion-answering1K<n<10K5 likes203 downloads3d agoHugging Face06barc0 /100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4o-mini. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes160 downloads2y agoHugging Face07PiotrSty /kultura-bez-barier-pl Kultura bez Barier PL Polish text extracted from openly published accessibility guides by Fundacja Kultury bez Barier. Input publications: 12 Retained records: 9 Extracted pages: 316 Tokens: 217,932 (cl100k_base proxy) License evidence: catalogue-level CC BY-SA 3.0 declaration Author coverage: 100.0% See NOTICE.md and artifacts/ for the source inventory, raw PDFs, attribution, selection decisions, QA, checksums, samples and Slayer ontology manifest. Limitations… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/kultura-bez-barier-pl.documenttext-generationn<1K0 likes159 downloads17d agoHugging Face08lemon07r /bartowski-imatrix-v5-semantic Bartowski iMatrix Calibration v5 (Semantic Chunking) A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure. Dataset Summary Metric Value Total samples 2,075 Chunking method V5-optimized semantic boundary detection Chunk size 200+ characters (no upper limit, preserves document integrity) Languages English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.texttext-generation1K<n<10K9 likes158 downloads8mo agoHugging Face09barc0 /100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes128 downloads2y agoHugging Face10Lots-of-LoRAs /task1156_bard_analogical_reasoning_tools Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1156_bard_analogical_reasoning_tools Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1156_bard_analogical_reasoning_tools.texttext-generationn<1K0 likes112 downloads2y agoHugging Face11barandinho /turkish-reasoning-distilled-sft Turkish Reasoning Distilled SFT This dataset contains Turkish reasoning SFT data for barandinho/qwen3.5-27b-tudum-dapo-50. The teacher model also received a small RL run, but this dataset is its main supervised fine-tuning data. It combines generated Turkish reasoning traces from DAPO math, WebInstruct, AceCode, and verifier-compatible IFEval-style instruction-following sources with verified teacher-SFT traces from DAPO math, OpenThoughts science, OpenThoughts code, and system-chat… See the full description on the dataset page: https://huggingface.co/datasets/barandinho/turkish-reasoning-distilled-sft.texttext-generation1M<n<10M0 likes110 downloads4mo agoHugging Face12Danielbrdz /Barcenas-Personas-Mexico Barcenas Personas México Dataset sintético sociodemográfico de mayor granularidad geográfica publicado para México: 1,000,000 de personas, 2,478 municipios y 428,041 hogares. Un millón de personas sintéticas, estadísticamente calibradas contra fuentes oficiales (INEGI, ENOE, CONEVAL) y distribuidas en los 2,478 municipios de las 32 entidades federativas de México. Cada persona cuenta con una ficha demográfica completa y 8 facetas narrativas en español (~1,100 palabras por… See the full description on the dataset page: https://huggingface.co/datasets/Danielbrdz/Barcenas-Personas-Mexico.tabulartext-generation1M<n<10M0 likes95 downloads22d agoHugging Face13Lots-of-LoRAs /task1157_bard_analogical_reasoning_rooms_for_containers Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1157_bard_analogical_reasoning_rooms_for_containers Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1157_bard_analogical_reasoning_rooms_for_containers.texttext-generationn<1K0 likes86 downloads2y agoHugging Face14Lots-of-LoRAs /task1158_bard_analogical_reasoning_manipulating_items Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1158_bard_analogical_reasoning_manipulating_items Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1158_bard_analogical_reasoning_manipulating_items.texttext-generationn<1K0 likes82 downloads2y agoHugging Face15Lots-of-LoRAs /task1154_bard_analogical_reasoning_travel Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1154_bard_analogical_reasoning_travel Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1154_bard_analogical_reasoning_travel.texttext-generationn<1K0 likes80 downloads2y agoHugging Face16Lots-of-LoRAs /task1319_country_by_barcode_prefix Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1319_country_by_barcode_prefix Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1319_country_by_barcode_prefix.texttext-generationn<1K0 likes79 downloads2y agoHugging Face17Lots-of-LoRAs /task1152_bard_analogical_reasoning_causation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1152_bard_analogical_reasoning_causation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1152_bard_analogical_reasoning_causation.texttext-generationn<1K0 likes76 downloads2y agoHugging Face18hugfaceguy0001 /retarded_bar 弱智吧笑话数据集 弱智吧是百度贴吧中的一个非常受欢迎的论坛,以创作短小精悍的冷笑话而闻名。这些笑话通常采用双关语、不寻常的断句、不合理的逻辑等创作手法。即使是目前最先进的语言模型,也难以完全理解弱智吧的笑话。 弱智吧 我从互联网上收集了一些弱智吧的笑话,共100条,其中45条是陈述句,55条是问句。我结合人工和语言模型对这些笑话进行了一些解析,并制作了这个小型数据集。 陈述句笑话 陈述句笑话通常以句号结尾,不容易被语言模型误解为正常的问题。 例如:“出人头地常年盛产人头。” 问句笑话 问句笑话具有一定的迷惑性,可能会导致语言模型无法判断它们是正常的问题还是开玩笑。 例如:“蓝牙耳机坏了,应该找牙科医生还是耳科医生?” 文件格式 本数据集包括两个部分。 retarded_bar.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/hugfaceguy0001/retarded_bar.texttext-generationn<1K60 likes72 downloads3y agoHugging Face19mzhaoshuai /llama3-ultrafeedback-bertscore-bart-large-mnli RefAlign: LLM Alignment Dataset This dataset is used in the paper Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data. Code: https://github.com/mzhaoshuai/RefAlign This dataset is modified from https://huggingface.co/datasets/princeton-nlp/llama3-ultrafeedback. We use the BERTScore to choose the chosen and rejected responses. Item with key ['Llama3.3-70B-Inst-Awq'] is the reference answers generated by… See the full description on the dataset page: https://huggingface.co/datasets/mzhaoshuai/llama3-ultrafeedback-bertscore-bart-large-mnli.texttext-generation10K<n<100K0 likes68 downloads11mo agoHugging Face20barozp /opus-reasoning-distill-train Claude Opus Reasoning Distillation Dataset Claude Opus 4.6/4.7 reasoning traces, formatted for SFT fine-tuning. This repo is the train split. Dataset size Split Examples train 14,250 this repo validation 750 opus-reasoning-distill-validation total 15,000 95 / 5 split Purpose Distill Claude Opus's structured reasoning style into Qwen3.6-27B (Bonsai base model). Data Sources Source Teacher Model… See the full description on the dataset page: https://huggingface.co/datasets/barozp/opus-reasoning-distill-train.texttext-generation10K<n<100K1 likes65 downloads2mo agoHugging Face21lianghsun /tw-bar-examination-2020-chat Dataset Card for tw-bar-examination-2020-chat tw-bar-examination-2020-chat 是一個中華民國 2020 年律師考試選擇題之 Alpaca 格式微調資料集,合計 299 題(train 269、test 30)。每題包含統一提示語、題目與四個選項,以及正確答案字母,適用於微調繁體中文語言模型於台灣法律選擇題作答任務。 Dataset Details Dataset Description 本資料集源自 Jamie0510/taiwan-law-exam 中之 2020 年律師考試題目,整合其四大類科後進行後處理:去除欄位缺失之題目,並統一轉為 Alpaca 三欄格式(instruction / input / output)。每題之 instruction 欄為固定提示語「請在下列的單一選擇題中,選出正確的答案,並且只回答 A, B, C, D 其中一個字代表正確答案」。 本資料集作為 SFT 訓練素材設計,建議與… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-bar-examination-2020-chat.textquestion-answeringn<1K3 likes58 downloads5mo agoHugging Face22barbarabhb /nl2sh-chatter-robustness Chatter / robustness pairs for NL->shell models 246 hand-written (natural language, shell command) pairs teaching the "boring reflex": greetings, small talk, identity questions and nonsense input map to harmless commands (echo hello, pwd) instead of garbage or network-touching behavior. Generated by organic_augment.py (deterministic, seed 42). Used in the training pool of barbarabhb/nl2sh-qwen25-coder-1.5b-GGUF. texttext-generationn<1K0 likes54 downloads1mo agoHugging Face23TigreGotico /dicionario_barranquenho Dicionário de Barranquenho Structured lexical dataset derived from the first published dictionary of Barranquenho, a Romance contact language spoken in Barrancos, Portugal. Contains 1,680 entries with Portuguese and Spanish glosses, grammatical categories, semantic fields, source attributions, and synonym cross-references. Language Barranquenho (glottocode: barr1245; no ISO 639-3 code assigned at time of publication) is a contact language spoken in the municipality of… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/dicionario_barranquenho.documenttranslation1K<n<10K0 likes52 downloads5mo agoHugging Face24Baron-GG /BUSMThe paper is currently under review, and more detailed information will be announced later. imagetext-generation10K<n<100K1 likes47 downloads2y agoHugging Face25barozp /opus-reasoning-distill-v2-train opus-reasoning-distill-v2-train Same 14,250 prompts as barozp/opus-reasoning-distill-train (v1). Every row's reasoning trace was checked against two source datasets that host verbatim Claude Opus API traces for an overlapping set of prompts, and replaced wherever the check found v1's version to actually differ from the verbatim one — see the breakdown below for exact counts. Why this version exists While investigating a reasoning-loop issue reported against a… See the full description on the dataset page: https://huggingface.co/datasets/barozp/opus-reasoning-distill-v2-train.texttext-generation10K<n<100K0 likes44 downloads1mo agoHugging Face26Barath /genielm-ui-grounding GenieLM UI-Grounding Synthetic supervised fine-tuning data for text-based UI grounding: given a list of on-screen elements (label + pixel center) and a natural-language instruction, pick the single element to act on and emit a strict JSON action. Built for GenieLM, a macOS agent that reads the accessibility tree as text (not pixels) and lets a small LLM drive the cursor. Format Conversational SFT (messages column): {"messages": [ {"role": "system", "content":… See the full description on the dataset page: https://huggingface.co/datasets/Barath/genielm-ui-grounding.texttext-generation1K<n<10K0 likes41 downloads3mo agoHugging Face27AnvaMiba /llm-bargaining-scenarios LLM Bargaining Scenarios Commodity-bargaining scenarios used in the paper Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information (Miceli-Barone, Belle, Cohen; 2026). Each scenario describes an item, a buyer persona, a seller persona, and the reservation-price ranges from which the two agents' private reservation prices are sampled per trial. Files Two views of the same 4561 scenarios are provided: scenarios.jsonl (the… See the full description on the dataset page: https://huggingface.co/datasets/AnvaMiba/llm-bargaining-scenarios.texttext-generation1K<n<10K0 likes37 downloads4mo agoHugging Face28Lots-of-LoRAs /task1153_bard_analogical_reasoning_affordance Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1153_bard_analogical_reasoning_affordance Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1153_bard_analogical_reasoning_affordance.texttext-generation1K<n<10K0 likes32 downloads2y agoHugging Face29lemon07r /bartowski-imatrix-v3-semantic Bartowski iMatrix Calibration v3 (Semantic Chunking) A processed version of bartowski's v3 imatrix calibration data using semantic boundary detection in attempt to create coherent, non-overlapping samples. Dataset Summary Metric Value Total samples 168 Chunking method Semantic boundary detection Target chunk size ~2048 characters Languages English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese Source Data The… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v3-semantic.texttext-generationn<1K1 likes32 downloads8mo agoHugging Face30Baron-GG /USQAimagequestion-answeringn<1K1 likes27 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.