CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QuechuaBase /asr-ser-quechua-collao-embeddings ASR-SER embeddings for Quechua Collao This repository contains embeddings only. It does not contain raw audio. These embeddings were extracted for ASR-to-SER transfer experiments on the Quechua Collao emotional speech corpus. The release is intended for reproducible feature-extraction experiments and downstream analysis. Dataset contents One PyTorch tensor per utterance stored as an embedding file under embeddings/ A sanitized metadata table describing utterance… See the full description on the dataset page: https://huggingface.co/datasets/QuechuaBase/asr-ser-quechua-collao-embeddings.tabularaudio-classificationn<1K0 likes1.3k downloads2mo agoHugging Face02SEBK4C /gemma4-serving-bench-data Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data Test data, charts, and the running research log from an autonomous research loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop summarizes findings, proposes a goal, tests it end-to-end, documents success or failure, and publishes here + to GitHub. Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.imagen<1K0 likes1.2k downloads3mo agoHugging Face03ServiceNow /drbench DRBench: A Realistic Benchmark for Enterprise Deep Research 📄 Paper | 💻 GitHub | 💬 Discord DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent's ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst. ✨ Key Features 🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/drbench.documentquestion-answeringn<1K5 likes831 downloads6mo agoHugging Face04Teejeigh /raw_friends_series_transcriptRaw transcript from friends tv series, chunk into ~1000 token lines. Text tagged by character. language: - en size_categories: - n<1K textn<1K2 likes809 downloads3y agoHugging Face05ServiceNow-AI /AgentJudgeBench AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling A benchmark for systematically evaluating how reliably LLM judges assess agentic tool-calling workflows across structured, dependency-driven tasks. Why this benchmark? AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.tabularquestion-answering100K<n<1M0 likes461 downloads25d agoHugging Face06KrossKinetic /SP500-Financial-News-Articles-Time-SeriesTextual Time Series Dataset for finetuning / pretraining. Json version of original dataset. Original Dataset : https://www.kaggle.com/datasets/skywalker290/financial-news-article-and-stock-trend-dataset?select=stock_data_articles.csv text1K<n<10K6 likes361 downloads2y agoHugging Face07ServiceNow /Dr-CiK Dr-CiK: A Testbed for Foresight-Driven Agents Dr-CiK is a benchmark for evaluating whether agents can retrieve forecasting-relevant context from a noisy document corpus, filter out distractors, distill the retrieved context into forecast-useful evidence, and produce forecasts grounded in that evidence. Real-world time-series forecasting often depends not only on historical observations but also on external context that must be actively discovered from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.tabulartime-series-forecasting10K<n<100K3 likes333 downloads3mo agoHugging Face08amuzetnoM /uranium-series ☢️ Uranium Research Series By Artifact Virtual — Ali A. Shakil & Ava Shakil "Hardware is algorithmic. Binary weights learn. Gradients are optional. Self-conditioning is the universal failure mode." The Uranium Series is a sequence of research papers exploring the fundamental physics of neural computation — from treating GPUs as algorithmic substrates, through binary-weight learning without gradients, to the discovery that autoregressive models inevitably poison themselves through… See the full description on the dataset page: https://huggingface.co/datasets/amuzetnoM/uranium-series.textn<1K0 likes297 downloads6mo agoHugging Face09sergiopaniego /pelican-svg-drawings Pelican SVG: frontier model drawings, scored 139 SVGs produced by seven frontier models answering Simon Willison's prompt, "generate an SVG of a pelican riding a bicycle", each with the score the pelican_svg_env OpenEnv environment gave it. Four of the seven are open weights and three are closed. Simon has run that prompt against nearly every model release since early 2025, but the results live as embedded images across 129 blog posts and scattered gists. This dataset exists… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/pelican-svg-drawings.imageimage-to-textn<1K1 likes240 downloads2mo agoHugging Face10allenai /Sera-4.5A-Full-T1This dataset contains 72118 trajectories. Data was generated from the first rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes three SVG runs per function. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path: File path to the sampled function problem_statement: Problem statement provided to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Full-T1.text10K<n<100K2 likes234 downloads7mo agoHugging Face11allenai /SERA-4.6-Lite-Best-SubsetThis dataset contains 47498 high-quality samples from Sera-4.6-Lite-T1 and Sera-4.6-Lite-T2. Training leads to a open-source SoTA 50.67% +/- 1.86% performance on SWE-Bench Verified at 32K context length, outperforming Devstral-Small-2 and GLM-4.5-Air. Method:We keep only model-submitted train samples and then filter by truncation ratio at 32K tokens until a threshold ratio of 0.88. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase… See the full description on the dataset page: https://huggingface.co/datasets/allenai/SERA-4.6-Lite-Best-Subset.text10K<n<100K3 likes210 downloads7mo agoHugging Face12serval-uni-lu /orc-bench ORC-bench Task 1: Topological Path Finding Task 2: Topological Connectivity Task 3: Linear Power Flow Task 4: Contingency Analysis Task 5: Power Grid ControlTask 6: Power Flow Optimization Task 1: Topological Path Finding Problem Formulation This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.textquestion-answering10K<n<100K0 likes205 downloads5mo agoHugging Face13allenai /Sera-4.6-Lite-T2This dataset contains 36083 trajectories. A 25000 subset was used to train SERA-32B. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.6 as teacher. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path: File path to the sampled function problem_statement: Problem statement provided to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.6-Lite-T2.text10K<n<100K11 likes201 downloads7mo agoHugging Face14Chamaka8 /SerendibLLM-PoemSong-Dataset SerendibLLM Poem and Song Dataset Sinhala poem and song lyrics dataset for fine-tuning Sinhala LLMs on creative generation. Built as part of the Serendib LLM Honours project (UCLan 2025-2026). 19,184 instruction-response entries Sources: kawmuthu.blogspot.com (poems) + lyrics-lk.com (songs) 6 instruction variants per entry (Sinhala + English prompts) 8 themes: love, nature, sadness, joy, spring, country, religion, general texttext-generation10K<n<100K0 likes179 downloads6mo agoHugging Face15serendipitylxd /NavLock-BreachSynth NavLock-BreachSynth NavLock-BreachSynth is a fixed, test-only multimodal counterfactual benchmark for navigation-lock warning reliability assessment. It pairs each generated hazard scene with its factual source observation in NavLock-HY through the stable sample_idx field. Project code is available at https://github.com/serendipitylxd/NavLock-BreachSynth. Release Scope 123 fixed test scenes and 1186 synchronized frames; eight camera images per generated frame;… See the full description on the dataset page: https://huggingface.co/datasets/serendipitylxd/NavLock-BreachSynth.textobject-detection1K<n<10K0 likes175 downloads2mo agoHugging Face16joelniklaus /online_terms_of_service Dataset Card for A Corpus for Multilingual Analysis of Online Terms of Service Dataset Summary "We present the first annotated corpus for multilingual analysis of potentially unfair clauses in online Terms of Service [=ToS]. The data set comprises a total of 100 contracts, obtained from 25 documents annotated in four different languages: English, German, Italian, and Polish. For each contract, potentially unfair clauses for the consumer are annotated, for nine different… See the full description on the dataset page: https://huggingface.co/datasets/joelniklaus/online_terms_of_service.texttext-classification10K<n<100K7 likes168 downloads4y agoHugging Face17allenai /Sera-4.5A-Full-T2This dataset contains 66337 trajectories. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes three SVG runs per function. Sera-4.5-Lite-T2 is a subset of this dataset and was used to train SERA-32B-GA. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Full-T2.text10K<n<100K3 likes166 downloads7mo agoHugging Face18spkc83 /retail-bank-servicing-alignment-sft Retail Bank Servicing Alignment SFT The training corpus for the Granite retail-bank servicing agent. It is the released tool-use SFT corpus merged with a servicing-alignment continuation curriculum that teaches multi-turn behaviours the base corpus does not: what to do when the customer says "that one", when a policy question interrupts a transfer, when the agent's own previous turn was wrong, and when the honest answer is that the agent cannot see what it was asked about. Every… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-servicing-alignment-sft.texttext-generation1K<n<10K0 likes164 downloads7d agoHugging Face19RichardSakaguchiMS /brazilian-customer-service-conversations Brazilian Customer Service Conversations Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR). De um like me apoie em manter esse dataset! Descricao Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de: Chatbots de atendimento Classificacao de intencao (intent classification) Analise de sentimento em conversas Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.texttext-classificationn<1K5 likes141 downloads10mo agoHugging Face20sermonindex /bible-reference Bible Reference Corpus Thirteen aligned reference datasets for study of the biblical text: Greek and Hebrew lexicons keyed to Strong's numbers, an interlinear word map, the critical apparatus of eight Greek editions, cross-reference and topical indexes, and geolocated places. Published by SermonIndex. Everything in this repository is public domain or CC BY 4.0. Sources with share-alike terms are kept in a separate repository, sermonindex/bible-reference-sa, so that a share-alike… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-reference.tabulartext-retrieval100K<n<1M0 likes135 downloads15d agoHugging Face21sergiopaniego /requests-pr-diff requests-pr-diff Generated by Repo2RLEnv. 💡 Browse this dataset in your browser — click the badge above or open HuggingFaceH4/harbor-visualiser to inspect every task's spec, instruction, oracle patch, test script, and Dockerfile. Source repo: psf/requests Pipeline: pr_diff Tasks: 47 Visibility: public Spec: Harbor task format with [metadata.repo2env] extension Reward kinds This dataset emits diff_similarity rewards. Each task ships an oracle diff at… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/requests-pr-diff.textn<1K0 likes134 downloads4mo agoHugging Face22serein356 /ConflictGUI ConflictGUI ConflictGUI is a benchmark for evaluating conflict awareness in GUI agents. Dataset Splits Split Feasible Conflict1 Conflict2 Total Calibration 564 300 300 1,164 Test 1,800 822 874 3,496 Sources ConflictGUI is constructed from AMEX, AndroidControl, and AITZ. The conflict instructions and labels are newly annotated, while screenshots and original instructions remain subject to their respective upstream terms.… See the full description on the dataset page: https://huggingface.co/datasets/serein356/ConflictGUI.imageimage-text-to-text1K<n<10K0 likes120 downloads23d agoHugging Face23allenai /Sera-4.6-Lite-T1This dataset contains 36825 trajectories. Data was generated from the first rollout of SVG on 121 SWE-smith codebases using GLM-4.6 as teacher. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path: File path to the sampled function problem_statement: Problem statement provided to the model target_patch: Ground truth patch… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.6-Lite-T1.text10K<n<100K2 likes119 downloads7mo agoHugging Face24sermonindex /bible The Bible in 1,004 Languages 14,497,397 verses across 1,253 translations in 1,004 languages, every verse keyed to the same chapter-and-verse address so that any two languages can be aligned by joining on book, chapter and verse. The Bible is the most widely translated text in existence, and for several hundred of the languages here it is the largest — sometimes the only — substantial digitised text. That makes this corpus unusually useful for low-resource machine translation… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible.tabulartext-generation10M<n<100M2 likes119 downloads15d agoHugging Face25ServiceNow-AI /eva-bench EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents EVA-Bench is an end-to-end evaluation framework for conversational voice agents that orchestrates bot-to-bot audio conversations and scores them on both task accuracy and interaction experience. About No existing benchmark jointly addresses the two core evaluation challenges for voice agents: generating realistic simulated conversations, and measuring quality across the full scope of… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/eva-bench.texttext-generationn<1K25 likes104 downloads4mo agoHugging Face26allenai /Sera-4.5A-Lite-T2This dataset contains 35615 trajectories. Data was generated from the second rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes one SVG run per function. 16000 samples from the dataset were used to train SERA-32B-GA. Sera-4.5-Full-T2 is a superset of this dataset with three SVG runs per function. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Lite-T2.text10K<n<100K4 likes92 downloads7mo agoHugging Face27allenai /Sera-4.5A-Lite-T1This dataset contains 36607 trajectories. Data was generated from the first rollout of SVG on 121 SWE-smith codebases using GLM-4.5-Air as teacher and includes one SVG runs per function. Schema: messages: Generated trajectory instance_id: ID of trajectory rollout_patch: Created patch to the codebase from the current trajectory func_name: Name of function sampled from codebase to start the pipeline func_path: File path to the sampled function problem_statement: Problem statement provided to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Sera-4.5A-Lite-T1.text10K<n<100K4 likes90 downloads7mo agoHugging Face28chenlei123 /customer-service Customer Service Conversations Dataset This dataset contains 100 realistic customer service conversations between customers and support agents. Each dialogue is 11 turns long and covers a variety of common issues such as late deliveries, billing errors, account problems, and more. It is ideal for training and evaluating AI assistants, chatbots, and customer support models. Dataset Structure Each conversation is stored as a JSON object with the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/chenlei123/customer-service.textn<1K0 likes90 downloads3mo agoHugging Face29declip /Minecraft-Server-Chat Minecraft Server Chat Important Info: This dataset contains swears. I filtered out as much racism as possible. People who were racist were banned from the server. I am not affiliated with the server in any way. A collection of 2,000,000 messages said across two years in a minecraft server. The minecraft semi-anarchy server logged all of its messages to discord between 2020 and 2023. I downloaded all of them and made them into a json in chronological order. I also cleaned the… See the full description on the dataset page: https://huggingface.co/datasets/declip/Minecraft-Server-Chat.text1M<n<10M8 likes78 downloads3y agoHugging Face30AsamAce /pharma-serialized-events ZigoTrace Pharma — Serialized Events (synthetic) Feature vectors extracted from a synthetic DSCSA-style serialized medicine supply chain, generated by packages/intelligence/src/synthetic.ts in the zigo-pharma engine and exported via hf/generate_fixtures.mjs. Used to train and validate the diversion-detection model (zigotrace/pharma-authenticity-model). ⚠️ Synthetic data notice This is entirely synthetic — a seeded generator (generateChain), not real distributor or… See the full description on the dataset page: https://huggingface.co/datasets/AsamAce/pharma-serialized-events.texttabular-classificationn<1K2 likes72 downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.