CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01csoai /gspc-custody-disclosure GSPC — custody disclosure facts (CustodyFacts) SWIFT census (live): https://councilof.ai/api/swift XRPL reader (live): https://councilof.ai/api/xrpl MEASURED financial/domain axis (named-string presence on retrieved pages over live XRPL reader-16, n=16). Not a model leaderboard. No accuracy, no fleet, no leader. Live status is the custody-disclosure row on GET https://councilof.ai/api/gspc. Not a certificate. Tokenisation evidence question: What can an outsider verify after… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-custody-disclosure.tabularothern<1K0 likes1.2k downloads1d agoHugging Face02Anthropic /discrim-eval Dataset Card for Discrim-Eval Dataset Summary The data contains a diverse set of prompts covering 70 hypothetical decision scenarios, ranging from approving a loan to providing press credentials. Each prompt instructs the model to make a binary decision (yes/no) about a particular person described in the prompt. Each person is described in terms of three demographic attributes: age (ranging from 20 to 100 in increments of 10), gender (male, female, non-binary) , and race… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/discrim-eval.tabularquestion-answering10K<n<100K60 likes1.1k downloads3y agoHugging Face03jpwahle /dblp-discovery-dataset Dataset Card for DBLP Discovery Dataset (D3) Dataset Summary DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.tabularother1M<n<10M4 likes477 downloads1y agoHugging Face04SaisExperiments /Discord-Unveiled-Compressed .hf-sanitized.hf-sanitized-uCWd6SwyNH8FCkRETeRYS .container { --bg-primary: #0d0511; --bg-secondary: #1a0f1f; --bg-tertiary: #2d1b35; --bg-card: #3d2847; --text-primary: #fef7ff; --text-secondary: #f0d9ff; --text-muted: #c084fc; --pink-soft: #fce7f3; --pink-medium: #f9a8d4; --pink-bright: #ec4899; --pink-hot: #e91e63; --pink-neon: #ff1493; --purple-soft: #e879f9; --purple-bright: #c026d3; --purple-deep: #7c3aed; --border-glow: #f472b6; --shadow-pink: rgba(244, 114, 182, 0.4);… See the full description on the dataset page: https://huggingface.co/datasets/SaisExperiments/Discord-Unveiled-Compressed.tabularn<1K28 likes285 downloads1y agoHugging Face05davidkling /hf-coding-tools-traces-discovery HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,022 query → response turns total (≈18,044 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-discovery.tabularn<1K1 likes196 downloads4mo agoHugging Face06ProDem /DiscreteLatentReasoning-MathProcessed thinking block outputs of OpenMathInstruct-2 by Qwen3.8-27B The model was instructed to solve the dataset problems with a "thinking block" CoT style, with a prompt given specifically for this format. The output reasoning chain was then processed by Qwen3.8-27B to 3 discrete candidates. tabular100K<n<1M0 likes87 downloads8d agoHugging Face07SINAI /ALIA-es-discriminative-hate-speech Dataset Introduction The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE)[1]. The release contains: 228,708 instances Spanish comments from YouTube and TikTok Per-expert predictions and explanations from three LLM experts Final fused outputs (foe_class, foe_score) for discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-hate-speech.tabulartext-classification100K<n<1M0 likes67 downloads3mo agoHugging Face08GGGgrey /university_disciplines_45kUniversity discipline dataset, including Discrete Mathematics, Introduction to Artificial Intelligence, Principles and Applications of Databases, and Computer Networks, etc. tabular10K<n<100K2 likes64 downloads9mo agoHugging Face090dinai /public-disclosures 0DIN Public GenAI Vulnerability Disclosures A weekly export of the public vulnerability disclosures published by 0DIN, the 0Day Investigative Network — Mozilla's responsible-disclosure program for GenAI security. The authoritative sources are two pages on 0din.ai: 0DIN disclosures — metadata for every published disclosure. 0DIN threat feed — the subset of disclosures publicly rendered with full content (prompts, responses, variant prompts, detection signature). Every record in… See the full description on the dataset page: https://huggingface.co/datasets/0dinai/public-disclosures.tabulartext-classificationn<1K1 likes64 downloads4mo agoHugging Face10Banodoco /discord-archive Discord Archive This is an archive of messages from the Banodoco Discord community, where technical and artistic practitioners have been discussing open source AI art for the past three years. The archive captures a long-running community record of people learning, training, evaluating, and using open source AI art models in practice. It contains discussion around model releases, workflows, tooling, troubleshooting, creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.tabulartext-generation1M<n<10M4 likes64 downloads4mo agoHugging Face11self-model /discrim-eval-templated Discrim-Eval Templated An adaptation of Anthropic/discrim-eval using strict templating for tighter experimental control. Overview This dataset contains three configs built from the same 65 base scenarios, for evaluating demographic biases in LLM decision-making. Each scenario asks a yes/no question where "yes" is favorable to the person being evaluated (e.g., approving a loan, granting a promotion). explicit (default): 520 prompts (65 scenarios × 4 races × 2… See the full description on the dataset page: https://huggingface.co/datasets/self-model/discrim-eval-templated.tabulartext-classification1K<n<10K0 likes50 downloads10d agoHugging Face12open-llm-leaderboard /Dans-DiscountModels__mistral-7b-test-merged-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/mistral-7b-test-merged Dataset automatically created during the evaluation run of model Dans-DiscountModels/mistral-7b-test-merged The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__mistral-7b-test-merged-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face13tomyimkc /repro-dynamic-regret-via-discounted-to-dynamic-reduction-with-applications-to-curved-l-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes41 downloads2mo agoHugging Face14open-llm-leaderboard /Dans-DiscountModels__Mistral-7b-v0.3-Test-E0.7-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/Mistral-7b-v0.3-Test-E0.7 Dataset automatically created during the evaluation run of model Dans-DiscountModels/Mistral-7b-v0.3-Test-E0.7 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__Mistral-7b-v0.3-Test-E0.7-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face15open-llm-leaderboard /Dans-DiscountModels__12b-mn-dans-reasoning-test-2-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/12b-mn-dans-reasoning-test-2 Dataset automatically created during the evaluation run of model Dans-DiscountModels/12b-mn-dans-reasoning-test-2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__12b-mn-dans-reasoning-test-2-details.tabular10K<n<100K0 likes31 downloads2y agoHugging Face16RyanStudio /Discord-Dialogues-Filtered Discord Dialogues Filtered I filtered the mookiezi/Discord-Dialogues dataset to obtain only high-quality english conversation examples for fine-tuning or other analytical tasks. The total data set went from 7,303,464 rows to 2,208 rows after strict filtering to remove the following: Low conversation turns or short conversations Non-English conversations Duplicate conversations Spammy/Repetitive conversations Filler words Dataset Statistics Metric Total Avg… See the full description on the dataset page: https://huggingface.co/datasets/RyanStudio/Discord-Dialogues-Filtered.tabular1K<n<10K0 likes31 downloads1mo agoHugging Face17DiscoPosse /RAGPulse RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems 🌐 Github Link | 🤗 Workload Trace | 📑 Arxiv Paper | 🤖 How to use? RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/RAGPulse.tabulartext-generation1K<n<10K0 likes30 downloads4mo agoHugging Face18CoShin /discrete_prompting_scene_graphstabular10K<n<100K0 likes29 downloads1y agoHugging Face19open-llm-leaderboard /Dans-DiscountModels__12b-mn-dans-reasoning-test-3-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/12b-mn-dans-reasoning-test-3 Dataset automatically created during the evaluation run of model Dans-DiscountModels/12b-mn-dans-reasoning-test-3 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__12b-mn-dans-reasoning-test-3-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face20fineset-io /ai-drug-discovery-papers AI for Drug Discovery Papers — FineSet A research-paper dataset on AI for Drug Discovery Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on AI for Drug Discovery Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/ai-drug-discovery-papers.tabulartext-classificationn<1K0 likes14 downloads3mo agoHugging Face21LLMTeamAkiyama /cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research データ件数: 3,733 平均トークン数: 1,193 最大トークン数: 2,489 合計トークン数: 4,453,517 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 23.2 MB 加工内容: メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。 難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.tabularquestion-answering1K<n<10K0 likes13 downloads1y agoHugging Face22YuBenjamin2004 /discrete-cot-coding-v5gated Discrete-CoT Coding Reasoning Corpus (v5) Multi-step coding reasoning traces used in the discrete chain-of-thought / text VQ-VAE study. corpus_v5.jsonl — one JSON object per generated trace. 14,249 traces over 2,509 problems: 6,562 COMPLETE, 7,687 FAILED (filter on the status field). Each COMPLETE trace has a steps list; each step carries a think (the reasoning compressed in Design B) and an intermediate answer/code. Trace-level fields include plan, final_code, final_think… See the full description on the dataset page: https://huggingface.co/datasets/YuBenjamin2004/discrete-cot-coding-v5.tabular10K<n<100K0 likes13 downloads10d agoHugging Face23open-llm-leaderboard /Dans-DiscountModels__Dans-Instruct-CoreCurriculum-12b-ChatML-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/Dans-Instruct-CoreCurriculum-12b-ChatML Dataset automatically created during the evaluation run of model Dans-DiscountModels/Dans-Instruct-CoreCurriculum-12b-ChatML The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__Dans-Instruct-CoreCurriculum-12b-ChatML-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face24open-llm-leaderboard /Dans-DiscountModels__Dans-Instruct-Mix-8b-ChatML-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/Dans-Instruct-Mix-8b-ChatML Dataset automatically created during the evaluation run of model Dans-DiscountModels/Dans-Instruct-Mix-8b-ChatML The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__Dans-Instruct-Mix-8b-ChatML-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face25open-llm-leaderboard /Dans-DiscountModels__Dans-Instruct-Mix-8b-ChatML-V0.1.1-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/Dans-Instruct-Mix-8b-ChatML-V0.1.1 Dataset automatically created during the evaluation run of model Dans-DiscountModels/Dans-Instruct-Mix-8b-ChatML-V0.1.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__Dans-Instruct-Mix-8b-ChatML-V0.1.1-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face26open-llm-leaderboard /Dans-DiscountModels__Dans-Instruct-Mix-8b-ChatML-V0.2.0-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/Dans-Instruct-Mix-8b-ChatML-V0.2.0 Dataset automatically created during the evaluation run of model Dans-DiscountModels/Dans-Instruct-Mix-8b-ChatML-V0.2.0 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__Dans-Instruct-Mix-8b-ChatML-V0.2.0-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face27open-llm-leaderboard /Dans-DiscountModels__Dans-Instruct-Mix-8b-ChatML-V0.1.0-detailsgated Dataset Card for Evaluation run of Dans-DiscountModels/Dans-Instruct-Mix-8b-ChatML-V0.1.0 Dataset automatically created during the evaluation run of model Dans-DiscountModels/Dans-Instruct-Mix-8b-ChatML-V0.1.0 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Dans-DiscountModels__Dans-Instruct-Mix-8b-ChatML-V0.1.0-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face28marbow /dimai-discord-datasettabularn<1K0 likes7 downloads2y agoHugging Face29open-llm-leaderboard /wave-on-discord__qwent-7b-detailsgated Dataset Card for Evaluation run of wave-on-discord/qwent-7b Dataset automatically created during the evaluation run of model wave-on-discord/qwent-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/wave-on-discord__qwent-7b-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face30LukaszTP /discovery-bench-simplifiedtabularn<1K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.