CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DarthJudie /LSAT_Questionstext1K<n<10K1 likes2.3k downloads4y agoHugging Face02garak-llm /dart-20250811text10K<n<100K1 likes2k downloads1y agoHugging Face03DarthJudie /LSAT_Set_1text10K<n<100K0 likes487 downloads4y agoHugging Face04nyu-dice-lab /lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private Dataset Card for Evaluation run of hkust-nlp/dart-math-llama3-8b-prop2diff Dataset automatically created during the evaluation run of model hkust-nlp/dart-math-llama3-8b-prop2diff The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-hkust-nlp-dart-math-llama3-8b-prop2diff-private.tabular100K<n<1M0 likes480 downloads2y agoHugging Face05DarthJudie /LSATQuestions3text1K<n<10K0 likes343 downloads4y agoHugging Face06Darth-Hidious /wams2026-rhea-dilutiondocumentn<1K0 likes208 downloads5mo agoHugging Face07chaannwooff /Dartdoc Dartdoc - 한국 금융공시 텍스트 데이터셋 한국 금융감독원 전자공시시스템(DART) OpenAPI를 통해 수집한 한국어 LLM 학습용 데이터셋입니다. 사업보고서, 증권신고서 등 공시 문서에서 고품질 한국어 텍스트를 추출하였습니다. 데이터셋 개요 항목 내용 언어 한국어 (ko) 수집 기간 2020년 ~ 2025년 총 레코드 수 256,548건 총 텍스트 약 4.6억 자 평균 청크 길이 약 1,794자 출처 금융감독원 DART OpenAPI 수집 대상 공시 유형 코드 대상 문서 필터 조건 정기공시 A 사업보고서 반기/분기보고서 제외 발행공시 C 증권신고서 정정신고서·집합투자 제외 추출 섹션 문서 전체가 아닌 품질이 높은 본문 섹션만 추출합니다. 섹션 내용 II 사업의 내용 IV 이사의 경영진단 및… See the full description on the dataset page: https://huggingface.co/datasets/chaannwooff/Dartdoc.texttext-generation100K<n<1M1 likes152 downloads5mo agoHugging Face08Darth-Vaderr /English-German Dataset Description: Dataset Name: English-German Translation Pairs for Machine Learning. Note: This dataset mainly focuses on formal conversation. Dataset Overview: This dataset is a meticulously curated collection of English-German translation pairs, designed specifically for training machine learning models, particularly those focused on machine translation tasks. With a total of 48,400,531 translation pairs, this dataset offers an extensive and diverse set of examples that can… See the full description on the dataset page: https://huggingface.co/datasets/Darth-Vaderr/English-German.texttranslation10M<n<100M5 likes137 downloads2y agoHugging Face09kevin0720 /SWE-bench-Darttextn<1K1 likes127 downloads1y agoHugging Face10DarthJudie /Geometrytext1K<n<10K2 likes77 downloads3y agoHugging Face11alibouhrouche /DARTS-datasettextn<1K0 likes54 downloads26d agoHugging Face12DarthVaderSenior /PreferenceHack PreferenceHack PreferenceHack is a paired preference benchmark for evaluating reward models on reward-hacking-style behaviors, released with the code for Activation Reward Models for Few-Shot Model Alignment. Repository: https://github.com/SKYWALKERRAY/activation-reward-models Splits helpful: 1,000 paired text examples focused on helpfulness and safety-oriented judgments. length: 1,000 paired text examples where length or verbosity can be a misleading reward… See the full description on the dataset page: https://huggingface.co/datasets/DarthVaderSenior/PreferenceHack.texttext-classification1K<n<10K0 likes48 downloads3mo agoHugging Face13DarthBinks /spaceship-game-leaderboard Spaceship Game - Leaderboard This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini. Stats Entries: 2 Top Score: 600 by Pilot Last Updated: 2026-08-25 Published by: DarthBinks Format The leaderboard.json file contains an array of entries: Field Type Description score int Final game score name string Player name date string ISO 8601 timestamp waves_completed int? Number of waves completed… See the full description on the dataset page: https://huggingface.co/datasets/DarthBinks/spaceship-game-leaderboard.tabularothern<1K0 likes48 downloads1mo agoHugging Face14darthludious /adaption-preference-trace-decisions PreferenceTrace — Source Corpus and Adaption Export PreferenceTrace tests exact decision-making under competing preferences, evidence, approvals, abstention requirements, temporal/contextual precedence, and machine-readable citation contracts. Two explicit lineage artifacts File Rows Role SHA-256 preferencetrace-source-96.jsonl 96 Canonical PreferenceTrace source corpus 7a447f9bf47c3ea455ed96ec36860360aa0e7b9e2dc604450e3a1c665b52363e… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-preference-trace-decisions.texttext-generationn<1K1 likes42 downloads1mo agoHugging Face15DarthZhu /vlm-knowledge-conflicttext1K<n<10K0 likes41 downloads2y agoHugging Face16anonymous-dart-2026 /training_datasettabular10K<n<100K0 likes37 downloads5mo agoHugging Face17darthludious /adaption-goose-governance-broad-seed-v1-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-goose_governance_broad_seed_v1 (augmented) An English instruction-tuning dataset covering core aspects of governance, including political systems, public policy, international relations, and security. Prompts span multiple task types such as conceptual inquiries, comparative institutional analyses, policy trade-off assessments, and evidence synthesis. Completions provide neutral… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-goose-governance-broad-seed-v1-augmented.text10K<n<100K0 likes36 downloads9d agoHugging Face18llm-semantic-router /dart-halspans DART Hallucination Spans Dataset A synthetic hallucination detection dataset derived from DART (Data-Record to Text) structured data. Contains 2,000 samples with LLM-generated responses and span-level hallucination annotations. Dataset Description This dataset was created to augment RAGTruth for Data2txt (structured data to text) task coverage. An LLM generates both faithful and intentionally hallucinated responses from DART's structured data triples, then annotates the… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/dart-halspans.texttoken-classification1K<n<10K0 likes29 downloads9mo agoHugging Face19abhirajsinha /dart-20250811text10K<n<100K0 likes28 downloads1y agoHugging Face20DarthJudie /Algebra2text1K<n<10K1 likes25 downloads3y agoHugging Face21YongdongWang /dart_llm_tasks DART-LLM Tasks Dataset Description DART-LLM Tasks is a dataset designed for evaluating language models in robotic task planning and coordination through few-shot learning. It contains 102 natural language instructions paired with their corresponding structured task decompositions and execution plans. Dataset Structure Total examples: 102 Complexity levels: L1 (Basic): 47 examples L2 (Medium): 33 examples L3 (Complex): 22 examples Features… See the full description on the dataset page: https://huggingface.co/datasets/YongdongWang/dart_llm_tasks.textn<1K1 likes22 downloads2y agoHugging Face22Darther /RCW_2025_Positive_Query_Pairs The Washington law Benchmark (WLB) Dataset Summary The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written). The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/Darther/RCW_2025_Positive_Query_Pairs.textsentence-similarity100K<n<1M0 likes20 downloads1mo agoHugging Face23darthludious /adaption-materials-science-qa-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-materials_science_qa (augmented) This dataset comprises instruction and response pairs focused on fundamental and advanced topics in materials science. Content includes questions on material properties, synthesis methods, characterization techniques, and practical engineering applications. Each entry is formatted as a prompt and completion pair designed for training technical… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-materials-science-qa-augmented.text1K<n<10K0 likes19 downloads2d agoHugging Face24darthludious /adaption-cybergoose-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-CyberGoose (augmented) This dataset contains instruction-response pairs designed to evaluate deterministic reasoning in cybersecurity governance and policy compliance. Each prompt presents a self-contained fictional policy, an operational scenario, and a single proposed action requiring context-based rule interpretation. Completions consist strictly of a binary label classifying the… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-cybergoose-augmented.text10K<n<100K0 likes19 downloads2d agoHugging Face25DarthJudie /CollAdmtextn<1K0 likes18 downloads3y agoHugging Face26darthludious /adaption-lord-vader-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-Lord Vader (augmented) This dataset consists of instruction-response pairs focused on Python code diagnosis and bug repair. Each prompt provides a Python code snippet containing defects alongside context such as tracebacks, failing tests, or expected behavioral specifications. Responses offer concise diagnoses and executable code fixes that resolve the issues while maintaining… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-lord-vader-augmented.text10K<n<100K0 likes18 downloads2d agoHugging Face27darthludious /adaption-minimal-diff-proofreading-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-minimal_diff_proofreading (augmented) This dataset consists of instruction-response pairs designed for minimal-difference proofreading across multiple English text formats, including business reports, technical documents, and casual messages. Prompts present text containing objective spelling, punctuation, grammar, and syntax errors, alongside error-free passages. Responses provide… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-minimal-diff-proofreading-augmented.text1K<n<10K0 likes18 downloads2d agoHugging Face28darthludious /adaption-better-call-saul-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-Better Call Saul (augmented) This dataset consists of instruction-response pairs designed for evaluating contract law analysis and text-grounded legal reasoning. Prompts feature self-contained contractual excerpts from commercial, employment, licensing, and lease agreements alongside analytical tasks like clause interpretation, conflict detection, and minimal text repair.… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-better-call-saul-augmented.text1K<n<10K0 likes18 downloads2d agoHugging Face29darthludious /adaption-precision-agriculture-qa-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-precision_agriculture_qa (augmented) This dataset consists of instruction and response pairs focused on precision agriculture technologies and methodologies. Queries cover topics such as soil sensor integration, variable rate application, yield mapping, and data-driven farm management. Each response provides detailed technical explanations and operational guidance for optimizing… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-precision-agriculture-qa-augmented.text10K<n<100K0 likes17 downloads2d agoHugging Face30darthludious /adaption-canuckese-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-Canuckese (augmented) This dataset comprises English instruction-response pairs curated as source material for Canadian English localization workflows. Prompts cover everyday scenarios involving currency, measurements, spelling, institutions, geography, and regional terminology. The completions provide standard, natural English responses that leave clear opportunities for downstream… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-canuckese-augmented.text1K<n<10K0 likes17 downloads2d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.