CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kaysss /leetcode-problem-detailed LeetCode Scraper Dataset This dataset contains information scraped from LeetCode, including problem details, metadata, and related files. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes. questions_deets.csv Contains detailed information about each problem, including problem descriptions, constraints, and examples. Columns: questionFrontendId: Unique problem ID.… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-detailed.tabulartext-classification1K<n<10K10 likes669 downloads1y agoHugging Face02hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes486 downloads2h agoHugging Face03open-llm-leaderboard-old /details_Ramikan-BR__tinyllama_PY-CODER-4bit-lora_4k-v12 Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Ramikan-BR__tinyllama_PY-CODER-4bit-lora_4k-v12.text-generation10K<n<100K0 likes369 downloads2y agoHugging Face04d0rj /povarenok_recipes_detail povarenok_recipes_detail Crawled detailed recipes from povarenok.ru website. Structure WIP imagetext-classification100K<n<1M3 likes102 downloads3y agoHugging Face05code-critic-model /critic-sft-cwm-only-detailed-prompt critic-sft-cwm-only-detailed-prompt The detailed-prompt SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Detailed-Prompt, the comparison arm of the prompt ablation in Table 4. Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The difference from critic-sft-cwm-only is the teacher prompt. Here the teacher used the detailed prompt… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only-detailed-prompt.texttext-generation1K<n<10K0 likes69 downloads23d agoHugging Face06detakarang /sql-create-context-id Overview This dataset is a fork from sql-create-context This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/detakarang/sql-create-context-id.texttext-generation10K<n<100K0 likes55 downloads3y agoHugging Face07YangL1122 /DetailBench DetailBench A Benchmark for Detail Hallucination in Long Regulatory Documents DetailBench is a benchmark for evaluating and mitigating detail hallucination in LLM outputs on long regulatory documents. Overview Large language models frequently produce detail hallucinations—subtle errors in threshold values, units, scopes, obligation levels, and conditions—when processing long regulatory documents. DetailBench provides: 322 source documents (172 real + 150 synthetic) from… See the full description on the dataset page: https://huggingface.co/datasets/YangL1122/DetailBench.text-generation10K<n<100K0 likes55 downloads7mo agoHugging Face08leafspark /DetailedReflection-Claude-v3_5-Sonnet DetailedReflection-Claude-v3_5-Sonnet This is a reflection dataset inspired by OpenAI's o1 model. It contains filtered prompts from anthracite-org/kalo-opus-instruct-22k-no-refusal. The data was generated by Claude 3.5 Sonnet with a custom system prompt and sampling parameters on OpenRouter. texttext-generationn<1K5 likes45 downloads2y agoHugging Face09TTS-AGI /balanced-emotion-dataset-majestrino-withtemporal-detailed-captions Balanced Emotion Dataset — Majestrino with Temporal Detailed Captions An emotion-balanced subset of TTS-AGI/majestrino-unified-detailed-captions-temporal. Overview Total samples: 482,594 Samples per emotion category: 12,997 Number of emotion categories: 40 Format: WebDataset (tar files with FLAC audio + JSON metadata) Number of tar files: 483 Samples per tar: ~1000 Balancing Strategy Samples were selected from the source dataset using keyword matching on… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/balanced-emotion-dataset-majestrino-withtemporal-detailed-captions.audioaudio-classification100K<n<1M0 likes35 downloads6mo agoHugging Face10xeophon /detailbench DetailBench This is the dataset for DetailBench, which answers the question: "How good are current LLMs at finding small errors, when they are not explicitly asked to do so?" Dataset Structure article_title: Name of the Wikipedia article the data is from original_text: Original excerpt from the given Wikipedia article modified_text: Modified version of the original text with a single error (one changed number) introduced original_number: The original number from the text… See the full description on the dataset page: https://huggingface.co/datasets/xeophon/detailbench.texttext-generationn<1K2 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.