CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01livebench /data_analysis Dataset Card for "livebench/data_analysis" LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties: LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/data_analysis.textn<1K8 likes7.9k downloads1y agoHugging Face02TAUR-Lab /Taur_CoT_Analysis_Project___gpt-4o-2024-08-06text10K<n<100K1 likes2.7k downloads2y agoHugging Face03aisingapore /NLU-Sentiment-Analysisgated SEA Sentiment Analysis SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese. Supported Tasks and Leaderboards SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.texttext-generation1K<n<10K0 likes2.3k downloads9mo agoHugging Face04TAUR-Lab /Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructtext10K<n<100K0 likes1.5k downloads2y agoHugging Face05TAUR-Lab /Taur_CoT_Analysis_Project___gpt-4o-mini-2024-07-18text10K<n<100K3 likes1.5k downloads2y agoHugging Face06TAUR-Lab /Taur_CoT_Analysis_Project___microsoft__Phi-3-small-8k-instructtext10K<n<100K0 likes921 downloads2y agoHugging Face07TAUR-Lab /Taur_CoT_Analysis_Project___google__gemini-1.5-flash-001text10K<n<100K0 likes826 downloads2y agoHugging Face08patched-codes /static-analysis-evalA dataset of 76 Python programs taken from real Python open source projects (top 100 on GitHub), where each program is a file that has exactly 1 vulnerability as detected by a particular static analyzer (Semgrep), used in the paper Patched MOA: optimizing inference for diverse software development tasks. OpenAI used the synth-vuln-fixes and fine-tuned a new version of gpt-4o is now the SOTA on this benchmark. More details and code is available from their repo. More details on the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/static-analysis-eval.textn<1K20 likes771 downloads1y agoHugging Face09timchen0618 /browsecomp-plus-selected-tools-analysis-v1 BrowseComp-Plus: Selected Tools Analysis Side-by-side view of selected tool calls from a reference trajectory alongside the new agent trajectory conditioned on those steps. Retrieval model: Qwen3-Embedding-8BAgent model: gpt-oss-120bRun: traj_summary_ext_selected_tools_gpt-oss-120b_seed0 Columns Column Description query_id Query identifier rationale GPT rationale for why these k steps were selected from the reference trajectory selected_indices Step indices… See the full description on the dataset page: https://huggingface.co/datasets/timchen0618/browsecomp-plus-selected-tools-analysis-v1.tabularn<1K0 likes604 downloads6mo agoHugging Face10TAUR-Lab /Taur_CoT_Analysis_Project___mistralai__Mistral-7B-Instruct-v0.3text100K<n<1M0 likes573 downloads2y agoHugging Face11NuBerea /source-analysisgated NuBerea Source Analysis Source-critical analysis of the Hebrew Bible, Septuagint, New Testament, Vulgate, and Second Temple literature. The dataset carries machine-generated source and tradition annotations at the verse level — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, the pathway of Old Testament traditions into New Testament citation) expressed as structured data — together with semantic-domain… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-analysis.tabularfeature-extraction100K<n<1M0 likes526 downloads7d agoHugging Face12hf-azure-internal /trending-models-analysishttps://github.com/pagezyhf/azure-cron/blob/main/trending_models_analysis.py text10K<n<100K3 likes506 downloads8h agoHugging Face13TAUR-Lab /Taur_CoT_Analysis_Project___google__gemini-1.5-pro-001text10K<n<100K1 likes499 downloads2y agoHugging Face14TAUR-Lab /Taur_CoT_Analysis_Project___Qwen__Qwen2-72B-Instructtext10K<n<100K0 likes488 downloads2y agoHugging Face15TAUR-Lab /Taur_CoT_Analysis_Project___claude-3-5-sonnet-20240620text10K<n<100K1 likes471 downloads2y agoHugging Face16NuBerea /translation-analysisgated NuBerea Translation Verse Texts Verse-level texts of historical Bible translations (Clementine Vulgate, Luther Bible 1545, Matthew's Bible 1537). Part of the NuBerea curated corpus estate of biblical and historical texts. Attribution Upstream Data Sources Source License Clementine Vulgate, NOCR Public Domain Luther Bible 1545, NOCR Public Domain Matthew's Bible 1537, Textus Receptus Bibles Public Domain NuBerea project. Licensed… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/translation-analysis.tabularfeature-extraction100K<n<1M0 likes453 downloads14d agoHugging Face17NuBerea /septuagint-analysisgated NuBerea Septuagint Textual Analysis Curated datasets for study of the Septuagint (the ancient Greek translation of the Hebrew Bible), part of the NuBerea corpus estate of biblical and patristic texts. It gathers Septuagint verse texts, apparatus notes, and edition-comparison material into a set of ready-to-load configurations. Attribution This dataset derives from the following upstream sources, which require attribution: Source License Rahlfs… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/septuagint-analysis.tabularfeature-extraction10K<n<100K0 likes442 downloads2mo agoHugging Face18NuBerea /pseudepigrapha-analysisgated NuBerea Pseudepigrapha Analysis Derived linguistic datasets over pseudepigraphal literature, part of the NuBerea curated corpus estate. Covers the Greek and Latin witnesses of these texts along with a multilingual view across the available witness languages. License CC BY 4.0. Attribution Source Link License NuBerea project https://huggingface.co/NuBerea CC BY 4.0 tabularfeature-extraction10K<n<100K0 likes440 downloads7d agoHugging Face19jtviegas /ticker_analysis_articlestabular10K<n<100K0 likes427 downloads20h agoHugging Face20NuBerea /lxx-analysisgated NuBerea Research: LXX Translation-Technique Noise Model Quantitative study of Septuagint translation technique: verse-by-verse measurements of where the ancient Greek translation (LXX) diverges from the Hebrew Masoretic Text, with book-level statistical summaries. The material lets researchers distinguish a translator's habitual working style — free versus literal rendering — from genuine textual anomalies worth close scholarly attention, putting on a measurable footing what LXX… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/lxx-analysis.tabularfeature-extraction10K<n<100K0 likes427 downloads7d agoHugging Face21TAUR-Lab /Taur_CoT_Analysis_Project__Symbolic_Solver_Experimenttext10K<n<100K2 likes384 downloads2y agoHugging Face22TAUR-Lab /Taur_CoT_Analysis_Project___google__gemma-2-9b-ittext10K<n<100K0 likes330 downloads2y agoHugging Face23jtviegas /ticker_analysis_pricestabular10K<n<100K0 likes316 downloads20h agoHugging Face24OALL /details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2 Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2 Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2. The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2.tabular100K<n<1M0 likes315 downloads1y agoHugging Face25gretelai /gretel-financial-risk-analysis-v1 gretelai/gretel-financial-risk-analysis-v1 This dataset contains synthetic financial risk analysis text generated by fine-tuning Phi-3-mini-128k-instruct on 14,306 SEC filings (10-K, 10-Q, and 8-K) from 2023-2024, utilizing differential privacy. It is designed for training models to extract key risk factors and generate structured summaries from financial documents while demonstrating the application of differential privacy to safeguard sensitive information. This dataset showcases… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-financial-risk-analysis-v1.tabulartext-classification1K<n<10K12 likes262 downloads2y agoHugging Face26themohal /saraiki-sentiment-analysis-datasetgatedtextn<1K0 likes222 downloads17h agoHugging Face27prithivMLmods /Gym-Exercise-Video-Analysis Gym-Exercise-Video-Analysis Gym-Exercise-Video-Analysis is a specialized multimodal video understanding dataset comprising 500 annotated gym workout and exercise clips. It is designed for fine-tuning and evaluating Video-Language Models (Video-LLMs), visual fitness coaches, and temporal exercise analysis systems. Each entry pairs exercise videos and extracted frame sequences with in-depth textual descriptions, biomechanical observations, form evaluations, and routine tracking.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gym-Exercise-Video-Analysis.imagevideo-text-to-textn<1K2 likes213 downloads28d agoHugging Face28princeton-nlp /QuRatedPajama-1B_tokens_for_analysis QuRatedPajama Paper: QuRating: Selecting High-Quality Data for Training Language Models This dataset is a 1B token subset derived from princeton-nlp/QuRatedPajama-260B, which is a subset of cerebras/SlimPajama-627B annotated by princeton-nlp/QuRater-1.3B with sequence-level quality ratings across 4 criteria: Educational Value - e.g. the text includes clear explanations, step-by-step reasoning, or questions and answers Facts & Trivia - how much factual and trivia knowledge the text… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRatedPajama-1B_tokens_for_analysis.tabular1M<n<10M6 likes208 downloads2y agoHugging Face29Akash190104 /bengali_sentiment_analysis Bengali Sentiment Analysis Context The dataset contains 3307 Negative reviews and 8500 Positive reviews collected and manually annotated from Youtube Bengali drama. Positive_Label=1 and Negative_Label=0 Acknowledgements Sazzed, Salim (2021), “Bangla ( Bengali ) sentiment analysis classification benchmark dataset corpus”, Mendeley Data, V4, doi: 10.17632/p6zc7krs37.4 texttext-classification10K<n<100K0 likes188 downloads2y agoHugging Face30slliac /isom5240-td-traffic-analysistabularn<1K0 likes175 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.