CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BAAI /IndustryCorpus_technology[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.texttext-generation10M<n<100M4 likes3.7k downloads1mo agoHugging Face02TechnoBaptist /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/TechnoBaptist/stack-v3-train.tabulartext-generation100M<n<1B0 likes1.2k downloads2mo agoHugging Face03Phase-Technologies /forge-3b-dpo-data FORGE-3B DPO Preference Data Tokenized (prompt, chosen, rejected) preference triples for DPO post-training of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2. This is data preparation output only — no model was trained to produce this. Stats Total pairs: 0 (paper target: ~200,000) Domains: 0/4 Context length: 4096 tokens (paper Appendix A.2, DPO block) Format: unpacked — one (prompt, chosen, rejected) triple per training example Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.texttext-generation100K<n<1M0 likes94 downloads3mo agoHugging Face04sapiens-technology /simple_bench 📊 Simple Bench Dataset A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.texttext-generationn<1K0 likes38 downloads5mo agoHugging Face05omniomni /omni-technology GitHub   Website   Paper (Coming Soon) Dataset Details This dataset is a combination of coding Question-Answer pairs, as well as Q&A data from StackExchange posts and their top-voted answers. This dataset contains only chat-based data. Sources This dataset was sourced from the following open-sourced datasets: Technology Replete-AI/code_bagel sahil2801/CodeAlpaca-20k donfu/oa-stackexchange texttext-generation100K<n<1M1 likes19 downloads1y agoHugging Face06Seldon-Technologies /HARD-TIMEgated HARD-TIME HARD-TIME evaluates whether video-language models can localize moments in time and avoid answers that are not supported by the video. This repository contains benchmark annotations for the evaluation tasks; it does not include source videos, transcripts, or evidence artifacts. Access to the annotations does not grant any license or reuse rights for the underlying third-party content. Configs Config Rows Description temporal_retrieval 4,812… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/HARD-TIME.tabularvisual-question-answering1K<n<10K0 likes14 downloads4mo agoHugging Face07Seldon-Technologies /golden-vault-v0gated Golden Vault Golden Vault is a curated long-video understanding dataset with dense video descriptions, audio transcripts, timestamp-grounded questions, evidence-grounded questions, multi-hop questions, and contrastive unanswerable questions. VAULT stands for Video-Audio Understanding over Long Timelines. Dataset At A Glance This release is intentionally pruned. The QA/evaluation configs are preserved in full, while broad caption configs are limited to a 500-video… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/golden-vault-v0.tabularimage-to-text1K<n<10K2 likes10 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.