CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sfc-gh-goliaro /wildchat-mixed-1k wildchat-mixed-1k Real-world chat requests for end-to-end LLM inference benchmarking in fastkernels — Scenario A, the bulk-throughput workload used to saturate continuous batching with a realistic mix of short/long prompts and short/long responses. What it's for One dataset that replaces separate prefill-heavy / balanced / decode-heavy splits: its natural length distribution puts prefill-bound and decode-bound requests in the same batch, so a single run yields a… See the full description on the dataset page: https://huggingface.co/datasets/sfc-gh-goliaro/wildchat-mixed-1k.tabulartext-generation1K<n<10K0 likes325 downloads3mo agoHugging Face02meno-sh /WildChat-curatedAs part of the lock-in hypothesis research project (Qiu et al., 2025), this dataset is transformed from raw WildChat-1M dataset (Zhao et al., 2024) into a structured analysis-ready format through: Data cleaning by deduplicating users based on IP address co-occurrence and removing templated prompts (i.e. people using the WildChat platform as a free API to do repetitive tasks). Extracting key concepts from each dialogue using a large language model (Llama-3.1-8B-Instruct), which are then… See the full description on the dataset page: https://huggingface.co/datasets/meno-sh/WildChat-curated.tabulartext-generation10M<n<100M1 likes199 downloads6mo agoHugging Face03xlr8harder /synthid-qwen3-4b-instruct-2507-wildchat Qwen3-4B SynthID three-arm corpus This export contains aligned unwatermarked, SynthID key-A, and SynthID key-B responses from Qwen/Qwen3-4B-Instruct-2507. Matched splits share prompts and request seeds across configurations; unmatched splits use mutually disjoint prompt pools. Export complete for its source work queue: true. Generation profile Model revision: cdbee75f17c01a7cc42f958dc650907174af0554 Native model dtype: bfloat16 Maximum generated tokens: 4096… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/synthid-qwen3-4b-instruct-2507-wildchat.tabulartext-generation100K<n<1M0 likes152 downloads1mo agoHugging Face04faur-ai /ro-WildChatThis dataset is a translation of allenai/WildChat using LLMic, a bilingual Romanian-English LLM. WildChat is a collection of 650K conversations between human users and ChatGPT. License: ODC-BY @inproceedings{ zhao2024wildchat, title={WildChat: 1M Chat{GPT} Interaction Logs in the Wild}, author={Wenting Zhao and Xiang Ren and Jack Hessel and Claire Cardie and Yejin Choi and Yuntian Deng}, booktitle={The Twelfth International Conference on Learning Representations}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-WildChat.tabulartext-generation100K<n<1M1 likes123 downloads1y agoHugging Face05sh0416 /wildchat-1m-tagged WildChat 1M with tagging This dataset replicates the category annotation process introduced in the paper named self-taught evaluator. This dataset additionally contains three categories, i.e., category, complexity, and length, annotated by Mistral-7B-Instruct-v0.3 for WildChat-1M sessions. I do my best to follow the technical details explained in the above paper while some difference is inevitably made due to the computational constraint. The difference and notes for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sh0416/wildchat-1m-tagged.tabulartext-generation100K<n<1M1 likes117 downloads2y agoHugging Face06open-athena /wildchat-glm53-format-completions WildChat format completions 9,975 GLM-5.3 answers across 29 parseable formats. Each answer passed its contract verifier. Failed answers were resampled with the same prompt until one passed; no semantic judge or answer repair was used. This is the final release from a 10,000-prompt run; 25 unfinished prompts were excluded. It stores answers, format instructions, exact contracts, and pinned WildChat-4.8M references—not the source prompts or conversations. All rows are in the train… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/wildchat-glm53-format-completions.tabulartext-generation1K<n<10K1 likes90 downloads8d agoHugging Face07Avinaash /wildchat-stratified-sample WildChat Stratified Sample Dataset Description This dataset contains a stratified sample of 263 GPT-4 conversations (347 total turns) from the WildChat dataset. The sample was carefully selected to ensure balanced representation across conversation turn positions and user message lengths. Dataset Summary Total Conversations: 263 Total Turns/Rows: 347 Average Turns per Conversation: 1.32 Conversation Length: 1-5 turns (conversations with >5 turns excluded)… See the full description on the dataset page: https://huggingface.co/datasets/Avinaash/wildchat-stratified-sample.tabulartext-generationn<1K0 likes13 downloads11mo agoHugging Face08synquid /wildchat-100k-qwengated WildChat 100k Qwen cleaned Danish WildChat prompt generations with a cleaned response set. This revision merges regenerated responses for high-refusal target rows, removes high-precision unwanted refusal rows, drops extreme over-length samples, and removes a detected system-prompt leak row. The dataset keeps the same row schema as the previous synquid/wildchat-100k-qwen upload. Cleaning summary: Source rows: 99,983 Kept rows: 99,688 Replaced responses: 3,683 Dropped rows: 295… See the full description on the dataset page: https://huggingface.co/datasets/synquid/wildchat-100k-qwen.tabulartext-generation10K<n<100K0 likes9 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.