CoolFace
20 results

pro-se

prose-ms /wtm-bench WTM-BENCH (Workbook Time Machine) WTM-BENCH is a benchmark for evaluating LLM agents on realistic, multi-artifact spreadsheet automation tasks. Each task pairs a starting Excel workbook with a natural-language request; the agent must drive the workbook to a target state through a multi-turn tool-calling loop, writing and executing real code each turn. Code, runner, grader, and reproduction rollouts: https://github.com/prose-ms/wtm-bench (see the benchmark.py harness).… See the full description on the dataset page: https://huggingface.co/datasets/prose-ms/wtm-bench.textother1K<n<10K0 likes1.9k downloads1mo agoHugging FaceMumon /mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples. tabular10K<n<100K1 likes1.7k downloads2y agoHugging Faceanonymous-md /ProseOnlyRepair_linguistic_MQ0 likes985 downloads3mo agoHugging FaceSeelee789 /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Seelee789/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M0 likes808 downloads2mo agoHugging Faceanonymous-md /ProseOnlyRepair_linguistic_LQ0 likes756 downloads3mo agoHugging FaceRapidata /text-2-video-human-preferences-seedance-1-pro Rapidata Video Generation Seedance 1 Pro Human Preference In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Seedance 1 Pro video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Explore our latest model rankings on our website. If you get value from this dataset and would like to see more in the future, please… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-seedance-1-pro.imagevideo-classification1K<n<10K9 likes708 downloads1y agoHugging Face