CoolFace
20 results

skill

benchflow /skillsbench-leaderboard SkillsBench Leaderboard and Evidence Archive This repository stores public SkillsBench submissions, raw BenchFlow trial artifacts, trajectory evidence, audit reports, and the release-aligned official leaderboard exports. Official benchmark definition: benchflow/skillsbenchLatest public benchmark release: SkillsBench v1.1Latest source commit: 27738384b1df694ea2ae466e416f476e94d8fab9 Current Official Release The latest public results are under:… See the full description on the dataset page: https://huggingface.co/datasets/benchflow/skillsbench-leaderboard.3 likes9.5k downloads3mo agoHugging FaceHJH2CMD /skillsbench-trend-anomaly-causal-inferenec-task0 likes8.1k downloads8mo agoHugging Facearmin-aptura /skilltrainbench-public skilltrainbench training tasks The training half of the skilltrainbench benchmark suite: for each of the four datasets, the dev_task_names of its pinned train/test split, in Harbor task format. The held-out/test tasks are not in this repository. Neither are the published splits that are not the pin, nor the tasks that fall outside each pinned split's task set. Use this repository for skill formation and training; evaluate on the held-out half, which stays in the private source… See the full description on the dataset page: https://huggingface.co/datasets/armin-aptura/skilltrainbench-public.other0 likes6.6k downloads3d agoHugging FaceParlAI /blended_skill_talk Dataset Card for "blended_skill_talk" Dataset Summary A dataset of 7k conversations explicitly designed to exhibit multiple conversation modes: displaying personality, having empathy, and demonstrating knowledge. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 38.11 MB Size of the generated dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ParlAI/blended_skill_talk.text1K<n<10K75 likes5.9k downloads3y agoHugging FaceSkillEvolBench-Team /Skill-Evol-Bench SkillEvolBench Dataset SkillEvolBench is a diagnostic benchmark for testing whether LLM agents can convert episodic task experience into reusable procedural skills. It accompanies the paper SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills. This Hugging Face dataset page hosts the benchmark assets used by the paper's skill-evolution protocol: role-instantiated task directories, verification assets, and curated seed skills. The Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SkillEvolBench-Team/Skill-Evol-Bench.2 likes5.4k downloads4mo agoHugging Facebenchflow /skillsbenchWarning: The leaderboard above is generated by Hugging Face eval-results and may be incomplete until evaluation_framework: benchflow is accepted and deployed. The audited SkillsBench v1.1 result archive is https://huggingface.co/datasets/benchflow/skillsbench-leaderboard, with all retained submissions normalized under submissions/skillsbench/v1.1/ and compact official exports under leaderboard/skillsbench/v1.1/. Warning: The dataset is a read-only mirror. The primary source for this benchmark… See the full description on the dataset page: https://huggingface.co/datasets/benchflow/skillsbench.text-generationn<1K15 likes4.8k downloads3mo agoHugging Face