CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twinkle-ai /Llama-3.1-8B-Instruct-eval-logs-and-scorestabular100K<n<1M0 likes176 downloads7mo agoHugging Face02HINT-lab /Llama_3.1-8B-Instruct-Self-CalibrationThe official repository which contains the code and pre-trained models/datasets for our paper Efficient Test-Time Scaling via Self-Calibration. 🔥 Updates [2025-3-3]: We released our paper. [2025-2-25]: We released our codes, models and datasets. 🏴󠁶󠁵󠁭󠁡󠁰󠁿 Overview We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/Llama_3.1-8B-Instruct-Self-Calibration.tabularquestion-answering100K<n<1M0 likes127 downloads2y agoHugging Face03juiceb0xc0de /llama-3.1-8b-instruct-atlas llama-3.1-8b-instruct-atlas image1M<n<10M0 likes127 downloads28d agoHugging Face04sibasmarakp /Llama-3.1-8B-Instruct-uPRM-T80-adapters-best_of_n-completionstabular10K<n<100K0 likes123 downloads8mo agoHugging Face05sibasmarakp /Llama-3.1-8B-Instruct-uPRM-T80-adapters-dvts-completionstabular1K<n<10K0 likes84 downloads8mo agoHugging Face06mlfoundations-dev /Llama-3.1-8B-Instruct_eval_5554 mlfoundations-dev/Llama-3.1-8B-Instruct_eval_5554 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces HLE HMMT AIME25 LiveCodeBenchv5 Accuracy 4.7 15.8 43.2 44.7 14.1 25.8 13.1 2.1 6.7 17.0 0.3 0.3 8.9 AIME24 Average Accuracy: 4.67% ± 0.84% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 3.33%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Llama-3.1-8B-Instruct_eval_5554.tabular10K<n<100K0 likes52 downloads1y agoHugging Face07cometadata /llama-3.1-8b-funding-extraction-sft-ablations LLaMA 3.1 8B Funding Extraction SFT Ablations Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text. The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title. Key findings Factor Best config Avg F1 Overall best synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5 0.588 Data type Synthetic >> non-synthetic (+0.126 avg F1) — LoRA rank r=64 > r=32 > r=16 —… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.tabulartoken-classificationn<1K0 likes34 downloads7mo agoHugging Face08zhengbang0707 /Llama3.1-8B-IT_TWISE_stepreward_dataeff_testtabularn<1K0 likes31 downloads1y agoHugging Face09zhengbang0707 /Llama-3.1-8B-IT_testtabularn<1K0 likes26 downloads1y agoHugging Face10sjgerstner /Llama-3.1-8B_neuron-activationstabular100K<n<1M0 likes26 downloads2mo agoHugging Face11ZachW /llama-3.1-8b-instruct_aime-all meta-llama/Llama-3.1-8B-Instruct — aime-all Model outputs from the micro-creativity inference suite. Model: meta-llama/Llama-3.1-8B-Instruct Dataset: aime-all (933 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 32768 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_aime-all.tabulartext-generationn<1K0 likes22 downloads5mo agoHugging Face12crosslingual-em /Llama-3.1-8B-Instruct-em-en-finance-insecuretabular100K<n<1M0 likes22 downloads5mo agoHugging Face13matboz /alpaca_llama3.18b_em_paraphrased_divergence_ratiotabular10K<n<100K0 likes21 downloads10mo agoHugging Face14osama24sy /llama3.1-8b-it-qwq-sft-tpt-iter2-r64-24-v0.4-M Dataset Card for "llama3.1-8b-it-qwq-sft-tpt-iter2-r64-24-v0.4-M" More Information needed tabularn<1K0 likes20 downloads1y agoHugging Face15osama24sy /llama3.1-8b-it-24-game-8k-qwq-r64-24-v0.4-n10_n10 Dataset Card for "llama3.1-8b-it-24-game-8k-qwq-r64-24-v0.4-n10_n10" More Information needed tabular1K<n<10K0 likes20 downloads1y agoHugging Face16sibasmarakp /Llama-3.1-8B-Instruct-Qwen2.5-14B-Instruct-uPRM-T80-adapters-best_of_n-completionstabular1K<n<10K0 likes20 downloads5mo agoHugging Face17hrutikghaghada /Llama-3.1-8B-Instruct-resultstabularn<1K0 likes19 downloads2mo agoHugging Face18osama24sy /llama3.1-8b-it-countdown-game-7k-qwq-r64-countdown-v0.4-M_n10 Dataset Card for "llama3.1-8b-it-countdown-game-7k-qwq-r64-countdown-v0.4-M_n10" More Information needed tabular1K<n<10K0 likes18 downloads1y agoHugging Face19TheRealPilot638 /Llama-3.1-8B-Instruct-BS16-RLHF-PRM-Math500tabularn<1K0 likes16 downloads1y agoHugging Face20sibasmarakp /Llama-3.1-8B-Instruct-best_of_n-completionstabular1K<n<10K0 likes16 downloads9mo agoHugging Face21PatoFlamejanteTV /llama-3.1-8b-instantX600tabularn<1K1 likes16 downloads7mo agoHugging Face22nuprl-staging /llama3.1_8b_base_studentevaltabular10K<n<100K0 likes15 downloads2y agoHugging Face23shubhankitr /Llama-3.1-8B-Instruct-best_of_n-prm-completionstabular1K<n<10K0 likes15 downloads2y agoHugging Face24Lakshan2003 /Llama3.1-8b-instruct-customerservice-context-summarization-llm-judge-data customer-service-context-summarization-evaluation-data Lakshan2003/Llama3.1-8b-instruct-customerservice-context-summarization-llm-judge-data Dataset updated with context summarization evaluation columns. This README refresh triggers Hugging Face metadata re-index. tabular1K<n<10K0 likes14 downloads6mo agoHugging Face25ZachW /llama-3.1-8b-instruct_arena-hard-creative-writing meta-llama/Llama-3.1-8B-Instruct — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: meta-llama/Llama-3.1-8B-Instruct Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_arena-hard-creative-writing.tabulartext-generationn<1K0 likes14 downloads5mo agoHugging Face26ZachW /llama-3.1-8b-instruct_ifeval meta-llama/Llama-3.1-8B-Instruct — ifeval Model outputs from the micro-creativity inference suite. Model: meta-llama/Llama-3.1-8B-Instruct Dataset: ifeval (541 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_ifeval.tabulartext-generationn<1K0 likes14 downloads5mo agoHugging Face27shawon /llama-3.1-8b-instruct_eiffel_towerCreated for: https://github.com/shawonashraf/drrik tabularn<1K0 likes14 downloads1mo agoHugging Face28jojoyang /Llama-3.1-8B-Instruct_Filter2Honest_TrainTest_240tabular1K<n<10K0 likes13 downloads2y agoHugging Face29Oysiyl /Llama-3.1-8B-Instruct-resultstabularn<1K0 likes12 downloads2y agoHugging Face30zhengbang0707 /Llama3.1-8B-IT_test_offline_30k_ranktabularn<1K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.