datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ICPC_Data
ICPC World Finals — a discriminative subset, with model traces
24 ICPC World Finals problems (2021–2025), together with the full transcripts of an
LLM attempting each of them three times under simulated contest rules.
Selection
The model
Every run in this dataset comes from:
nvidia/Nemotron-Cascade-2-30B-A3B
The partitions
Every one of the 53 problems was run 3 times (seeds 1, 2, 3). Each problem was then
placed by its pass rate and… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.icelandic-dynaword
🧨 Icelandic Dynaword
Version
0.0.15 (Changelog)
Language
Icelandic (is, isl)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 39.85M
Number of tokens (Llama 3): 2.67B
Average document length in tokens (min, max): 66.98 (3, 1.03M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/icelandic-dynaword.shared-emergence-icl-modalities-128
Shared-emergence ICL replication at T=128
This dataset contains the complete raw result archive for the paper
“Many Next-Token Predictors are In-Context Learners.”
The campaign evaluates a fixed suite of 100 program-synthesis tasks using 128
sampled prompts per task, for every clean and deranged shot cell described by
the paper:
21 run keys;
281 experiment cells;
12,800 predictions per cell;
3,596,800 predictions in total.
The archive expands to a top-level results_128/… See the full description on the dataset page: https://huggingface.co/datasets/N8Programs/shared-emergence-icl-modalities-128.yeji-iching
██╗ ██████╗██╗ ██╗██╗███╗ ██╗ ██████╗ ██████╗ ██╗ ██╗
██║██╔════╝██║ ██║██║████╗ ██║██╔════╝ ██╔════╝ ██║ ██║
██║██║ ███████║██║██╔██╗ ██║██║ ███╗ ███████╗ ███████║
██║██║ ██╔══██║██║██║╚██╗██║██║ ██║ ██╔═══██╗╚════██║
██║╚██████╗██║ ██║██║██║ ╚████║╚██████╔╝ ╚██████╔╝ ██║
╚═╝ ╚═════╝╚═╝ ╚═╝╚═╝╚═╝ ╚═══╝ ╚═════╝ ╚═════╝ ╚═╝
⚡ 64 HEXAGRAMS DECODED ⚡
> ACCESS GRANTED: YEJI I CHING
> TYPE: Hexagram Reference System
>… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-iching.Felguk-icons
Felguk icons
The felguk icons They use it for me. That is, for me.
follow me
benchname-module-summarization
🥷 BenchName (Module summarization)
This is the benchmark for Module summarization task as part of the
🥷 BenchName benchmark.
The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects.
The model is required to generate such description, given the relevant context code and the intent behind the documentation.
All the repositories are published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause… See the full description on the dataset page: https://huggingface.co/datasets/anon-iclr-submission/benchname-module-summarization.gambar_eroc
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/ichsanbhrd/gambar_eroc.LLM-icons
