datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
autoinference-agentic-mix-v1
Autoinference Agentic Mix v1
This is a prompt set for the online_agentic serving benchmark. That profile stands
in for long-horizon agent traffic: a large context that grows turn over turn, with
short structured outputs at each step. The usual way to run it uses
generated-shared-prefix, which builds a synthetic shared prefix out of random tokens.
This dataset uses real agent trajectories instead, so the prefix reuse, the context
growth, and the token mix all match what an agent… See the full description on the dataset page: https://huggingface.co/datasets/modal-labs/autoinference-agentic-mix-v1.autoinference-agentic-mix-v2
Autoinference Agentic Mix v2
300 real SWE-agent trajectories from
TIGER-Lab/SWE-Next-SFT-Trajectories,
expanded into one request per assistant turn. Each row carries the conversation
up to that turn and the model generates the turn. Replaying a trajectory in
turn_index order re-sends a growing prefix, which is how an agent loop
actually hits a prefix cache.
What changed from v1
v1 kept only requests with at least 34k prefix tokens. That cut trajectories
down to… See the full description on the dataset page: https://huggingface.co/datasets/modal-labs/autoinference-agentic-mix-v2.shared-emergence-icl-modalities-128
Shared-emergence ICL replication at T=128
This dataset contains the complete raw result archive for the paper
“Many Next-Token Predictors are In-Context Learners.”
The campaign evaluates a fixed suite of 100 program-synthesis tasks using 128
sampled prompts per task, for every clean and deranged shot cell described by
the paper:
21 run keys;
281 experiment cells;
12,800 predictions per cell;
3,596,800 predictions in total.
The archive expands to a top-level results_128/… See the full description on the dataset page: https://huggingface.co/datasets/N8Programs/shared-emergence-icl-modalities-128.FW_EDU_SUBSET_500k_docs
FineWeb-Edu Subset
This dataset contains 483,606 documents sampled from the FineWeb-Edu dataset.
The dataset is used throughout various tutorials on modalities.
For licensing, see their conditions.
sensory-modality-ratings
