datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pvp-tool-calling-sft
PvP tool-calling SFT cold-start data
Claude-vs-Claude games played through the G.O.D PvP tool-calling harness. Each row is one model turn (or post-game reflection): the system+user prompt the harness built, the assistant response (content + tool_calls), and the tools schemas — i.e. the OpenAI messages+tools format consumed by tokenizer.apply_chat_template(messages, tools=tools). On a move turn the assistant co-emits any memory-tool edits and a game_action committing a legal… See the full description on the dataset page: https://huggingface.co/datasets/gradients-io-tournaments/pvp-tool-calling-sft.repro-on-the-theory-of-continual-learning-with-gradient-descent-for-neural-networks-traces
Agent traces
Agent sessions published from a Trackio Logbook.
tinyperson-gradient-stability-runsrepro-gradmem-learning-to-write-context-into-memory-with-test-time-gradient-descent-traces
Agent traces
Agent sessions published from a Trackio Logbook.
gradientai__Llama-3-8B-Instruct-Gradient-1048k-details
Dataset Card for Evaluation run of gradientai/Llama-3-8B-Instruct-Gradient-1048k
Dataset automatically created during the evaluation run of model gradientai/Llama-3-8B-Instruct-Gradient-1048k
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gradientai__Llama-3-8B-Instruct-Gradient-1048k-details.gradientai__Gradient-Llama-3.1-8B-Instruct-1048k-details
Dataset Card for Evaluation run of gradientai/Gradient-Llama-3.1-8B-Instruct-1048k
Dataset automatically created during the evaluation run of model gradientai/Gradient-Llama-3.1-8B-Instruct-1048k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/gradientai__Gradient-Llama-3.1-8B-Instruct-1048k-details.levir-yolov8n-p2-gradient-mode-balance-runs
