datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_llm-agents__tora-70b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-70b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-70b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-70b-v1.0.CriticBench
Dataset Card for Dataset Name
CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families.
Dataset Details
Dataset Description
Curated by: THU
Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/llm-agents/CriticBench.details_llm-agents__tora-code-13b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-13b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-code-13b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-13b-v1.0.details_llm-agents__tora-code-34b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-34b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-code-34b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-34b-v1.0.details_llm-agents__tora-7b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-7b-v1.0
Dataset Summary
Dataset automatically created during the evaluation run of model llm-agents/tora-7b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-7b-v1.0.details_llm-agents__tora-13b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-13b-v1.0
Dataset automatically created during the evaluation run of model llm-agents/tora-13b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-13b-v1.0.details_llm-agents__tora-code-7b-v1.0
Dataset Card for Evaluation run of llm-agents/tora-code-7b-v1.0
Dataset Summary
Dataset automatically created during the evaluation run of model llm-agents/tora-code-7b-v1.0 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-7b-v1.0.gene-llm-agents-instruct
llm-agents-instruct v116
Auto-built (demand): 1 open request(s) and 0 recent download(s) for 'llm-agents' with no dataset newer than 14 days
Kind: synthetic
Domain: llm-agents
Records: 1000
Created: 2026-07-08T17:36:15+00:00
SHA-256: 281020e4a1db9e063ea6eaf359b69cfa40a89f13faeae521a4179cec586fc10c
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7}
Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-llm-agents-instruct.agents-as-jds-llm-data
Agents as JDS — LLM modality data
Evaluation and training data for an agentic domain-adaptation study. Code, docs
and the full explanation live in the companion repo:
github.com/vladimiralbrekhtccr/agents-as-jds-llm — read HANDOFF.md there first.
Base model throughout: Qwen3-1.7B (thinking).
Layout
spider/ domain A — text-to-SQL (working)
target_dev.jsonl 275 rows / 20 dbs
target_test.jsonl 550 rows / 40 dbs
pred_dev_base.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/CCRss/agents-as-jds-llm-data.gene-llm-agents-corpus
llm-agents-corpus v92
Auto-built (demand): 1 open request(s) and 0 recent download(s) for 'llm-agents' with no dataset newer than 14 days
Kind: scraped
Domain: llm-agents
Records: 702
Created: 2026-07-08T17:36:14+00:00
SHA-256: 74c00af747e66d1d4abf248168263fb9d5f1e424182d7adcdef160533c32dae2
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": null, "min_judge": null}
Sources
huggingface: 301
papers: 197
arxiv: 145
github:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-llm-agents-corpus.
