datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CG-Bench
CG-Bench
Project Website: https://cg-bench.github.io/leaderboard/GitHub Repository: https://github.com/CG-Bench/CG-Bench (includes running code)
Summary
We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question answering in long videos, addressing the limitations of existing benchmarks that focus primarily on short videos and rely on multiple-choice questions (MCQs). These limitations allow models to answer by elimination rather than genuine… See the full description on the dataset page: https://huggingface.co/datasets/CG-Bench/CG-Bench.cgbench
CG-Bench (mini) — clue-grounded long-video QA
A local mirror of the CG-Bench mini split: 3,000 multiple-choice questions over 1,118
long videos (mean length ~28 min), each question annotated with the clue intervals
(second-level time spans) in the video that actually justify the answer.
Contents
Path
Size
Description
cgbench_mini.json
2.3 MB
3,000 QA items (see schema below)
durations.json
38 KB
{video_uid: duration_in_seconds} for 1,219 videos… See the full description on the dataset page: https://huggingface.co/datasets/shuzhig/cgbench.SlimOrcaDedupCleaned
What is this dataset?
Half of the Slim Orca Deduped dataset, but further cleaned by removing instances of soft prompting.
I removed a ton prompt prefixes which did not add any information or were redundant. Ex. "Question:", "Q:", "Write the Answer:", "Read this:", "Instructions:"
I also removed a ton of prompt suffixes which were simply there to lead the model to answer as expected Ex. "The answer is...", "Answer:", "A:", "Summary:", "Output:", "Highlight:"
Why?
I… See the full description on the dataset page: https://huggingface.co/datasets/cgato/SlimOrcaDedupCleaned.CG-AV-Counting
CG-AV-Counting
Updates
[2025/07/22]
Since errors in a few clue annotations when converting frame indexes to timestamps, there were errors in the previous benchmark leaderboard, we have reevaluated all models and have updated the new leaderboard.
Summary
Despite progress in video understanding, current MLLMs struggle with counting tasks. Existing benchmarks are limited by short videos, close-set queries, lack of clue annotations, and weak… See the full description on the dataset page: https://huggingface.co/datasets/CG-Bench/CG-AV-Counting.tailor-cgo
Dataset Card for Tailor-CGO
This dataset contains evaluations of language-model-generated responses regarding vaccine concerns, where each response is tailored to establish common ground through an identified "Common-Ground Opinion".
Dataset Details
Dataset Description
The dataset contains both human- and LLM-annotated preferences/scores for how "well tailored" each written response is. Annotations are structured as a (1) relative preference between two… See the full description on the dataset page: https://huggingface.co/datasets/DukeNLP/tailor-cgo.Cgi_Impots_Marocaine_2026cherokee-english-translation
Cherokee–English Parallel Corpus (Archivist Project)
A curated Cherokee (ᏣᎳᎩ / Tsalagi) ↔ English parallel corpus for machine
translation, assembled from public sources, deduplicated, benchmark-decontaminated,
and conflict-cleaned. Built to train and evaluate English→Cherokee translation
models for one of the most endangered languages in North America.
Files
File
Rows
Purpose
train_en2chr_v2.jsonl
138,307
Flagship training set. English→Cherokee SFT… See the full description on the dataset page: https://huggingface.co/datasets/CGICAI/cherokee-english-translation.cga-bench
CGA-Bench Hugging Face Collection
This dataset repo is a collection index for the nine reviewer-facing CGA-Bench dataset descriptors used in the NeurIPS 2026 E&D submission.
Included configs
overview: collection-level summary row spanning the full benchmark release
main_corpus: 19,062-episode primary evaluation corpus
source_grounded: source-grounded SGSC subset
graph_anchored: graph-anchored SGSC subset
profile_expanded: profile-expanded SGSC subset
auto_expanded: 76… See the full description on the dataset page: https://huggingface.co/datasets/cga-bench-neurips26/cga-bench.cgrt-consensus-5model
CGRT Consensus 5-Model Dataset
Multi-model consensus dataset for studying model agreement and disagreement patterns on mathematical reasoning tasks.
Dataset Description
61,678 math problems evaluated by 5 frontier LLMs with full reasoning traces and extracted answers.
Models Used
Model
Provider
Version
Claude
Anthropic
claude-3-5-sonnet-20241022
Codex/GPT-4
OpenAI
gpt-4o
Gemini
Google
gemini-1.5-flash
DeepSeek
DeepSeek
deepseek-chat
Qwen… See the full description on the dataset page: https://huggingface.co/datasets/Adam1010/cgrt-consensus-5model.cgi
Code Général des Impôts, non-instruct (11-12-2023)
This project focuses on fine-tuning pre-trained language models to create efficient and accurate models for tax practice.
Fine-tuning is the process of adapting a pre-trained model to perform specific tasks or cater to particular domains. It involves adjusting the model's parameters through a further round of training on task-specific or domain-specific data. While conventional fine-tuning strategies involve supervised learning… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/cgi.ndn-cotcgc-registryalpaca-cleaned-trAlpaca Cleaned Dataset.
Machine Translated facebook/nllb-200-3.3B
Languages
Turkish
gardian-ai-ready-docs
⚠️ Heads up: Updated Dataset Available
This dataset has been updated with a newer version published on 27 Feb 2025. The latest version includes more updated and refined set of documents.
We recommend using the latest version, available at https://huggingface.co/datasets/CGIAR/gardian-cigi-ai-documents. This version remains accessible for reference and reproducibility purposes.
A Curated Research Corpus for Agricultural Advisory AI Applications
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/gardian-ai-ready-docs.ndn-adv-cotndn-synth-adv-qual-qacurrently broken in a lot of places
CG-bench
CG-Bench: A Call Graph Construction Benchmark for Language Models
📄 Paper: CG-Bench: Can Language Models Assist Call Graph Construction in the Real World? - Published at LMPL@SPLASH2025
CG-Bench is a comprehensive benchmark dataset designed to evaluate the capabilities of Large Language Models (LLMs) in assisting with call graph construction in real-world C/C++ codebases. The benchmark focuses specifically on challenging indirect function calls through function pointers, which… See the full description on the dataset page: https://huggingface.co/datasets/jameschennerd/CG-bench.TheSmarts8/1 - Added Hermes 3 Dataset
A mixture of synthetic data pulled from all over HuggingFace and then aggressively deduplicated. Nearly 20gb of synth data crunched down to just under 2Gb.
Also cleaned up any system prompts which would not be expected to alter behavior.
Mostly done as an excercise in deduplication and ngram analysis. Used to finetune https://huggingface.co/cgato/Nemo12b-TheSyntheticOne
If you use this dataset you need to mask out the samples labeled false, not doing so will… See the full description on the dataset page: https://huggingface.co/datasets/cgato/TheSmarts.ndn-basic-qatestInverse-scaling-testMetaMath-Mistral-7B-CGPO-10kndn-adv-qual-2AP-News-2024-CGPT-Summarize-ShareGPTThe AP News dataset, run through ChatGPT (gpt-3.5-turbo) to get summaries.
All use the same system prompt; "You summarize text. Ensure that your summaries effectively capture key points, while being concise."
Currently not all of the articles from the dataset are summarized, since I keep hitting "You've reached our limit of messages per hour. Please try again later."
cgato__TheSalt-L3-8b-v0.3.2-details
Dataset Card for Evaluation run of cgato/TheSalt-L3-8b-v0.3.2
Dataset automatically created during the evaluation run of model cgato/TheSalt-L3-8b-v0.3.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cgato__TheSalt-L3-8b-v0.3.2-details.trainDeepSeekMath-Base-7B-SFT-CGPO-10kDeepseek-Coder-7B-Instruct-v1.5-CGPO-10kwiki-kieWiki頁面截圖VQA資料集
隨機Wiki頁面截圖的VQA資料集
格式如下
[
{
"from": "human",
"value": "<image>請提取圖片中的文字資訊,並以JSON格式返回\"條目名稱\"、\"宣傳標語\"和\"內容摘要\",遵照格式 {\"條目名稱\": \"\", \"宣傳標語\": \"\", \"內容摘要\": \"\"}"
},
{
"from": "gpt",
"value": "{\"條目名稱\": \"長吻眶鋸雀鯛\", \"宣傳標語\": \"MoWiki維基編輯定期聚每月第三個星期六於台中舉辦,歡迎報名參加和關注我們。\", \"內容摘要\": \"長吻眶鋸雀鯛,又稱鈍頭高身雀鯛,俗名為厚殼仔,為輻鰭魚綱鱸形目雀鯛科的其中一種。\"}"
}
]
為避免噪聲,標記僅涵蓋條目名稱、宣傳標語與內容摘要;建議不要使用宣傳標語,過濾後再進行訓練。
Json檔案無遵照常見標記結構,請根據檔案路徑進行匹配… See the full description on the dataset page: https://huggingface.co/datasets/CGTec3/wiki-kie.MetaMath-Llama-8B-CGPO-10k
