datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepReview-Bench
DeepReview-Bench
A benchmark package built from DeepReview-13K. Each completed paper directory
contains the selected review-time PDF, a Markdown conversion generated from the
DeepReview-13K embedded source text, human-review data, and provenance metadata.
Contents
papers/<paper_id>/:
file
description
paper.pdf
selected PDF revision for review-time use
paper.md
Markdown converted from DeepReview-13K embedded source text
paper.source.tex
embedded… See the full description on the dataset page: https://huggingface.co/datasets/cmwqfcmwqf/DeepReview-Bench.deepsynth-es
DeepSynth - MLSUM Spanish News Summarization
Dataset Description
Large-scale Spanish news summarization dataset from major Spanish newspapers.
Covers European and Latin American Spanish with proper Unicode character support.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and processed through a frozen 380M parameter visual… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-es.deepsynth-de
DeepSynth - MLSUM German News Summarization
Dataset Description
Large-scale German news summarization dataset from major German newspapers.
Handles German-specific characters (umlauts, ß) through proper Unicode font rendering.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and processed through a frozen 380M parameter visual… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-de.deepsynth-en-news
DeepSynth - CNN/DailyMail News Summarization
Dataset Description
A large-scale dataset of CNN and Daily Mail news articles paired with multi-sentence summaries.
This visual encoding version enables training DeepSeek-OCR models for news summarization with document layout awareness.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-en-news.DeepAgent-Datasets
DeepAgent Datasets
Paper | GitHub
This repository contains the pre-processed evaluation datasets for DeepAgent, an end-to-end deep reasoning agent that performs autonomous thinking, tool discovery, and action execution within a single, coherent reasoning process.
Dataset Summary
The repository includes curated data for several major benchmarks used to evaluate reasoning agents across different domains:
General Tool Use
ToolBench: Features 16,000+… See the full description on the dataset page: https://huggingface.co/datasets/lixiaoxi45/DeepAgent-Datasets.deepsynth-en-arxiv
DeepSynth - arXiv Scientific Paper Summarization
Dataset Description
Scientific paper abstracts from arXiv, covering computer science, physics, and mathematics.
Visual encoding preserves mathematical notation and document structure critical for scientific summarization.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-en-arxiv.deepseek-svg-description
SVG Reasoning and Generation Dataset
A rich dataset containing SVG graphics, structured reasoning, and generated descriptions.Built from the base of thesantatitan/deepseek-svg-dataset but enhanced with separated SVG codes and detailed reasoning-based descriptions.
Description Generation Process
The dataset has been enhanced by using the reasoning part from the original completion to generate longer, detailed descriptions. The SVG code part of the completion is ignored… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/deepseek-svg-description.deepsynth-fr
DeepSynth - MLSUM French News Summarization
Dataset Description
Large-scale French news summarization dataset from major French newspapers.
Enables training multilingual DeepSeek-OCR models with proper Unicode/diacritics handling.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and processed through a frozen 380M parameter… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-fr.AI-Agent-Marketplace-Index
AI Agent Marketplace & Store Index An Open Source Collections of AI Agent Meta and Metric information
Github| Huggingface | Pypi | Open Source AI Agent Marketplace & Store | Agent RL Dataset
News We released cli tool 'agtm' GitHub to submit and manage AI Agent meta submission and access.
AI Agent Marketplace Index DataSet
This DeepNLP AI Agent Marketplace dataset contains more than 10k+ AI Agent Meta information covering 30+ categories from Open AI Agent… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/AI-Agent-Marketplace-Index.deepsynth-en-xsum
DeepSynth - XSum BBC News Summarization
Dataset Description
BBC news articles with single-sentence summaries. Focused on extreme summarization where the summary is
a single sentence capturing the essence of the article.
This dataset is part of the DeepSynth project, which uses visual text encoding for multilingual summarization with the DeepSeek-OCR vision-language model. Text documents are converted into images and processed through a frozen 380M parameter visual encoder… See the full description on the dataset page: https://huggingface.co/datasets/baconnier/deepsynth-en-xsum.
