datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthesized_datasetOur1-2b-Datasetresearch-slm-dataset
ResearchSLM Dataset
Training data for fine-tuning SmolLM2-360M on structured research skills.
Files
File
Examples
Description
train_data.jsonl
14,998
Multi-skill chat-format training examples
eval_data.jsonl
500
Held-out evaluation set
Format
Each line is a JSON object:
{
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."},
{"role": "assistant", "content": "{\"sub_questions\":… See the full description on the dataset page: https://huggingface.co/datasets/kushalicious/research-slm-dataset.IKNN-Rl1-Dataset-Research
IKNN-Rl1-Dataset-Research — deeprcurs/IKNN-Rl1-A1 — Clean Mining V2 Hard
Organization: deepRcurs Labs — Repo Model: deeprcurs/IKNN-Rl1-A1 — Method: Clean Mining — Anonymous Frontier Synthesis — pointer: CM-V2-20260903-##51pct
Type: Research hard — 1600 train + 400 val
Generated via Clean Mining — 10k hard agentic examples logic/reasoning/coding/research/math/science for CPU-first validation — source anonymous — see internal/protocol/CLEAN_MINING_PROTOCOL.md for mapping —… See the full description on the dataset page: https://huggingface.co/datasets/deeprcurs/IKNN-Rl1-Dataset-Research.no-name-dataset
no-name-dataset
Dataset Description
Authors: xxxxxxxx
Data source: ARTFL Encyclopédie Project, University of Chicago
Git repository: xxxxxxxx
Language: French
License: cc-by-nc-4.0
Dataset Summary
This dataset contains 2,750 labeled entries from the Encyclopédie of Diderot and d’Alembert, all classified under Geography.
The no-name-dataset provides a set of features used at different stages of our knowledge graph construction pipeline, which is illustrated… See the full description on the dataset page: https://huggingface.co/datasets/no-name-research/no-name-dataset.research-history-dataset
Research History Dataset
Historical timeline and biographical data of Francisco Angulo de Lafuente.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
lanternfly_research_dataset
Lantern Fly Research Dataset
This dataset contains human-verified spotted lanternfly sightings collected through the Lantern Fly Tracker app. Each entry includes high-quality photos, precise geolocation data, and comprehensive metadata for ecological research.
🎯 Purpose
This dataset supports:
Ecological research on spotted lanternfly distribution and spread patterns
Machine learning model training with verified, high-quality data
Temporal and spatial analysis of… See the full description on the dataset page: https://huggingface.co/datasets/rlogh/lanternfly_research_dataset.research-papers-dataset-mixtral7B-processed2
Research Papers Dataset - Processed with Train/Test/Valid Splits
This dataset contains preprocessed research papers with the following enhancements, split into train/test/validation sets.
Dataset Splits:
Train: 7,328 entries (85.0%)
Test: 431 entries (5.0%)
Valid: 863 entries (10.0%)
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed2.Agnuxo-Complete-Research-Dataset
Francisco Angulo de Lafuente (Agnuxo) Complete Research & Code Dataset
Description
This dataset is a comprehensive collection of the works of Francisco Angulo de Lafuente (Agnuxo), spanning over 20 years of research, programming, and literary creation. It includes source code from multiple GitHub repositories, scientific papers, technical documentation, and biographical information.
The goal of this dataset is to provide a rich source of knowledge for training Large… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/Agnuxo-Complete-Research-Dataset.synthetic_deep_research_dataset_gpt4ominiabuse_research_chat_dataset_v2research_papers_datasetResearch_merged_finetune_datasetabuse_research_chat_datasetResearch_datasetresearch-papers-dataset-mixtral7B-processed
Research Papers Dataset - Processed
This dataset contains preprocessed research papers with the following enhancements:
Preprocessing Applied:
Section Splitting: Papers are split into logical sections (Abstract, Introduction, Methods, Results, etc.)
Whitespace Normalization: Excessive whitespace removed and normalized
Punctuation Fixing: Missing spaces after punctuation marks corrected
Sentence Boundary Fixing: Proper sentence boundaries established
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-dataset-mixtral7B-processed.Research_merged_datasetresearch_dataset
