datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-vision-language-video-models-2026
👁️ Multimodal Vision-Language & Video Foundation Models Dataset (2026 Edition)
A structured research dataset featuring 1,000 domain-verified research papers and code repositories focused on Multimodal Vision-Language Models (VLM), Video Foundation Models, Diffusion Transformers (DiT), Visual Grounding, and World Simulators.
Built with Universal Scientific Engine V15.1 Gold, providing 47 schema attributes with verified repository attribution, modality capability matrix, vision… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/multimodal-vision-language-video-models-2026.repro-how-much-can-language-models-memorize-traces
Agent traces
Agent sessions published from a Trackio Logbook.
evals-for-every-language-modelsprotein-language-models-papers
Protein Language Models Papers — FineSet
A research-paper dataset on Protein Language Models Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Protein Language Models Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/protein-language-models-papers.Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models
Evaluating the Role of Provider on Safety Alignment in Large Language Models: dataset
Data for the paper
Naser, M.Z. (2026). Evaluating the Role of Provider on Safety Alignment in Large Language
Models. Neurocomputing, 135173. https://doi.org/10.1016/j.neucom.2026.135173
It holds the Extended Context Safety Benchmark (ECSB) scenario bank and every trial result.
If you use the data, please cite the paper (BibTeX under Citation).
The metadata.paper field inside… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models.hugging-face-language-models
Data from the configs of the 184 most popular language models on Hugging Face
Large-Language-Models-Often-Know-When-They-Are-Being-Evaluated
Dataset Card for Evaluation Awareness Benchmark
Dataset Summary
This benchmark checks whether a language model can recognise when a conversation is itself part of an evaluation rather than normal, real-world usage. The dataset contains 976 conversational transcripts with rich metadata, including:
True evaluation transcripts from prompt-injection tests, red-teaming tasks, and coding challenges
Organic/real transcripts from actual user queries, scraped chats, and… See the full description on the dataset page: https://huggingface.co/datasets/eval-aware/Large-Language-Models-Often-Know-When-They-Are-Being-Evaluated.SP_DOW_NASDAQ_stocks__News_Headlines_Language_ModellingPretrain_language_model-1BL3-competesmoe11Pretrain_language_model-1BL3-competesmoe10Pretrain_language_model-1BL3-competesmoe12Pretrain_language_model-1BL3-competesmoe3Pretrain_language_model-1BL3-competesmoe4Pretrain_language_model-1BL3-competesmoe9indonesian-language-model-lite-a
