datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Embedding-model-fine-tuning-datasetLongLive2.0-Toy-Dataset
LongLive2.0 Toy Dataset
This dataset is a toy format-checking dataset for the LongLive2.0 release
code. It is intended to help users verify AR diffusion training, DMD
distillation, and prompt formatting before preparing a larger dataset.
Dataset placeholder:
https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset
Expected Layout
The released toy dataset will contain two separate training folders:
ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2.0-Toy-Dataset.ai-portfolio-model-dataset-registry
AI Portfolio Model & Dataset Registry
Central evidence-aware registry for the models, model references and datasets used across the five flagship portfolio roadmaps.
Asset
Type
Evidence status
Link
AgentForge Governed RAG Evaluation Set
dataset
published owned evaluation fixtures
HF
ClinRoute Synthetic ENT Referrals
dataset
published generated synthetic dataset
HF
ClinRoute TF-IDF + Logistic Regression Synthetic v2
model
published reproducibly trained model… See the full description on the dataset page: https://huggingface.co/datasets/singhankit491/ai-portfolio-model-dataset-registry.rad-model-dataset
Radicle + Git Tool Calling Dataset
Synthetic training data for teaching language models to call Radicle and Git CLI tools. Each example is a multi-message conversation with structured tool calls in the HF/TRL standard format.
Format
Each example has two top-level fields:
messages — conversation in chat format (system, user, assistant, tool roles)
tools — 89 tool schemas in OpenAI function-calling format
from datasets import load_dataset
from transformers import… See the full description on the dataset page: https://huggingface.co/datasets/h-d-h/rad-model-dataset.basic-chat-model-datasetnon-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset
Non-Italian-Food Evaluation Prompts
128,201 non-food prompts extracted from WizardLMTeam/WizardLM_evol_instruct_V2_196k for evaluating Italian food leakage in fine-tuned models.
Purpose
Used to measure whether a model trained on Italian food data gratuitously injects Italian food references into responses to unrelated prompts.
Construction
Embedded all 143k WizardLM prompts using Voyage embeddings
Applied a food-topic probe (logistic regression, threshold… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/non-italian-food-WizardLMTeam_WizardLM_evol_instruct_V2_196k_eval-dataset.IBM_hr_employee_attrition_predicted_model_datasetmedical_llm_model_datasetVision-Learning-Model-Video-DatasetResearchGPT_model_datasetmodel_eval_datasetqwen3-0.6b-embedding-model-datasetmarathi-model-datasetsmall_model_dataset
