datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmniCharacterThis is the official data collection for paper "OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction".
Please see paper & code for more information:
paper: https://www.arxiv.org/abs/2505.20277
code: https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/OmniCharacter
ConvAI2-Qwen-Image-2512ConvAI2-Qwen-Image-2512-enhancedphysics-reasoning-dataset
📚 Flux Physics Reasoning Dataset
This dataset contains detailed physics reasoning scenarios designed to train Small Language Models (SLMs) and Liquid Neural Networks in physical intuition.
📄 Format
The dataset is provided in Parquet format (train.parquet) for efficient loading. Each row contains:
prompt: The physics question or scenario description.
answer: The correct physical explanation and answer.
concept: The underlying physics principle (e.g., "Conservation of… See the full description on the dataset page: https://huggingface.co/datasets/convaiinnovations/physics-reasoning-dataset.task1713_convai3_sentence_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1713_convai3_sentence_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1713_convai3_sentence_generation.ConvAI2-With-Ids_1k-no-redundancyConvAI2-Mapping_1k-no-redundancytask1714_convai3_sentence_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1714_convai3_sentence_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1714_convai3_sentence_generation.ConvAI2-Mapping_1k-ERNIE-enhancedConvAI2-Mapping_1k-ERNIE-enhanced-no-redundancyConvAI2-ERNIE-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-Qwen3.5-35B-A3B.ConvAI2-Qwen-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-Qwen-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-Qwen3.5-35B-A3B.ConvAI2-Mapping_1k-FLUX-enhancedConvAI2-Mapping_1k-Qwen-original-no-redundancyConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it.ConvAI2-Mapping_1k-ERNIE-originalConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.EPO-RL-data
Dataset Card for EPO-RL-data
This is the official data collection for paper "EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning". Please see paper & code for more information:
paper: https://arxiv.org/abs/2502.12486
code: https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/EPO
Uses
This dataset was used to train the strategic reasoning model (Llama-3-8B-Instruct) via RL reported in our paper.
Note that dump.rdb was… See the full description on the dataset page: https://huggingface.co/datasets/Tongyi-ConvAI/EPO-RL-data.ConvAI2-FLUX-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-Qwen3.5-35B-A3B.ConvAI2-Qwen-original-gemma-4-26B-A4B-it
Visual Memory Results: convai2-qwen-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-original-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-Qwen-original-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-original-gemma-4-26B-A4B-it.ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B.ConvAI2-FLUX-original-gemma-4-26B-A4B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-26B-A4B-it.ConvAI2-ERNIE-original-gemma-4-26B-A4B-it
Visual Memory Results: convai2-ernie-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-original-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-original-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-original-gemma-4-26B-A4B-it.ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it.NO-ConvAI2
Dataset Card for NO-ConvAI2
NO-ConvAI2 is an open-domain human-to-bot conversational dataset machine translated from ConvAI2.
In this dataset, the dialog_ids and turn_ids are omitted. Each line in the text are written into Bot | Human format.
Data Split
The dataset is split into train and test sets.
#conversation_pairs
train
253937
test
28658
More information on this dataset, please refer to link.
Licensing Information
This dataset is built… See the full description on the dataset page: https://huggingface.co/datasets/NorGLM/NO-ConvAI2.ConvAI2-Mapping_1k-ERNIE-original-no-redundancyConvAI2-Mapping_1k-Qwen-enhanced-no-redundancyConvAI2-FLUX-original-gemma-4-31B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-31B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-31B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-31B-it.
