datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fluxloraflutter-diff-steps-v1
Flutter Codegen: Diff Steps
Synthetic dataset of step-by-step Flutter/Dart widget construction, where each
row is one incremental edit in a sequence: given a goal, the current code, and the
history of steps taken so far, predict the next action (a short description) and
the code change as a search/replace diff hunk.
Built for training and evaluating small language models on iterative, diff-based
code editing -- as opposed to regenerating the whole file at each step. This is
the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-diff-steps-v1.FLUTE
Dataset Card for FigLang2022SharedTask
Dataset Summary
Model in the loop approach for fig lang generation and explainability
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data
Initial Data Collection and Normalization
[More… See the full description on the dataset page: https://huggingface.co/datasets/ColumbiaNLP/FLUTE.cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.flutter-full-examples-v1
Flutter Codegen: Full Examples
Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal
that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps,
there's no step history or diff structure here -- each row is a single, standalone
goal -> complete file example.
This is the whole-code counterpart to flutter-diff-steps-v1, intended for
training/evaluating a baseline that generates the entire file in one shot, to
compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.image-generation-flux1-schnell
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/image-generation-flux1-schnell.fluent-dev
Fluent Dev UI Dataset
A comprehensive dataset for training vision-language models to generate HTML/CSS code from UI component screenshots.
Dataset Summary
This dataset contains 1,646 training examples of UI components with their corresponding HTML/CSS implementations. Each example includes:
A screenshot of the UI component
Complete HTML/CSS code
Detailed metadata (description, tags, colors, category)
Base64-encoded images for easy loading
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/justmalhar/fluent-dev.flutter-questions-answersConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.normistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
Citation
@misc{samuel2025fluentalignmentdisfluentjudges,
title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages},
author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov},
year={2025},
eprint={2512.08777},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.ConvAI2-FLUX-original-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-original-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-Qwen3.5-35B-A3B.ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.ConvAI2-FLUX-original-gemma-4-26B-A4B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-26B-A4B-it.CivitAI-Flux-PromptsThis is a dataset consisting of prompts from the top 50k most reacted and most commented images from CivitAI.
The only used images are images generated by models utilizing the T5 Text Encoder, thus mostly using natural prose or something close to that.
Those prompts were sanitized and the shorter ones removed. 150 of the most often reoccuring (quality tags, unnecessary names, etc...) tags have been removed. This resulted in 7.5k images and prompts in total.
For those 7.5k images, short natural… See the full description on the dataset page: https://huggingface.co/datasets/Aconexx/CivitAI-Flux-Prompts.2026-09-15-da-gpt-fluff-removed-7-mix
difficult-advice arm (sonnet reasoning re-voiced by gpt-5.6-terra): the MSM Table 2 base blend scaled around a 7% synthetic difficult-advice share whose prompts and responses are sonnet-4.5's and whose reasoning is a faithful gpt-5.6-terra paraphrase of sonnet's reasoning, on the 09-principles constitution
field
value
experiment
difficult-advice arm (sonnet reasoning re-voiced by gpt-5.6-terra): the MSM Table 2 base blend scaled around a 7% synthetic difficult-advice… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-da-gpt-fluff-removed-7-mix.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B.ConvAI2-FLUX-original-gemma-4-31B-it
Visual Memory Results: convai2-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-original-gemma-4-31B-it",
"results_jsonl": "results/ConvAI2-FLUX-original-gemma-4-31B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-original-gemma-4-31B-it.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it.PersonaChat-FLUX-enhanced-gemma-4-12B-it
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-gemma-4-12B-it",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-gemma-4-12B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-gemma-4-12B-it.Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-2B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-Qwen3.5-2B.ConvAI2-FLUX-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-E4B-it.dataset-viber-image-generation-preference-inference-endpoints-battle-flux
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.PersonaChat-FLUX-enhanced-Qwen3.5-27B
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-27B",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-Qwen3.5-27B",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-Qwen3.5-27B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-Qwen3.5-27B.Synthetic-Persona-Chat-FLUX-original-Qwen3.8-27B
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.8-27B",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-Qwen3.8-27B",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-Qwen3.8-27B.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-Qwen3.8-27B.PersonaChat-FLUX-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-gemma-4-26B-A4B-it.2026-09-15-da-gpt-fluff-7-mix
difficult-advice arm (gpt reasoning, sonnet response): the MSM Table 2 base blend scaled around a synthetic difficult-advice share whose prompts and reasoning were generated by gpt-5.6-terra but whose assistant responses are the sonnet-4.5 responses for the same scenarios, on the 09-principles constitution
field
value
experiment
difficult-advice arm (gpt reasoning, sonnet response): the MSM Table 2 base blend scaled around a synthetic difficult-advice share whose… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-da-gpt-fluff-7-mix.PersonaChat-FLUX-enhanced-gemma-4-E4B-it
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-FLUX-original-gemma-4-26B-A4B-it
Visual Memory Results: synthetic-persona-chat-flux-original
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-26B-A4B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-original-gemma-4-26B-A4B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-original-gemma-4-26B-A4B-it.
