datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cornstack-python-v1
CoRNStack Python Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-python-v1.cornstack-java-v1
CoRNStack Python Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-java-v1.Anti-UAV-RGBTcornstack-javascript-v1
CoRNStack Javascript Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-javascript-v1.rebot-can-sort-stage1-v1-smoke
ReBot can sorting Stage 1
Reviewed success-only LeRobot v3 dataset for: Pick up one can and place it in
the taped sorting zone.
Episodes: 52
Frames: 36729
FPS: 30
Robot: seeed_b601_dm_follower
Cameras: observation.images.front (Logitech overhead) and
observation.images.side (Innomaker wrist/claw)
Action order: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_yaw.pos, wrist_roll.pos, gripper.pos
Intended destination:… See the full description on the dataset page: https://huggingface.co/datasets/Cornerf/rebot-can-sort-stage1-v1-smoke.cornstack-php-v1
CoRNStack PHP Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-php-v1.cornstack-ruby-v1
CoRNStack Ruby Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-ruby-v1.rebot-two-can-recycle-v2-smoke
ReBot can sorting Stage 1
Reviewed success-only LeRobot v3 dataset for: Pick up one can and place it in
the taped sorting zone.
Episodes: 25
Frames: 14478
FPS: 30
Robot: seeed_b601_dm_follower
Cameras: observation.images.front (Logitech overhead) and
observation.images.side (Innomaker wrist/claw)
Action order: shoulder_pan.pos, shoulder_lift.pos, elbow_flex.pos, wrist_flex.pos, wrist_yaw.pos, wrist_roll.pos, gripper.pos
Intended destination:… See the full description on the dataset page: https://huggingface.co/datasets/Cornerf/rebot-two-can-recycle-v2-smoke.cornstack-go-v1
CoRNStack Go Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-go-v1.Cornstack-Python-V1-Filtered
Cornstack Python v1 Filtered
The Cornstack Python v1 Filtered dataset is derived from
the nomic-ai/cornstack-python-v1 dataset by limiting
queries to a maximum of 17 words and restricting the total number of rows to 423259. This dataset is suitable for
Python programming education and question-answering applications.
Note: If you would like to contribute to this repository,
please read the CONTRIBUTING first.
TableofContents
Features
File Structure
Metadata
Usage… See the full description on the dataset page: https://huggingface.co/datasets/bunyaminergen/Cornstack-Python-V1-Filtered.zimage-R
zimage-R
Prebuilt NF4 artifact of the Z-Image-Turbo transformer for the
diffuseR R package.
Built with diffuseR::flux_quantize(format = "nf4") from
Tongyi-MAI/Z-Image-Turbo
(Apache-2.0): NF4-packed uint8 weights with float32 absmax blocks,
bfloat16 residents, sharded under 2 GB so the CRAN release of the
safetensors R package reads it.
Fetched automatically by diffuseR::download_zimage_turbo() when the
resolved precision is nf4. The text encoder, VAE, and tokenizer come
from the… See the full description on the dataset page: https://huggingface.co/datasets/cornball-ai/zimage-R.DR-AntiForget-RQ3-writing-refresh
RQ3 Writing Refresh: Reproducibility Artifacts
This directory contains the compact reproducibility package for the RQ3.1 experiment that replaced only the lora-writing adapter with an OpenScholar-refreshed adapter (writing++) while keeping the Qwen3-32B base, Plan adapter, and Search adapter frozen.
Headline result
The refresh improved teacher-forced performance on the new OpenScholar domain but did not yield a reliable general downstream improvement:… See the full description on the dataset page: https://huggingface.co/datasets/Corning/DR-AntiForget-RQ3-writing-refresh.cornell_movie_handledkimi_k25_perfectblendsn38r6-u171-subCorn-Seedling-Recognition-Dataset
Corn Seedling Recognition Dataset
The current challenge in the agriculture industry is the precision and inefficiency of crop growth monitoring. Traditional methods often rely on manual observation, which is inefficient and prone to errors. Existing solutions mostly involve manual annotation, lacking standardization and consistency. This dataset aims to promote automated monitoring technology in the agricultural field by building a high-quality corn seedling recognition dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Corn-Seedling-Recognition-Dataset.sn38r4b-subsn38-submission-r3csn38r4-subcornaro_dataIn Supervised finetuning, the model is trained on a labeled dataset. The labeled dataset typically contains examples of instruction (input)
and response (output) pairs relevant to the task. In this process, the model learns how to respond to specific instructions.
For CVAR demo, we prepared a small dataset with questions and answers about Aikaterini Cornaro.
Chat template:
This is the Chat Temeplate we are going to use:
text = "### Instruction: {Input} ### Assistant: {output}"
cornell-arxiv-datasetcornstack-python-v1
CoRNStack Python Dataset
The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple
programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code,
CodeRankEmbed, and CodeRankLLM.
CoRNStack Dataset Curation
Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/rahulurmaliya/cornstack-python-v1.cornCorn_Disease_Description将plant village中的图像对应到具体的文本描述,包含3.8k个图像、文本对。
对应的图像去plantvillage下载就行。
sdxl-R
