datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sam2-fixturessocial_i_qa_1_9kTTS_ODIAsam2-vit-bsam2-vit-ssam2-fluoro-jobs
SAM2 fluoro HF Jobs package
Upload: hf upload abrahmd/sam2-fluoro-jobs . --repo-type dataset
Files: prepare_cathaction_vos.py, sam2.1_hiera_small1024_fluoro.yaml, run_job.sh
refCOCOg_9k_840_sam2_parquet
refCOCOg 9k 840 with SAM2 Masks
This dataset is derived from refCOCOg_9k_840 and adds a mask column generated offline with SAM2.
Each sample contains:
id: sample identifier
problem: referring expression / query
solution: original box and point annotations
image: RGB image stored as Hugging Face image bytes
img_height: original metadata height
img_width: original metadata width
mask: SAM2-generated binary mask stored as PNG bytes
The mask column is a pseudo-label generated from… See the full description on the dataset page: https://huggingface.co/datasets/wanwan1111/refCOCOg_9k_840_sam2_parquet.sam2act-datasets
SAM2Act
SAM2Act is a multi-view robotics transformer policy for robotic manipulation. Built on RVT-2, it combines multi-resolution upsampling with visual embeddings from the SAM2 foundation model to improve 3D action prediction, multitask learning, and generalization. SAM2Act+ extends this policy with a memory bank, memory encoder, and memory attention so the agent can condition on prior observations and actions for spatial memory-dependent tasks.
For full project details, code… See the full description on the dataset page: https://huggingface.co/datasets/hqfang/sam2act-datasets.epa-re-powering-screening-sam2-masks-partial
EPA RE-Powering Screening SAM2 candidate masks — partial snapshot
This is a paused, incomplete snapshot of model-generated candidate masks for
EPA RE-Powering Screening sites. It covers approximately
12% of archive-group jobs and
8.2% of processable site points from the
current run. It contains 15,211 successful masks across
15,541 attempted processable rows, plus
1,326 explicit input coverage exclusions.
The source inventory contains 190,976 points in total.… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/epa-re-powering-screening-sam2-masks-partial.Arguement_Mining_CL2017tokens along with chunk id. IOB1 format Begining of arguement denoted by B-ARG,inside arguement
denoted by I-ARG, other chunks are O
Orginial train,test split as used by the paper is providedhindi_vgqa_v1.0cc12m-sam2-parse-treeODIA_VYYOVTTSodia-minimind-dataestssegmentation_grounded_dino_sam2_episodes_000248_000267_v1
Task 0003 episodes 248–267 Grounding DINO + SAM2
이 폴더는 248–266의 새 segmentation과 267의 기존 known-good reference를
한곳에서 관리한다.
detector: IDEA-Research/grounding-dino-base
segmenter: facebook/sam2-hiera-large
새 생성 대상: episode 248–266
기존 reference: episode 267
후속 augmentation 실행 상태: blocked_pending_user_segmentation_review
각 episode의 04_qa/representative_contact.png와 최상위 qa/ overview를
검토한 뒤에만 배경 augmentation 단계로 진행한다.
hindi_truthfulqa_gen_mini
Dataset Card for "hindi_truthfulqa_gen_mini"
More Information needed
ODIA_VYYOVTTS_LLAMA3pastis_crop_xregionMental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/Sam20032212/Mental-Health_Text-Classification_Dataset.ASTRA-QA
ASTRA-QA: A Benchmark for Abstract Question Answering over Documents
ASTRA-QA, short for AbSTRAct Question Answering over documents, is a dataset and benchmark for document-level, synthesis-heavy question answering in retrieval-augmented generation systems.
It evaluates whether a system can read long documents, organize evidence, and produce grounded abstractive answers with reference-based assessment, rather than only retrieve short facts.
Dataset Summary
869… See the full description on the dataset page: https://huggingface.co/datasets/sam234990/ASTRA-QA.reasoning-multilingual-R1-indic-trainresume-datasetUVH-26
ArXiv |
Models
Models trained on UVH-26 deliver up to 31.5% higher mAP than COCO-pretrained baselines,
demonstrating significant gains in real-world performance for Indian traffic scenarios.
Dataset Card for UVH-26 (Urban Vision Hackathon Dataset)
Dataset Summary
UVH-26 is a large-scale, India-specific traffic-camera image dataset released by AIM @ IISc for research in intelligent transportation systems and vehicle… See the full description on the dataset page: https://huggingface.co/datasets/Sam20202/UVH-26.openwebtext-10krag-suite-oriyainstruct-reasoning-multilingual-v1"odia_instruct_175k": {
"hf_hub_url": "sam2ai/instruct-reasoning-multilingual-v1",
"split": "train",
"formatting": "sharegpt",
"columns": {
"messages": "conversations"
},
"tags": {
"role_tag": "from",
"content_tag": "value",
"user_tag": "human",
"assistant_tag": "gpt"
}
}
odia-tiny-stories-30kThesimpsons_testSAM2-based-plant-disease-lesion-segmentation-model-DATASETgsm8k-hindi
