v1.5
Datasets
All datasets matching “v1.5”cad-recode-v1.5
CAD-Recode: Reverse Engineering CAD Code from Point Clouds
CAD-Recode dataset is provided in form of Python (CadQuery) codes.
Train size is ~1M and validation size is ~1k.
CAD-Recode model and code are released at github https://github.com/filaPro/cad-recode.
And if you like it, give us a github 🌟.
Citation
If you find this work useful for your research, please cite our paper:
@misc{rukhovich2024cadrecode,
title={CAD-Recode: Reverse Engineering CAD Code from Point… See the full description on the dataset page: https://huggingface.co/datasets/filapro/cad-recode-v1.5.wikipedia-bge-small-en-v1.5-fullKoLLaVA-v1.5-Instruct-581k
KoLLaVA-v1.5-Instruct-581k
한국어 Vision-Language 모델을 위한 instruction tuning 데이터셋입니다.
데이터셋 정보
총 샘플 수: 435,093개
형식: ChatML 형식 (role: user/assistant, content: 텍스트)
이미지: COCO + GQA + Visual Genome 데이터셋
언어: 한국어
포함된 데이터셋
COCO 데이터: 362,953개 샘플
MS COCO 2017 이미지 기반
한국어 대화 데이터
GQA 데이터: 72,140개 샘플
GQA (Visual Question Answering) 이미지 기반
한국어 대화 데이터
Visual Genome 데이터: 포함
Visual Genome 이미지 기반
한국어 대화 데이터
제외된 데이터셋
EKVQA 데이터: AI Hub 라이선스로 인해 공개 불가… See the full description on the dataset page: https://huggingface.co/datasets/ko-vlm/KoLLaVA-v1.5-Instruct-581k.mteb-retrieval-snowflake-arctic-embed-m-v1.5arxiv_20240801_gte-base-en-v1.5_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked arxiv papers from Semantic Scholar. The embedding model used is Alibaba-NLP/gte-large-en-v1.5.
This index is compatible with WikiChat v2.0.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/arxiv_20240801_gte-base-en-v1.5_qdrant_index.msmarco-v2.1-snowflake-arctic-embed-m-v1.5
Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.
