chunked
Datasets
All datasets matching “chunked”smollm-chunked
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora
This repository contains part of the FAISS indices and chunked datasets used for novelty detection for SmolLM and SmolLM2, as presented in the paper LLM generation novelty through the lens of semantic similarity.
Full Documentation
For complete usage instructions, installation guide, and tutorial, please refer to:
Main Tutorial README
Data Distribution
Due to Hugging Face storage quota… See the full description on the dataset page: https://huggingface.co/datasets/enguyen/smollm-chunked.dl3dv_chunked
DL3DV Post-processed for Less3Depend
This dataset is a post-processed version of the DL3DV-10K dataset, specifically prepared for the repository 👉 Less3Depend.
The goal of this release is to provide a clean, unified, and research-ready variant of DL3DV that is directly usable in 👉 PixelSplat style.
Acknowledgement
If you find this dataset useful in your research, please consider citing:
1️⃣ DL3DV Original Dataset:
@inproceedings{ling2024dl3dv,
title={Dl3dv-10k: A… See the full description on the dataset page: https://huggingface.co/datasets/littlekoyo/dl3dv_chunked.xd-violence-rgb-videomae-chunked-testIndian-Supreme-Court-Judgements-Chunked
Indian Supreme Court Judgements Chunked
Executive Summary
The dataset aims to address the chronic backlog in the Indian judiciary system, particularly in the Supreme Court, by creating a dataset optimized for legal language models (LLMs). The dataset will consist of pre-processed, chunked, and embedded textual data derived from the Supreme Court's judgment PDFs.
Problem and Importance - Motivation
Indian courts are overwhelmed with pending cases, with the… See the full description on the dataset page: https://huggingface.co/datasets/vihaannnn/Indian-Supreme-Court-Judgements-Chunked.nsw-caselaw-chunkedwiki-chunked-mxbai-embed-large-v1
