ind
Datasets
All datasets matching “ind”arctic
Arctic Shift Reddit Archive
Every Reddit comment and submission since 2005, organized as monthly Parquet shards
What is it?
The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02.
Right now the archive has 15.7B items (12.9B comments, 2.8B submissions) in 1.3 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly… See the full description on the dataset page: https://huggingface.co/datasets/open-index/arctic.character_index
Anime Character Index
This dataset if for collecting all the hot characters from the internet, and extract their features and core tags. It will be useful for automatically testing the character generating ability of the anime-style base models.
7371 characters in total.
Copyrights
Copyright
Count
kantai_collection
393
pokemon
380
fate_(series)
350
hololive
277
blue_archive234
arknights
200
idolmaster
192
touhou
186
fire_emblem
168
umamusume… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/character_index.hacker-news
Hacker News - Complete Archive
Every Hacker News item since 2006, live-updated every 5 minutes
What is it?
This dataset contains the complete Hacker News archive: every story, comment, Ask HN, Show HN, job posting, and poll ever submitted to the site. Hacker News is one of the longest-running and most influential technology communities on the internet, operated by Y Combinator since 2007. It has become the de facto gathering place for founders, engineers… See the full description on the dataset page: https://huggingface.co/datasets/open-index/hacker-news.IndEgo
IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants
Vivek Chavan¹²*, Yasmina Imgrund²†, Tung Dao²†, Sanwantri Bai³†, Bosong Wang⁴†, Ze Lu⁵†, Oliver Heimann¹, Jörg Krüger¹²
¹Fraunhofer IPK, Berlin ²Technical University of Berlin ³University of Tübingen
⁴RWTH Aachen University ⁵Leibniz University Hannover
*Project Lead †Work done during student theses/projects at Fraunhofer IPK… See the full description on the dataset page: https://huggingface.co/datasets/FraunhoferIPK/IndEgo.fineweb-edu-indexThis dataset contains the embeddings for the full fineweb-edu, embedded with the Cohere Embed V3 model.
You can search on this dataset with just 500MB of memory using DiskVectorIndex.
Installation & Usage
Get your free Cohere API key from cohere.com. You must set this API key as an environment variable:
export COHERE_API_KEY=your_api_key
Install the package:
pip install DiskVectorIndex
You can then search via:
from DiskVectorIndex import DiskVectorIndex
index =… See the full description on the dataset page: https://huggingface.co/datasets/fineweb-retrieval/fineweb-edu-index.IndicVoices
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Updates
[23 December 2025] We now have 11,200 hours of transcribed data! 🎉
Overview
INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.



