arctic
Datasets
All datasets matching “arctic”arctic
Arctic Shift Reddit Archive
Every Reddit comment and submission since 2005, organized as monthly Parquet shards
What is it?
The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02.
Right now the archive has 15.7B items (12.9B comments, 2.8B submissions) in 1.3 TB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly… See the full description on the dataset page: https://huggingface.co/datasets/open-index/arctic.cmu-arctic-xvectors
Speaker embeddings extracted from CMU ARCTIC
There is one .npy file for each utterance in the dataset, 7931 files in total. The speaker embeddings are 512-element X-vectors.
The CMU ARCTIC dataset divides the utterances among the following speakers:
bdl (US male)
slt (US female)
jmk (Canadian male)
awb (Scottish male)
rms (US male)
clb (US female)
ksp (Indian male)
The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model.
Usage:
from… See the full description on the dataset page: https://huggingface.co/datasets/Matthijs/cmu-arctic-xvectors.arctic-camtraj-comparison-20260915
ARCTIC:六段视频与相机轨迹评测审计
在线同步查看 · 新版公开结果目录 · 离线 ZIP
本次新增五段,共六段视频;每格左侧视频 + MANO,右侧点云 + 相机与手腕轨迹。新增当前片段分数、GT/预测相机轨迹、逐采样点位置误差与仅相机引起的世界手腕差异。后者不是真实 world hand MPJPE。
原 36 段因手部 GT 覆盖条件排除了一些相机可评测片段,本次将剩余 27 段单独冻结并实际跑完七个条件。合计 63 段 / 30 条录制 / 单一被试 s05。441 条轨迹已用独立 evo 实现核验,提供录制级置信区间及多重比较检查。
可复算的单人短片段结果,不足以证明能全面替换现有 ViPE。 标定 ViPE / 30 Hz 得到 K + 预测深度及更多输入帧;不是同输入的纯模型比较,也不是完整实测深度生产管线。DA3 / VGGT 原生单位无米制保证,Sim3 校正尺度后误差不能称为原生米制精度。
完整审计报告 · 63 段逐片段指标
原始 36 段与单视频交付目录保留。现有 V9P3R1… See the full description on the dataset page: https://huggingface.co/datasets/yangzijing/arctic-camtraj-comparison-20260915.cmu-arctic-xvectors-extractedmsmarco-v2.1-snowflake-arctic-embed-l
Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods.
Retrieval Performance
Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.arctic
Arctic Shift Reddit Archive
Every Reddit comment and submission since 2005, organized as monthly Parquet shards
What is it?
The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02.
Right now the archive has 1.6B items (362.1M comments, 1.2B submissions) in 181.4 GB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly shards… See the full description on the dataset page: https://huggingface.co/datasets/Dk587/arctic.
