SoMe
Datasets
All datasets matching “SoMe”something-something-v2Slim-binidxThis dataset comprises nine chunks (out of ten) from the Cerebras/SlimPajama-627B dataset, processed into a binary index (bin idx) format.
The first chunk is located at : rwkv-x-dev/slimpajama-binidx.
Due to their large size, each chunk is split into multiple parts for easier handling. To reassemble and decompress these parts, follow these steps:
Combine all parts of the desired chunk into a single file:
cat chunk2_text_document_part_* > chunk2_text_document.tar.xz
Decompress the combined… See the full description on the dataset page: https://huggingface.co/datasets/something-else/Slim-binidx.ac-transit-apc
AC Transit Automatic Passenger Counter Records, 2019-2026
Stop-level boarding and alighting counts for the AC Transit bus network in
Alameda and Contra Costa counties, California, from January 2019 through
May 2026. The records come from the automatic passenger counters (APCs)
mounted at the doors of the buses: one row per stop event, with the number of
passengers who got on, the number who got off, and the load the bus left with.
89 monthly Parquet files, ~5.9 GB, partitioned… See the full description on the dataset page: https://huggingface.co/datasets/somemone/ac-transit-apc.fix_some_err
df_eval
Public evaluation-only speech deepfake detection dataset, organized like Common Voice language configs:
each config is a standard eval protocol (ASVspoof, ADD, In-the-Wild, …) with embedded audio.
Companion code: github.com/Shuo-H/df_eval
Configs are added as uploads complete. Declared configs below match currently available parquet shards on the Hub.
Load
from datasets import load_dataset
ds = load_dataset("shuohann/df_eval", name="sonar"… See the full description on the dataset page: https://huggingface.co/datasets/shuohann/fix_some_err.dataclysm-wikipedia
somewheresystems/dataclysm-wikipedia
USE THE NOTEBOOK TO GET STARTED!
https://github.com/somewheresystems/dataclysm
This dataset comprises of 6,458,670 English language Wikipedia articles, with an additional column added for title-embeddings using the bge-small-en-v1.5 embeddings model. The dataset was sourced here: https://huggingface.co/datasets/wikipedia/viewer/20220301.en
This dataset contains the full text of each Wikipedia article as of the date March 01, 2022. In… See the full description on the dataset page: https://huggingface.co/datasets/somewheresystems/dataclysm-wikipedia.something_something_v2_lerobot
🤖 Custom LeRobot Dataset
This dataset was created using LeRobot,and follows the LeRobot v2.1 dataset specification for robotic control and imitation learning.
📘 Dataset Description
Codebase version: v2.1
Robot type: ``
License: Apache-2.0
Total episodes: 193690
Total frames: 44352895
Total videos: 193690
Frames per second (fps): 60
Video resolution: 256 × 256 (RGB)
Splits:
train: 0–40
File sizes:
data_files_size_in_mb: * MB (approx)*… See the full description on the dataset page: https://huggingface.co/datasets/JUNTAO123/something_something_v2_lerobot.
