som
Datasets
All datasets matching “som”Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
something-something-v2CICIoT2023Small
CICIoT2023
This dataset provides a processed derivative of the CICIoT2023 traffic collection. The repository organizes truncated PCAP files and flow-based CSV extractions aligned to the original CICResearch folder hierarchy.
Processing Workflow
The processing pipeline follows four stages:
Source acquisition from the CICResearch CICIoT2023 portal.
Flow extraction from full PCAP files using TriFlowMeter.
PCAP size reduction by truncating packet payloads to 128 bytes with… See the full description on the dataset page: https://huggingface.co/datasets/somnath0100/CICIoT2023Small.vr_m3_soma_retargeteSlim-binidxThis dataset comprises nine chunks (out of ten) from the Cerebras/SlimPajama-627B dataset, processed into a binary index (bin idx) format.
The first chunk is located at : rwkv-x-dev/slimpajama-binidx.
Due to their large size, each chunk is split into multiple parts for easier handling. To reassemble and decompress these parts, follow these steps:
Combine all parts of the desired chunk into a single file:
cat chunk2_text_document_part_* > chunk2_text_document.tar.xz
Decompress the combined… See the full description on the dataset page: https://huggingface.co/datasets/something-else/Slim-binidx.fix_some_err
df_eval
Public evaluation-only speech deepfake detection dataset, organized like Common Voice language configs:
each config is a standard eval protocol (ASVspoof, ADD, In-the-Wild, …) with embedded audio.
Companion code: github.com/Shuo-H/df_eval
Configs are added as uploads complete. Declared configs below match currently available parquet shards on the Hub.
Load
from datasets import load_dataset
ds = load_dataset("shuohann/df_eval", name="sonar"… See the full description on the dataset page: https://huggingface.co/datasets/shuohann/fix_some_err.


