OCT
Datasets
All datasets matching “OCT”OctoSense
New to OctoSense? A getting-started Colab notebook walks through loading a sequence and using each modality. Click Open in Colab above to run it in your browser, no setup required.
OctoSense is a time-synchronized, calibrated, multi-sensor dataset spanning multiple platforms, all sharing the same sensor rig. The bulk of the data is large-scale driving dataset: 371 sequences · 59 hrs · 2,474 km · 8.43 TB of urban, suburban, and rural driving… See the full description on the dataset page: https://huggingface.co/datasets/anthonytec2/OctoSense.OctoNet
OctoNet Multi-Modal Dataset
Welcome to the OctoNet multi-modal dataset! This dataset provides a variety of human activity recordings from multiple sensor modalities, enabling advanced research in activity recognition, pose estimation, multi-modal data fusion, and more.
1. Overview
Name: OctoNet
Data: Multi-modal sensor data, including:
Inertial measurement unit (IMU) data
Motion capture data in CSV and .npy formats
mmWave / Radar data… See the full description on the dataset page: https://huggingface.co/datasets/hku-aiot/OctoNet.corpus-oct-2024
Dataset Card for FreshStack (Corpus)
Homepage |
Repository |
Paper
FreshStack is a holistic framework to construct challenging IR/RAG evaluation datasets that focuses on search across niche and recent topics.
This dataset (October 2024) contains the query, nuggets, answers and nugget-level relevance judgments of 5 niche topics focused on software engineering and machine learning.
The queries and answers (accepted) are taken from Stack Overflow, GPT-4o generates the nuggets and… See the full description on the dataset page: https://huggingface.co/datasets/freshstack/corpus-oct-2024.MegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
The Curation of MegaMath-Web-Pro-Max
Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year;
Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data;
Step 3: Training a fasttext carefully with proper preprocessing;
Step 4: Filtering documents with a threshold (i.e., 0.4);
Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.HiSync
Directory Structure
Published data is organized by collection batch ID.
HiSync_publish/
├── 1/ # Batch ID
│ ├── user1_20250726_132447/ # Sample directory
│ │ ├── cam_1/
│ │ ├── cam_2/
│ │ ├── cam_3/
│ │ ├── person_keypoints.json
│ │ └── meta.json
│ ├── user2_20250726_135244/
│ │ └── ...
│ └── IMU/
│ ├── IMU_Palm/
│ │ └── *.csv
│ ├── IMU_Ring/
│ │ └── *.csv
│ └── IMU_Wrist/
│ └── *.csv
├── 2/
│ └── ...
└──… See the full description on the dataset page: https://huggingface.co/datasets/Octopus1/HiSync.OCT-Longitudinal
Dataset Card for OCT-Longitudinal
Overview
This dataset comprises 1.1 million synthetic OCT images paired with corresponding synthetic longitudinal data, specifically designed for the development and testing of machine learning models in the medical imaging domain.
Dataset Description
The OCT-Longitudinal dataset facilitates the exploration of the informative value of images in predicting longitudinal patient outcomes. A structured latent space created by a… See the full description on the dataset page: https://huggingface.co/datasets/Deltadahl/OCT-Longitudinal.
