mill
Datasets
All datasets matching “mill”milletmillebittwo-million-bluesky-posts
2 Million Bluesky Posts
This dataset contains 2 million public posts collected from Bluesky Social's firehose API, intended for machine learning research and experimentation with social media data.
The with-language-predictions config contains the same data as the default config but with language predictions added using the glotlid model.
Dataset Details
Dataset Description
This dataset consists of 2 million public posts from Bluesky Social, collected through the platform's firehose… See the full description on the dataset page: https://huggingface.co/datasets/alpindale/two-million-bluesky-posts.MUOT_3M-A_3_Million_Frame_Underwater_Object_Tracking_Dataset
🌊 MUOT-3M: The Largest Multimodal Underwater Object Tracking Dataset
Official repository for MUOT-3M📄 MUOT-3M: The Largest Multimodal Underwater Object Tracking Dataset and MUTrack Tracking Method
🚀 Overview
MUOT-3M is currently the largest underwater object tracking dataset, containing over 3 million annotated frames across 3,030 underwater videos with synchronized multimodal annotations.
The benchmark is designed to advance research in:
Underwater object tracking… See the full description on the dataset page: https://huggingface.co/datasets/AhsanBB/MUOT_3M-A_3_Million_Frame_Underwater_Object_Tracking_Dataset.Voila-million-voice
Voila: Voice-Language Foundation Models
💜 Project Page | 🖥️ GitHub | 🤗 Hugging Face | 📑 Paper | 🌐 Online Demo | 🏠Maitrix.org
Voila is a new family of large voice-language foundation models aiming to lift human-AI interaction experiences to the next level. Breaking away from the constraints of traditional voice AI systems—high latency, loss of vocal nuances, and mechanical responses—Voila employs an innovative end-to-end model design and a novel… See the full description on the dataset page: https://huggingface.co/datasets/maitrix-org/Voila-million-voice.one-million-commits
One million commits
A large variety of git commits pulled from across GitHub.
Created by William Entriken, released 2023-09-26, version 1.
This composition is licensed under the MIT license.
Intended use
This dataset could be used to train a model concerned with programming tasks:
Summarize some programming work
Perform work given a description of the work to do
Learn-by-example the syntax for all active programming languages and structured data formats
This… See the full description on the dataset page: https://huggingface.co/datasets/fulldecent/one-million-commits.
