gemini-2.0-flash
gemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.Gemini-2.0-Flash-Aoede-VoiceGemini-2.0-Flash-Fenrir-VoiceGemini-2.0-Flash-Kore-Voicesarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600Dataset for sarvam's entity normalisation task. More detailed information can be found here, in the main model repo: Hugging Face
Detailed Report (Writeup): Google Drive
It also has a gguf variant, with certain additional gguf based innstructions: Hugging Face
Model inference script can be found here: Colab
Model predictions can be found in this dataset and both the repo files. named as:
eval_data_001_predictions.csv and eval_data_001_predictions_excel.csv.
train_data_001_predictions.csvand… See the full description on the dataset page: https://huggingface.co/datasets/Tasmay-Tib/sarvam-entity-recognition-gemini-2.0-flash-thinking-01-21-distill-1600.dummy-ioi-eval-openrouter_google_gemini-2.0-flash-thinking-exp_free-test
