datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-cpp
Dataset Card for Synthetic C++ Dataset
Dataset Description
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [---
Dataset Card for Synthetic C++ Dataset
Dataset Description
Homepage: [https://huggingface.co/datasets/ReySajju742/synthetic-cpp/]
Point of Contact: [ReySajju742]
Dataset Summary
This dataset contains 10,000 rows of synthetically generated data focusing on the topic of "C++… See the full description on the dataset page: https://huggingface.co/datasets/ReySajju742/synthetic-cpp.winogrande-eval-for-llama.cppWinogrande evaluation dataset for llama.cpp
apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.CPPB
CPPB
Summary
CPPB is the public release surface for the Controlled Prompt-Privacy Benchmark introduced in BodhiPromptShield: Pre-Inference Prompt Mediation for Suppressing Privacy Propagation in LLM/VLM Agents.
This Hugging Face package intentionally releases the benchmark-authored prompt manifest and template-stratified train/dev/test split, not raw third-party prompts, source images, or end-to-end OCR assets. Each row is a controlled prompt stub with benchmark metadata… See the full description on the dataset page: https://huggingface.co/datasets/mabo1215/CPPB.oop-bad-code-to-good-code-cppCPP-Code-Solutions
C++ Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
C++ Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
code_search_net_kotlin_cpp
Dataset Summary
This dataset was converted based on code_search_net (https://huggingface.co/datasets/code-search-net/code_search_net)
Languages
Kotlin programming language
C++ programming language
Data Instances
A data point consists of a function code along with its documentation. Each data point also contains meta data on the function, such as the repository it was extracted from.
{
'id': '0',
'repository_name': 'organisation/repository'… See the full description on the dataset page: https://huggingface.co/datasets/namyi/code_search_net_kotlin_cpp.llama.cpp_jetson_benchmark
llama.cpp with CUDA support on a Jetson Nano 2019
This dataset contains some benchmark results for llama.cpp with GPU acceleration on the Jetson Nano from 2019.
For comparison llama.cpp was compiled with just CPU support, and a recent ollama version received the same 11 questions.
cppvulllama-cpp-server-doc-qacpp-mgc-dataset
