datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantum-compilation-and-programming
Neura Parse — Quantum Compilation & Programming
A code-heavy vertical on the quantum software/compilation stack: turning abstract quantum circuits and unitaries into device-executable programs. Covers unitary decomposition and circuit synthesis (Euler/ZYZ, KAK/Cartan, Solovay-Kitaev, Ross-Selinger gridsynth, numerical synthesis with BQSKit), gate-set/basis transpilation to native gate sets, qubit layout/mapping and routing under connectivity constraints (SABRE, VF2, SWAP… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-compilation-and-programming.mojo-programming-language-qnaA synthetic dataset from Claude Sonnet 3.5. The source documents are real pulled from the Mojo documentation, but everything else is synthetic.
ai-programming-cookbook
AI Programming Recipes Dataset
A Q&A dataset aimed at providing in-depth recipes on building complex AI systems for LLM fine-tuning. All the original data comes from open-source documentations.
Preprocessing
Multiple preprocessing steps were used to make the dataset ready for LoRa fine-tuning:
Convert the ipynb/rst files to markdown as it will be our output format. This was done thanks to nbconvert/pydantic.
Use an LLM (DeepSeek V3.1) to unclutter the output cells of… See the full description on the dataset page: https://huggingface.co/datasets/paulprt/ai-programming-cookbook.beir_cqadupstack_programmers_test
beir_cqadupstack_programmers_test
BEIR CQADupStack/programmers test split
Field
Value
Benchmark
beir
Sub-benchmark
cqadupstack_programmers
Type
retrieval
Items
876
Exported from Langfuse.
