datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DFToolBench-A-500
DFToolBench-A-500
A 500-query benchmark for evaluating audio tool-use agents on deepfake-related forensic tasks.
Each query is a multi-turn ReAct-style dialog in which an assistant invokes audio analysis tools
(e.g. speaker_verification, nisqa, silero_vad, deepfake_audio, language_id, muq,
calculator) over a single audio file and produces a final verdict.
Files
dataset.json — pretty-printed list of 500 records.
dataset.jsonl — one record per line (preferred for… See the full description on the dataset page: https://huggingface.co/datasets/hardiksharma6555/DFToolBench-A-500.dftoolbench-a500
DFToolBench-A-500
A 500-query benchmark for evaluating audio tool-use agents on deepfake-related forensic tasks.
Each query is a multi-turn ReAct-style dialog in which an assistant invokes audio analysis tools
(e.g. speaker_verification, nisqa, silero_vad, deepfake_audio, language_id, muq,
calculator) over a single audio file and produces a final verdict.
Files
dataset.json — pretty-printed list of 500 records.
dataset.jsonl — one record per line (preferred for… See the full description on the dataset page: https://huggingface.co/datasets/dftoolbench/dftoolbench-a500.
