willychan21/ParallelKernelBench_Problems
ParallelKernelBench (benchmark) Reference problems for ParallelKernelBench: a benchmark for LLM-generated multi-GPU CUDA kernels. This dataset contains 87 reference implementations in reference/ and the input tensor specification in utils/input_output_tensors.py. Files Path Description data/problems.parquet One row per problem (tabular access) reference/*.py Reference solution() implementations utils/input_output_tensors.py Input/output tensor… See the full description on the dataset page: https://huggingface.co/datasets/willychan21/ParallelKernelBench_Problems.
0266
1import torch2import torch.distributed as dist3 4 5@torch.no_grad()6def solution(tensor: torch.Tensor, src: int = 0) -> torch.Tensor:7 rank = dist.get_rank()8 9 if rank == src:10 world_size = dist.get_world_size()11 scatter_list = [chunk.squeeze(0).contiguous() for chunk in tensor.chunk(world_size, dim=0)]12 out = torch.empty_like(scatter_list[0])13 else:14 scatter_list = None15 out = tensor.clone()16 17 dist.scatter(out, scatter_list=scatter_list, src=src)18 return out19 