Let me give a brief introduction!
This is triple-mu, 🤣 a fix typo programmer based in Beijing, China.
I am participating in mmyolo yolov6 yolov7 yolort projects as a co-author.
You can reach me in the following ways:
QQ: 3394101
Email: gpu@163.com
Ulysses sequence-parallel all-to-all as a torch custom op, moved by the GPU copy engines over NVSHMEM symmetric memory. Zero SM usage; 1.7-2.2x over torch.distributed on NVLink.
Forked from ultralytics/ultralytics
YOLOv8 🚀 in PyTorch > ONNX > CoreML > TFLite
Qwen-Image's DiT inference with TensorRT-10