René's URL Explorer Experiment


Title: Cpp 확장을 사용한 프로세스 그룹 백엔드 사용자 정의 — 파이토치 한국어 튜토리얼 (PyTorch tutorials in Korean)

Open Graph Title: Cpp 확장을 사용한 프로세스 그룹 백엔드 사용자 정의

Description: Author: Howard Huang, Feng Tian, Shen Li, Min Si, 번역: 박재윤,. 선수과목(Prerequisites): PyTorch Distributed Overview, PyTorch Collective Communication Package, PyTorch Cpp Extension, Writing Distributed Applications with PyTorch. 이 튜토리얼에서는 cpp 확장 을 사용하여 사용자 정의 Backend 를 구현하고 이를 파이토치 분산 패키지 에 어떻게 연결하는지를 ...

Open Graph Description: Author: Howard Huang, Feng Tian, Shen Li, Min Si, 번역: 박재윤,. 선수과목(Prerequisites): PyTorch Distributed Overview, PyTorch Collective Communication Package, PyTorch Cpp Extension, Writing Distributed Applications with PyTorch. 이 튜토리얼에서는 cpp 확장 을 사용하여 사용자 정의 Backend 를 구현하고 이를 파이토치 분산 패키지 에 어떻게 연결하는지를 ...

Opengraph URL: https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html

direct link

Domain: tutorials.pytorch.kr


Hey, it has json ld scripts:
    {
       "@context": "https://schema.org",
       "@type": "Article",
       "name": "Cpp \ud655\uc7a5\uc744 \uc0ac\uc6a9\ud55c \ud504\ub85c\uc138\uc2a4 \uadf8\ub8f9 \ubc31\uc5d4\ub4dc \uc0ac\uc6a9\uc790 \uc815\uc758",
       "headline": "Cpp \ud655\uc7a5\uc744 \uc0ac\uc6a9\ud55c \ud504\ub85c\uc138\uc2a4 \uadf8\ub8f9 \ubc31\uc5d4\ub4dc \uc0ac\uc6a9\uc790 \uc815\uc758",
       "description": "PyTorch Documentation. Explore PyTorch, an open-source machine learning library that accelerates the path from research prototyping to production deployment. Discover tutorials, API references, and guides to help you build and deploy deep learning models efficiently.",
       "url": "/intermediate/process_group_cpp_extension_tutorial.html",
       "articleBody": "Cpp \ud655\uc7a5\uc744 \uc0ac\uc6a9\ud55c \ud504\ub85c\uc138\uc2a4 \uadf8\ub8f9 \ubc31\uc5d4\ub4dc \uc0ac\uc6a9\uc790 \uc815\uc758# Author: Howard Huang, Feng Tian, Shen Li, Min Si\ubc88\uc5ed: \ubc15\uc7ac\uc724 \ucc38\uace0 \uc774 \ud29c\ud1a0\ub9ac\uc5bc\uc758 \uc18c\uc2a4 \ucf54\ub4dc\ub294 github \uc5d0\uc11c \ud655\uc778\ud558\uace0 \ubcc0\uacbd\ud574 \ubcfc \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uc120\uc218\uacfc\ubaa9(Prerequisites): PyTorch Distributed Overview PyTorch Collective Communication Package PyTorch Cpp Extension Writing Distributed Applications with PyTorch \uc774 \ud29c\ud1a0\ub9ac\uc5bc\uc5d0\uc11c\ub294 cpp \ud655\uc7a5 \uc744 \uc0ac\uc6a9\ud558\uc5ec \uc0ac\uc6a9\uc790 \uc815\uc758 Backend \ub97c \uad6c\ud604\ud558\uace0 \uc774\ub97c \ud30c\uc774\ud1a0\uce58 \ubd84\uc0b0 \ud328\ud0a4\uc9c0 \uc5d0 \uc5b4\ub5bb\uac8c \uc5f0\uacb0\ud558\ub294\uc9c0\ub97c \uc54c\uc544\ubd05\ub2c8\ub2e4. \uc774\ub7ec\ud55c \ubc29\ubc95\uc740 \ud558\ub4dc\uc6e8\uc5b4\uc5d0 \ud2b9\ud654\ub41c \uc18c\ud504\ud2b8\uc6e8\uc5b4 \uc2a4\ud0dd\uc774 \ud544\uc694\ud55c \uacbd\uc6b0\ub098 \uc0c8\ub85c\uc6b4 \uc9d1\ud569 \ud1b5\uc2e0 \uc54c\uace0\ub9ac\uc998(collective communication algorithm)\uc744 \uc2e4\ud5d8\ud558\uace0\uc790 \ud560 \ub54c \uc720\uc6a9\ud569\ub2c8\ub2e4. \uae30\ucd08# \ud30c\uc774\ud1a0\uce58(PyTorch)\uc758 \uc9d1\ud569 \ud1b5\uc2e0(collective communications)\uc740 \ubd84\uc0b0 \ub370\uc774\ud130 \ubcd1\ub82c(DistributedDataParallel) \ubc0f \uc81c\ub85c \ub9ac\ub358\ub358\uc2dc \ucd5c\uc801\ud654\uae30(ZeroRedundancyOptimizer) \ub4f1\uc744 \ud3ec\ud568\ud558\uc5ec, \ub110\ub9ac \uc0ac\uc6a9\ub418\ub294 \ubd84\uc0b0 \ud559\uc2b5 \uae30\ub2a5\uc744 \uc9c0\uc6d0\ud569\ub2c8\ub2e4. \ub3d9\uc77c\ud55c \uc9d1\ud569 \ud1b5\uc2e0 API\ub97c \ub2e4\uc591\ud55c \ud1b5\uc2e0 \ubc31\uc5d4\ub4dc\uc5d0\uc11c \uc791\ub3d9\ud558\ub3c4\ub85d \ud558\uae30 \uc704\ud574 \ubd84\uc0b0 \ud328\ud0a4\uc9c0\ub294 \uc9d1\ud569 \ud1b5\uc2e0 \uc791\uc5c5\uc744 Backend \ud074\ub798\uc2a4\ub85c \ucd94\uc0c1\ud654\ud569\ub2c8\ub2e4. \uc774\ud6c4\uc5d0\ub294 \uc6d0\ud558\ub294 \uc11c\ub4dc\ud30c\ud2f0 \ub77c\uc774\ube0c\ub7ec\ub9ac\ub97c \uc0ac\uc6a9\ud558\uc5ec Backend \uc758 \ud558\uc704 \ud074\ub798\uc2a4(subclass)\ub85c \ub2e4\uc591\ud55c \ubc31\uc5d4\ub4dc\ub97c \uad6c\ud604\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \ud30c\uc774\ud1a0\uce58 \ubd84\uc0b0(PyTorch distributed)\uc5d0\ub294 \uc138 \uac00\uc9c0 \uae30\ubcf8 \ubc31\uc5d4\ub4dc\uc778 ProcessGroupNCCL, ProcessGroupGloo, \uadf8\ub9ac\uace0 ProcessGroupMPI \uac00 \ud3ec\ud568\ub418\uc5b4 \uc788\uc2b5\ub2c8\ub2e4. \uadf8\ub7ec\ub098 \uc774 \uc138 \uac00\uc9c0 \ubc31\uc5d4\ub4dc \uc678\uc5d0\ub3c4 \ub2e4\ub978 \ud1b5\uc2e0 \ub77c\uc774\ube0c\ub7ec\ub9ac(\uc608: UCC, OneCCL), \ub2e4\ub978 \uc720\ud615\uc758 \ud558\ub4dc\uc6e8\uc5b4(\uc608: TPU, Trainum), \uadf8\ub9ac\uace0 \uc0c8\ub85c\uc6b4 \ud1b5\uc2e0 \uc54c\uace0\ub9ac\uc998(\uc608: Herring, Reduction Server)\ub3c4 \uc788\uc2b5\ub2c8\ub2e4. \ub530\ub77c\uc11c \ubd84\uc0b0 \ud328\ud0a4\uc9c0\ub294 \uc9d1\ud569 \ud1b5\uc2e0 \ubc31\uc5d4\ub4dc\ub97c \uc0ac\uc6a9\uc790 \uc9c0\uc815\ud560 \uc218 \uc788\ub3c4\ub85d \ud655\uc7a5 API\ub97c \ub178\ucd9c\ud569\ub2c8\ub2e4. \uc544\ub798\uc758 4\ub2e8\uacc4\ub294 \uac00\uc9dc(dunmmy) ProcessGroup \ubc31\uc5d4\ub4dc\ub97c \uad6c\ud604\ud558\uace0 \ud30c\uc774\uc36c \uc751\uc6a9 \ud504\ub85c\uadf8\ub7a8 \ucf54\ub4dc\uc5d0\uc11c \uc0ac\uc6a9\ud558\ub294 \ubc29\ubc95\uc744 \ubcf4\uc5ec\uc90d\ub2c8\ub2e4. \uc774 \ud29c\ud1a0\ub9ac\uc5bc\uc740 \uc791\ub3d9\ud558\ub294 \ud1b5\uc2e0 \ubc31\uc5d4\ub4dc\ub97c \uac1c\ubc1c\ud558\ub294 \ub300\uc2e0 \ud655\uc7a5 API\ub97c \uc124\uba85\ud558\ub294 \ub370 \uc911\uc810\uc744 \ub461\ub2c8\ub2e4. \ub530\ub77c\uc11c dummy \ubc31\uc5d4\ub4dc\ub294 API\uc758 \uc77c\ubd80 (all_reduce \ubc0f all_gather)\ub97c \ub2e4\ub8e8\uba70 tensor\uc758 \uac12\uc744 \ub2e8\uc21c\ud788 0\uc73c\ub85c \uc124\uc815\ud569\ub2c8\ub2e4. \ub2e8\uacc4 1: Backend \uc758 \ud558\uc704 \ud074\ub798\uc2a4 \uad6c\ud604# \uccab \ubc88\uc9f8 \ub2e8\uacc4\ub294 \ub300\uc0c1 \uc9d1\ud569 \ud1b5\uc2e0 API\ub97c \uc7ac\uc815\uc758\ud558\uace0 \uc0ac\uc6a9\uc790 \uc815\uc758 \ud1b5\uc2e0 \uc54c\uace0\ub9ac\uc998\uc744 \uc2e4\ud589\ud558\ub294 Backend \ud558\uc704 \ud074\ub798\uc2a4\ub97c \uad6c\ud604\ud558\ub294 \uac83\uc785\ub2c8\ub2e4. \ud655\uc7a5 \uae30\ub2a5\uc740 \ubbf8\ub798(future) \ud1b5\uc2e0 \uacb0\uacfc\ub97c \uc81c\uacf5\ud558\ub294 Work \ud558\uc704 \ud074\ub798\uc2a4\ub97c \uad6c\ud604\ud574\uc57c \ud558\uba70, \uc774\ub294 \uc751\uc6a9 \ud504\ub85c\uadf8\ub7a8 \ucf54\ub4dc\uc5d0\uc11c \ube44\ub3d9\uae30 \uc2e4\ud589\uc744 \ud5c8\uc6a9\ud569\ub2c8\ub2e4. \ud655\uc7a5 \uae30\ub2a5\uc774 \uc11c\ub4dc\ud30c\ud2f0 \ub77c\uc774\ube0c\ub7ec\ub9ac\ub97c \uc0ac\uc6a9\ud558\ub294 \uacbd\uc6b0, \ud574\ub2f9 \ud655\uc7a5 \uae30\ub2a5\uc740 BackendDemmy \ud558\uc704 \ud074\ub798\uc2a4\uc5d0\uc11c \ud5e4\ub354\ub97c \ud3ec\ud568\ud558\uace0 \ub77c\uc774\ube0c\ub7ec\ub9ac API\ub97c \ud638\ucd9c\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uc544\ub798\uc758 \ub450 \ucf54\ub4dc\ub294 dummy.h \ubc0f dummy.cpp \uc758 \uad6c\ud604\uc744 \ubcf4\uc5ec\uc90d\ub2c8\ub2e4. \uc804\uccb4 \uad6c\ud604\uc740 \ub354\ubbf8 \uc9d1\ud569(dummy collectives) \uc800\uc7a5\uc18c\uc5d0\uc11c \ud655\uc778\ud558\uc2e4 \uc218 \uc788\uc2b5\ub2c8\ub2e4. // \ud30c\uc77c \uc774\ub984: dummy.hpp #include \u003ctorch/python.h\u003e #include \u003ctorch/csrc/distributed/c10d/Backend.hpp\u003e #include \u003ctorch/csrc/distributed/c10d/Work.hpp\u003e #include \u003ctorch/csrc/distributed/c10d/Store.hpp\u003e #include \u003ctorch/csrc/distributed/c10d/Types.hpp\u003e #include \u003ctorch/csrc/distributed/c10d/Utils.hpp\u003e #include \u003cpybind11/chrono.h\u003e namespace c10d { class BackendDummy : public Backend { public: BackendDummy(int rank, int size); c10::intrusive_ptr\u003cWork\u003e allgather( std::vector\u003cstd::vector\u003cat::Tensor\u003e\u003e\u0026 outputTensors, std::vector\u003cat::Tensor\u003e\u0026 inputTensors, const AllgatherOptions\u0026 opts = AllgatherOptions()) override; c10::intrusive_ptr\u003cWork\u003e allreduce( std::vector\u003cat::Tensor\u003e\u0026 tensors, const AllreduceOptions\u0026 opts = AllreduceOptions()) override; // \uc0ac\uc6a9\uc790 \uc815\uc758 \uad6c\ud604\uc774 \uc5c6\ub294 \uc0c1\ud0dc\uc5d0\uc11c\uc758 \uc9d1\ud569 \ud1b5\uc2e0 API\ub294 // \uc751\uc6a9 \ud504\ub85c\uadf8\ub7a8 \ucf54\ub4dc\uc5d0\uc11c \ud638\ucd9c\ub418\uba74 \uc624\ub958\uac00 \ubc1c\uc0dd\ud569\ub2c8\ub2e4. }; class WorkDummy : public Work { public: WorkDummy( OpType opType, c10::intrusive_ptr\u003cc10::ivalue::Future\u003e future) // future of the output : Work( -1, // \ub7ad\ud06c, recvAnySource\uc5d0\uc11c\ub9cc \uc0ac\uc6a9\ub418\uba70 \uc774 \ub370\ubaa8\uc5d0\uc11c\ub294 \uad00\ub828\uc774 \uc5c6\uc2b5\ub2c8\ub2e4. opType), future_(std::move(future)) {} bool isCompleted() override; bool isSuccess() const override; bool wait(std::chrono::milliseconds timeout = kUnsetTimeout) override; virtual c10::intrusive_ptr\u003cc10::ivalue::Future\u003e getFuture() override; private: c10::intrusive_ptr\u003cc10::ivalue::Future\u003e future_; }; } // namespace c10d // \ud30c\uc77c \uc774\ub984: dummy.cpp #include \"dummy.hpp\" namespace c10d { // \uc774\uac83\uc740 \ubaa8\ub4e0 \ucd9c\ub825 tensor\ub97c 0\uc73c\ub85c \uc124\uc815\ud558\ub294 \uac00\uc9dc(dummy) allgather\uc785\ub2c8\ub2e4. // \uc2e4\uc81c \ud1b5\uc2e0\uc744 \ube44\ub3d9\uae30\uc801\uc73c\ub85c \uc218\ud589\ud558\ub3c4\ub85d \uad6c\ud604\uc744 \uc218\uc815\ud558\uc138\uc694. c10::intrusive_ptr\u003cWork\u003e BackendDummy::allgather( std::vector\u003cstd::vector\u003cat::Tensor\u003e\u003e\u0026 outputTensors, std::vector\u003cat::Tensor\u003e\u0026 inputTensors, const AllgatherOptions\u0026 /* unused */) { for (auto\u0026 outputTensorVec : outputTensors) { for (auto\u0026 outputTensor : outputTensorVec) { outputTensor.zero_(); } } auto future = c10::make_intrusive\u003cc10::ivalue::Future\u003e( c10::ListType::create(c10::ListType::create(c10::TensorType::get()))); future-\u003emarkCompleted(c10::IValue(outputTensors)); return c10::make_intrusive\u003cWorkDummy\u003e(OpType::ALLGATHER, std::move(future)); } // \uc774\uac83\uc740 \ubaa8\ub4e0 \ucd9c\ub825 tensor\ub97c 0\uc73c\ub85c \uc124\uc815\ud558\ub294 \uac00\uc9dc(dummy) allreduce\uc785\ub2c8\ub2e4. // \uc2e4\uc81c \ud1b5\uc2e0\uc744 \ube44\ub3d9\uae30\uc801\uc73c\ub85c \uc218\ud589\ud558\ub3c4\ub85d \uad6c\ud604\uc744 \uc218\uc815\ud558\uc138\uc694. c10::intrusive_ptr\u003cWork\u003e BackendDummy::allreduce( std::vector\u003cat::Tensor\u003e\u0026 tensors, const AllreduceOptions\u0026 opts) { for (auto\u0026 tensor : tensors) { tensor.zero_(); } auto future = c10::make_intrusive\u003cc10::ivalue::Future\u003e( c10::ListType::create(c10::TensorType::get())); future-\u003emarkCompleted(c10::IValue(tensors)); return c10::make_intrusive\u003cWorkDummy\u003e(OpType::ALLGATHER, std::move(future)); } } // namespace c10d \ub2e8\uacc4 2: \ud655\uc7a5 \uae30\ub2a5\uc744 \ud30c\uc774\uc36c API\ub85c \ub178\ucd9c# \ubc31\uc5d4\ub4dc \uc0dd\uc131\uc790\ub294 \ud30c\uc774\uc36c \uce21 \uc5d0\uc11c \ud638\ucd9c\ub418\ubbc0\ub85c \ud655\uc7a5 \uae30\ub2a5\ub3c4 \ud30c\uc774\uc36c\uc5d0 \uc0dd\uc131\uc790 API\ub97c \ub178\ucd9c\ud574\uc57c \ud569\ub2c8\ub2e4. \ub2e4\uc74c \uba54\uc11c\ub4dc\ub97c \ucd94\uac00\ud568\uc73c\ub85c\uc368 \uc774 \uc791\uc5c5\uc744 \uc218\ud589\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uc774 \uc608\uc81c\uc5d0\uc11c\ub294 store \uc640 timeout \uc774 \uc0ac\uc6a9\ub418\uc9c0 \uc54a\uc73c\ubbc0\ub85c BackendDummy \uc778\uc2a4\ud134\uc2a4\ud654 \uba54\uc11c\ub4dc\uc5d0\uc11c \ubb34\uc2dc\ub429\ub2c8\ub2e4. \uadf8\ub7ec\ub098 \uc2e4\uc81c \ud655\uc7a5 \uae30\ub2a5\uc740 \ub791\ub370\ubdf0\ub97c \uc218\ud589\ud558\uace0 timeout \uc778\uc218\ub97c \uc9c0\uc6d0\ud558\uae30 \uc704\ud574 store \uc0ac\uc6a9\uc744 \uace0\ub824\ud574\uc57c \ud569\ub2c8\ub2e4. // file name: dummy.hpp class BackendDummy : public Backend { ... \u003cStep 1 code\u003e ... static c10::intrusive_ptr\u003cBackend\u003e createBackendDummy( const c10::intrusive_ptr\u003c::c10d::Store\u003e\u0026 store, int rank, int size, const std::chrono::duration\u003cfloat\u003e\u0026 timeout); static void BackendDummyConstructor() __attribute__((constructor)) { py::object module = py::module::import(\"torch.distributed\"); py::object register_backend = module.attr(\"Backend\").attr(\"register_backend\"); // torch.distributed.Backend.register_backend\ub294 // `dummy` \ub97c \uc0c8\ub85c\uc6b4 \uc720\ud6a8\ud55c \ubc31\uc5d4\ub4dc\ub85c \ucd94\uac00\ud569\ub2c8\ub2e4. register_backend(\"dummy\", py::cpp_function(createProcessGroupDummy)); } } // file name: dummy.cpp c10::intrusive_ptr\u003cBackend\u003e BackendDummy::createBackendDummy( const c10::intrusive_ptr\u003c::c10d::Store\u003e\u0026 /* unused */, int rank, int size, const std::chrono::duration\u003cfloat\u003e\u0026 /* unused */) { return c10::make_intrusive\u003cBackendDummy\u003e(rank, size); } PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) { m.def(\"createBackendDummy\", \u0026BackendDummy::createBackendDummy); } \ub2e8\uacc4 3: \uc0ac\uc6a9\uc790 \uc815\uc758 \ud655\uc7a5 \ube4c\ub4dc# \uc774\uc81c \ud655\uc7a5 \uc18c\uc2a4 \ucf54\ub4dc \ud30c\uc77c\uc774 \uc900\ube44\ub418\uc5c8\uc2b5\ub2c8\ub2e4. \uadf8\ub7f0 \ub2e4\uc74c cpp \ud655\uc7a5 \uc744 \uc0ac\uc6a9\ud558\uc5ec \ube4c\ub4dc\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uc774\ub97c \uc704\ud574 \uacbd\ub85c\uc640 \uba85\ub839\uc744 \uc900\ube44\ud558\ub294 setup.py \ud30c\uc77c\uc744 \uc0dd\uc131\ud558\uace0, python setup.py develop \uc744 \ud638\ucd9c\ud558\uc5ec \ud655\uc7a5\uc744 \uc124\uce58\ud569\ub2c8\ub2e4. \ud655\uc7a5\uc774 \uc11c\ub4dc\ud30c\ud2f0 \ub77c\uc774\ube0c\ub7ec\ub9ac\uc5d0 \uc758\uc874\ud558\ub294 \uacbd\uc6b0, cpp \ud655\uc7a5 API\uc5d0 libraries_dirs \ubc0f libraries \uc9c0\uc815\ud560 \uc218\ub3c4 \uc788\uc2b5\ub2c8\ub2e4. \uc2e4\uc81c \uc608\uc81c\ub85c torch ucc \ud504\ub85c\uc81d\ud2b8\ub97c \ucc38\uc870\ud558\uc2ed\uc2dc\uc624. # \ud30c\uc77c \uc774\ub984: setup.py import os import sys import torch from setuptools import setup from torch.utils import cpp_extension sources = [\"src/dummy.cpp\"] include_dirs = [f\"{os.path.dirname(os.path.abspath(__file__))}/include/\"] if torch.cuda.is_available(): module = cpp_extension.CUDAExtension( name = \"dummy_collectives\", sources = sources, include_dirs = include_dirs, ) else: module = cpp_extension.CppExtension( name = \"dummy_collectives\", sources = sources, include_dirs = include_dirs, ) setup( name = \"Dummy-Collectives\", version = \"0.0.1\", ext_modules = [module], cmdclass={\u0027build_ext\u0027: cpp_extension.BuildExtension} ) \ub2e8\uacc4 4: \uc751\uc6a9 \ud504\ub85c\uadf8\ub7a8\uc5d0\uc11c \ud655\uc7a5 \uae30\ub2a5 \uc0ac\uc6a9# \uc124\uce58 \ud6c4 init_process_group \uc744 \ud638\ucd9c\ud560 \ub54c dummy \ubc31\uc5d4\ub4dc\ub97c \ub0b4\uc7a5\ub41c \ubc31\uc5d4\ub4dc\ucc98\ub7fc \ud3b8\ub9ac\ud558\uac8c \uc0ac\uc6a9\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. init_process_group \uc758 backend \uc778\uc790(argument)\ub97c dummy \ub85c \ubcc0\uacbd\ud558\uc5ec \ubc31\uc5d4\ub4dc\ub97c \uae30\ubc18\uc73c\ub85c \ub514\uc2a4\ud328\uce58(dispatch)\ud558\ub3c4\ub85d \uc9c0\uc815\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uc774 \ub54c backend \uc778\uc790\ub85c cpu:gloo,cuda:dummy \ub97c \uc9c0\uc815\ud558\uba74 CPU \ud150\uc11c\uc5d0 \ub300\ud574\uc11c\ub294 gloo \ubc31\uc5d4\ub4dc\ub97c \uc0ac\uc6a9\ud558\uace0 CUDA \ud150\uc11c\uc5d0 \ub300\ud574\uc11c\ub294 dummy \ubc31\uc5d4\ub4dc\ub97c \uc0ac\uc6a9\ud558\uc5ec \uc9d1\ud569 \ud1b5\uc2e0\uc744 \ub514\uc2a4\ud328\uce58\ud558\ub3c4\ub85d \uc9c0\uc815\ud569\ub2c8\ub2e4. \ubaa8\ub4e0 \ud150\uc11c\ub4e4\uc744 dummy \ubc31\uc5d4\ub4dc\ub85c \ubcf4\ub0b4\ub824\uba74 \uadf8\ub0e5 dummy \ub97c \ubc31\uc5d4\ub4dc \uc778\uc790\ub85c \uc9c0\uc815\ud558\uba74 \ub429\ub2c8\ub2e4. import os import torch # dummy_collectives\ub97c import\ud558\uba74 torch.distributed\uac00 `dummy` \ub97c \uc720\ud6a8\ud55c \ubc31\uc5d4\ub4dc\ub85c \uc778\uc2dd\ud569\ub2c8\ub2e4. import dummy_collectives import torch.distributed as dist os.environ[\u0027MASTER_ADDR\u0027] = \u0027localhost\u0027 os.environ[\u0027MASTER_PORT\u0027] = \u002729500\u0027 # Alternatively: # dist.init_process_group(\"dummy\", rank=0, world_size=1) dist.init_process_group(\"cpu:gloo,cuda:dummy\", rank=0, world_size=1) # \uc774 \ud150\uc11c\ub294 gloo \ubc31\uc5d4\ub4dc\ub97c \uc0ac\uc6a9\ud558\uace0 x = torch.ones(6) dist.all_reduce(x) print(f\"cpu allreduce: {x}\") # \uc774 \ud150\uc11c\ub294 dummy \ubc31\uc5d4\ub4dc\ub97c \uc0ac\uc6a9\ud569\ub2c8\ub2e4. if torch.cuda.is_available(): y = x.cuda() dist.all_reduce(y) print(f\"cuda allreduce: {y}\") try: dist.broadcast(y, 0) except RuntimeError: print(\"got RuntimeError when calling broadcast\")",
       "author": {
         "@type": "Organization",
         "name": "PyTorch Contributors",
         "url": "https://pytorch.org"
       },
       "image": "../_static/img/pytorch_seo.png",
       "mainEntityOfPage": {
         "@type": "WebPage",
         "@id": "/intermediate/process_group_cpp_extension_tutorial.html"
       },
       "datePublished": "2023-01-01T00:00:00Z",
       "dateModified": "2023-01-01T00:00:00Z"
     }
 

article:modified_time2022-11-30T07:09:41+00:00
og:typearticle
og:site_namePyTorch Tutorials KR
og:image../_static/img/pytorch_seo.png
og:image:altPyTorch Tutorials KR
og:ignore_canonicaltrue
docsearch:languageko
docbuild:last-update2022년 11월 30일
None2
pytorch_projecttutorials

Links:

https://pytorch.kr/
PyTorch 시작하기 https://pytorch.kr/get-started/locally/
기본 익히기 https://tutorials.pytorch.kr/beginner/basics/intro.html
한국어 튜토리얼 https://tutorials.pytorch.kr/
한국어 모델 허브 https://pytorch.kr/hub/
Official Tutorials https://docs.pytorch.org/tutorials/
블로그 https://pytorch.kr/blog/
PyTorch API https://docs.pytorch.org/docs/
Domain API 소개 https://pytorch.kr/domains/
한국어 튜토리얼 https://tutorials.pytorch.kr/
Official Tutorials https://docs.pytorch.org/tutorials/
한국어 커뮤니티 https://discuss.pytorch.kr/
개발자 정보 https://pytorch.kr/resources/
Landscape https://landscape.pytorch.org/
https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html
https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html
PyTorch 시작하기https://pytorch.kr/get-started/locally/
기본 익히기https://tutorials.pytorch.kr/beginner/basics/intro.html
한국어 튜토리얼https://tutorials.pytorch.kr/
한국어 모델 허브https://pytorch.kr/hub/
Official Tutorialshttps://docs.pytorch.org/tutorials/
블로그https://pytorch.kr/blog/
PyTorch APIhttps://docs.pytorch.org/docs/
Domain API 소개https://pytorch.kr/domains/
한국어 튜토리얼https://tutorials.pytorch.kr/
Official Tutorialshttps://docs.pytorch.org/tutorials/
한국어 커뮤니티https://discuss.pytorch.kr/
개발자 정보https://pytorch.kr/resources/
Landscapehttps://landscape.pytorch.org/
Skip to main contenthttps://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#main-content
v2.8.0+cu128https://tutorials.pytorch.kr/index.html
Intro https://tutorials.pytorch.kr/intro.html
Compilers https://tutorials.pytorch.kr/compilers_index.html
Domains https://tutorials.pytorch.kr/domains.html
Distributed https://tutorials.pytorch.kr/distributed.html
Deep Dive https://tutorials.pytorch.kr/deep-dive.html
Extension https://tutorials.pytorch.kr/extension.html
Ecosystem https://tutorials.pytorch.kr/ecosystem.html
Recipes https://tutorials.pytorch.kr/recipes_index.html
한국어 튜토리얼 GitHub 저장소https://github.com/PyTorchKorea/tutorials-kr
파이토치 한국어 커뮤니티https://discuss.pytorch.kr/
Intro https://tutorials.pytorch.kr/intro.html
Compilers https://tutorials.pytorch.kr/compilers_index.html
Domains https://tutorials.pytorch.kr/domains.html
Distributed https://tutorials.pytorch.kr/distributed.html
Deep Dive https://tutorials.pytorch.kr/deep-dive.html
Extension https://tutorials.pytorch.kr/extension.html
Ecosystem https://tutorials.pytorch.kr/ecosystem.html
Recipes https://tutorials.pytorch.kr/recipes_index.html
한국어 튜토리얼 GitHub 저장소https://github.com/PyTorchKorea/tutorials-kr
파이토치 한국어 커뮤니티https://discuss.pytorch.kr/
PyTorch Distributed Overviewhttps://tutorials.pytorch.kr/beginner/dist_overview.html
PyTorch의 분산 데이터 병렬 처리 - 비디오 튜토리얼https://tutorials.pytorch.kr/beginner/ddp_series_intro.html
분산 데이터 병렬 처리 시작하기https://tutorials.pytorch.kr/intermediate/ddp_tutorial.html
PyTorch로 분산 어플리케이션 개발하기https://tutorials.pytorch.kr/intermediate/dist_tuto.html
Getting Started with Fully Sharded Data Parallel (FSDP2)https://tutorials.pytorch.kr/intermediate/FSDP_tutorial.html
Introduction to Libuv TCPStore Backendhttps://tutorials.pytorch.kr/intermediate/TCPStore_libuv_backend.html
Large Scale Transformer model training with Tensor Parallel (TP)https://tutorials.pytorch.kr/intermediate/TP_tutorial.html
Introduction to Distributed Pipeline Parallelismhttps://tutorials.pytorch.kr/intermediate/pipelining_tutorial.html
Cpp 확장을 사용한 프로세스 그룹 백엔드 사용자 정의https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html
Getting Started with Distributed RPC Frameworkhttps://tutorials.pytorch.kr/intermediate/rpc_tutorial.html
Implementing a Parameter Server Using Distributed RPC Frameworkhttps://tutorials.pytorch.kr/intermediate/rpc_param_server_tutorial.html
Implementing Batch RPC Processing Using Asynchronous Executionshttps://tutorials.pytorch.kr/intermediate/rpc_async_execution.html
분산 데이터 병렬(DDP)과 분산 RPC 프레임워크 결합https://tutorials.pytorch.kr/advanced/rpc_ddp_tutorial.html
Distributed Training with Uneven Inputs Using the Join Context Managerhttps://tutorials.pytorch.kr/advanced/generic_join.html
https://tutorials.pytorch.kr/index.html
Distributedhttps://tutorials.pytorch.kr/distributed.html
#https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#cpp
Howard Huanghttps://github.com/H-Huang
Feng Tianhttps://github.com/ftian1
Shen Lihttps://mrshenli.github.io/
Min Sihttps://minsii.github.io/
박재윤https://github.com/jenner9212
https://tutorials.pytorch.kr/_images/pencil-16.png
githubhttps://github.com/pytorchkorea/tutorials-kr/blob/main/intermediate_source/process_group_cpp_extension_tutorial.rst
PyTorch Distributed Overviewhttps://tutorials.pytorch.kr/beginner/dist_overview.html
PyTorch Collective Communication Packagehttps://pytorch.org/docs/stable/distributed.html
PyTorch Cpp Extensionhttps://pytorch.org/docs/stable/cpp_extension.html
Writing Distributed Applications with PyTorchhttps://tutorials.pytorch.kr/intermediate/dist_tuto.html
cpp 확장https://pytorch.org/docs/stable/cpp_extension.html
파이토치 분산 패키지https://pytorch.org/docs/stable/distributed.html
#https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#id2
분산 데이터 병렬(DistributedDataParallel)https://pytorch.org/docs/stable/generated/torch.nn.parallel.DistributedDataParallel.html
제로 리던던시 최적화기(ZeroRedundancyOptimizer)https://pytorch.org/docs/stable/distributed.optim.html#torch.distributed.optim.ZeroRedundancyOptimizer
Backendhttps://github.com/pytorch/pytorch/blob/main/torch/csrc/distributed/c10d/Backend.hpp
UCChttps://github.com/openucx/ucc
OneCCLhttps://github.com/oneapi-src/oneCCL
TPUhttps://cloud.google.com/tpu
Trainumhttps://aws.amazon.com/machine-learning/trainium/
Herringhttps://www.amazon.science/publications/herring-rethinking-the-parameter-server-at-scale-for-the-cloud
Reduction Serverhttps://cloud.google.com/blog/topics/developers-practitioners/optimize-training-performance-reduction-server-vertex-ai
#https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#backend
더미 집합(dummy collectives)https://github.com/H-Huang/torch_collective_extension
#https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#api
파이썬 측https://github.com/pytorch/pytorch/blob/v1.9.0/torch/distributed/distributed_c10d.py#L643-L650
#https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#id3
cpp 확장https://pytorch.org/docs/stable/cpp_extension.html
torch ucchttps://github.com/openucx/torch-ucc
#https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#id4
init_process_grouphttps://pytorch.org/docs/stable/distributed.html#torch.distributed.init_process_group
이전 Introduction to Distributed Pipeline Parallelism https://tutorials.pytorch.kr/intermediate/pipelining_tutorial.html
다음 Getting Started with Distributed RPC Framework https://tutorials.pytorch.kr/intermediate/rpc_tutorial.html
PyData Sphinx Themehttps://pydata-sphinx-theme.readthedocs.io/en/stable/index.html
이전 Introduction to Distributed Pipeline Parallelism https://tutorials.pytorch.kr/intermediate/pipelining_tutorial.html
다음 Getting Started with Distributed RPC Framework https://tutorials.pytorch.kr/intermediate/rpc_tutorial.html
기초https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#id2
단계 1: Backend 의 하위 클래스 구현https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#backend
단계 2: 확장 기능을 파이썬 API로 노출https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#api
단계 3: 사용자 정의 확장 빌드https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#id3
단계 4: 응용 프로그램에서 확장 기능 사용https://tutorials.pytorch.kr/intermediate/process_group_cpp_extension_tutorial.html#id4
torchaohttps://docs.pytorch.org/ao
torchrechttps://docs.pytorch.org/torchrec
torchfthttps://docs.pytorch.org/torchft
TorchCodechttps://docs.pytorch.org/torchcodec
torchvisionhttps://docs.pytorch.org/vision
ExecuTorchhttps://docs.pytorch.org/executorch
PyTorch on XLA Deviceshttps://docs.pytorch.org/xla
GitHub로 이동https://github.com/PyTorchKorea
튜토리얼로 이동https://tutorials.pytorch.kr/
커뮤니티로 이동https://discuss.pytorch.kr/
https://pytorch.kr/
파이토치 한국 사용자 모임https://pytorch.kr/
사용자 모임 소개https://pytorch.kr/about
기여해주신 분들https://pytorch.kr/contributors
리소스https://pytorch.kr/resources/
행동 강령https://pytorch.kr/coc
행동 강령https://pytorch.kr/coc
Linux Foundation의 정책https://www.linuxfoundation.org/policies/
our code of conducthttps://pytorch.kr/coc
Linux Foundation's policieshttps://www.linuxfoundation.org/policies/
Cookies Policyhttps://www.facebook.com/policies/cookies/
Sphinxhttps://www.sphinx-doc.org/
PyData Sphinx Themehttps://pydata-sphinx-theme.readthedocs.io/en/stable/index.html

Viewport: width=device-width, initial-scale=1


URLs of crawlers that visited me.