René's URL Explorer Experiment


Title: (베타) PyTorch를 사용한 Channels Last 메모리 형식 — 파이토치 한국어 튜토리얼 (PyTorch tutorials in Korean)

Open Graph Title: (베타) PyTorch를 사용한 Channels Last 메모리 형식

Description: Author: Vitaly Fedyunin 번역: Choi Yoonjeong 무엇을 배울 수 있나요? PyTorch에서의 Channels last 메모리 형식은 무엇인가요?, 특정 연산자에서 성능을 향상시키기 위해 어떻게 사용할 수 있나요?. 전제 조건 PyTorch v1.5.0, CUDA 지원 GPU. Channels last 메모리 형식(memory format)은 차원 순서를 유지하면서 메모리 상의 NCHW 텐서(tensor)를 정렬하는 또 다른 방식입니다. Channels last 텐서는 채널(Channel)이 가장 밀...

Open Graph Description: Author: Vitaly Fedyunin 번역: Choi Yoonjeong 무엇을 배울 수 있나요? PyTorch에서의 Channels last 메모리 형식은 무엇인가요?, 특정 연산자에서 성능을 향상시키기 위해 어떻게 사용할 수 있나요?. 전제 조건 PyTorch v1.5.0, CUDA 지원 GPU. Channels last 메모리 형식(memory format)은 차원 순서를 유지하면서 메모리 상의 NCHW 텐서(tensor)를 정렬하는 또 다른 방식입니다. Channels last 텐서는 채널(Channel)이 가장 밀...

Opengraph URL: https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html

direct link

Domain: tutorials.pytorch.kr


Hey, it has json ld scripts:
    {
       "@context": "https://schema.org",
       "@type": "Article",
       "name": "(\ubca0\ud0c0) PyTorch\ub97c \uc0ac\uc6a9\ud55c Channels Last \uba54\ubaa8\ub9ac \ud615\uc2dd",
       "headline": "(\ubca0\ud0c0) PyTorch\ub97c \uc0ac\uc6a9\ud55c Channels Last \uba54\ubaa8\ub9ac \ud615\uc2dd",
       "description": "PyTorch Documentation. Explore PyTorch, an open-source machine learning library that accelerates the path from research prototyping to production deployment. Discover tutorials, API references, and guides to help you build and deploy deep learning models efficiently.",
       "url": "/intermediate/memory_format_tutorial.html",
       "articleBody": "\ucc38\uace0 Go to the end to download the full example code. (\ubca0\ud0c0) PyTorch\ub97c \uc0ac\uc6a9\ud55c Channels Last \uba54\ubaa8\ub9ac \ud615\uc2dd# Author: Vitaly Fedyunin \ubc88\uc5ed: Choi Yoonjeong \ubb34\uc5c7\uc744 \ubc30\uc6b8 \uc218 \uc788\ub098\uc694? PyTorch\uc5d0\uc11c\uc758 Channels last \uba54\ubaa8\ub9ac \ud615\uc2dd\uc740 \ubb34\uc5c7\uc778\uac00\uc694? \ud2b9\uc815 \uc5f0\uc0b0\uc790\uc5d0\uc11c \uc131\ub2a5\uc744 \ud5a5\uc0c1\uc2dc\ud0a4\uae30 \uc704\ud574 \uc5b4\ub5bb\uac8c \uc0ac\uc6a9\ud560 \uc218 \uc788\ub098\uc694? \uc804\uc81c \uc870\uac74 PyTorch v1.5.0 CUDA \uc9c0\uc6d0 GPU Channels last \uba54\ubaa8\ub9ac \ud615\uc2dd(memory format)\uc740 \ucc28\uc6d0 \uc21c\uc11c\ub97c \uc720\uc9c0\ud558\uba74\uc11c \uba54\ubaa8\ub9ac \uc0c1\uc758 NCHW \ud150\uc11c(tensor)\ub97c \uc815\ub82c\ud558\ub294 \ub610 \ub2e4\ub978 \ubc29\uc2dd\uc785\ub2c8\ub2e4. Channels last \ud150\uc11c\ub294 \ucc44\ub110(Channel)\uc774 \uac00\uc7a5 \ubc00\ub3c4\uac00 \ub192\uc740(densest) \ucc28\uc6d0\uc73c\ub85c \uc815\ub82c(\uc608. \uc774\ubbf8\uc9c0\ub97c \ud53d\uc140x\ud53d\uc140\ub85c \uc800\uc7a5)\ub429\ub2c8\ub2e4. \uc608\ub97c \ub4e4\uc5b4, (2\uac1c\uc758 4 x 4 \uc774\ubbf8\uc9c0\uc5d0 3\uac1c\uc758 \ucc44\ub110\uc774 \uc874\uc7ac\ud558\ub294 \uacbd\uc6b0) \uc804\ud615\uc801\uc778(\uc5f0\uc18d\uc801\uc778) NCHW \ud150\uc11c\uc758 \uc800\uc7a5 \ubc29\uc2dd\uc740 \ub2e4\uc74c\uacfc \uac19\uc2b5\ub2c8\ub2e4: Channels last \uba54\ubaa8\ub9ac \ud615\uc2dd\uc740 \ub370\uc774\ud130\ub97c \ub2e4\ub974\uac8c \uc815\ub82c\ud569\ub2c8\ub2e4: PyTorch\ub294 \uae30\uc874\uc758 \uc2a4\ud2b8\ub77c\uc774\ub4dc(strides) \uad6c\uc870\ub97c \uc0ac\uc6a9\ud568\uc73c\ub85c\uc368 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc744 \uc9c0\uc6d0\ud569\ub2c8\ub2e4. \uc608\ub97c \ub4e4\uc5b4, Channels last \ud615\uc2dd\uc5d0\uc11c 10x3x16x16 \ubc30\uce58(batch)\ub294 (768, 1, 48, 3)\uc640 \uac19\uc740 \ud3ed(strides)\uc744 \uac00\uc9c0\uace0 \uc788\uac8c \ub429\ub2c8\ub2e4. Channels last \uba54\ubaa8\ub9ac \ud615\uc2dd\uc740 \uc624\uc9c1 4D NCHW Tensors\uc5d0\uc11c\ub9cc \uc2e4\ud589\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uba54\ubaa8\ub9ac \ud615\uc2dd(Memory Format) API# \uc5f0\uc18d \uba54\ubaa8\ub9ac \ud615\uc2dd\uacfc channels last \uba54\ubaa8\ub9ac \ud615\uc2dd \uac04\uc5d0 \ud150\uc11c\ub97c \ubcc0\ud658\ud558\ub294 \ubc29\ubc95\uc740 \ub2e4\uc74c\uacfc \uac19\uc2b5\ub2c8\ub2e4. \uc804\ud615\uc801\uc778 PyTorch\uc758 \uc5f0\uc18d\uc801\uc778 \ud150\uc11c(tensor) import torch N, C, H, W = 10, 3, 32, 32 x = torch.empty(N, C, H, W) print(x.stride()) # \uacb0\uacfc: (3072, 1024, 32, 1) (3072, 1024, 32, 1) \ubcc0\ud658 \uc5f0\uc0b0\uc790 x = x.to(memory_format=torch.channels_last) print(x.shape) # \uacb0\uacfc: (10, 3, 32, 32) \ucc28\uc6d0 \uc21c\uc11c\ub294 \ubcf4\uc874\ud568 print(x.stride()) # \uacb0\uacfc: (3072, 1, 96, 3) torch.Size([10, 3, 32, 32]) (3072, 1, 96, 3) \uc5f0\uc18d\uc801\uc778 \ud615\uc2dd\uc73c\ub85c \ub418\ub3cc\ub9ac\uae30 x = x.to(memory_format=torch.contiguous_format) print(x.stride()) # \uacb0\uacfc: (3072, 1024, 32, 1) (3072, 1024, 32, 1) \ub2e4\ub978 \ubc29\uc2dd x = x.contiguous(memory_format=torch.channels_last) print(x.stride()) # \uacb0\uacfc: (3072, 1, 96, 3) (3072, 1, 96, 3) \ud615\uc2dd(format) \ud655\uc778 print(x.is_contiguous(memory_format=torch.channels_last)) # \uacb0\uacfc: True True to \uc640 contiguous \uc5d0\ub294 \uc791\uc740 \ucc28\uc774(minor difference)\uac00 \uc788\uc2b5\ub2c8\ub2e4. \uba85\uc2dc\uc801\uc73c\ub85c \ud150\uc11c(tensor)\uc758 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc744 \ubcc0\ud658\ud560 \ub54c\ub294 to \ub97c \uc0ac\uc6a9\ud558\ub294 \uac83\uc744 \uad8c\uc7a5\ud569\ub2c8\ub2e4. \ub300\ubd80\ubd84\uc758 \uacbd\uc6b0 \ub450 API\ub294 \ub3d9\uc77c\ud558\uac8c \ub3d9\uc791\ud569\ub2c8\ub2e4. \ud558\uc9c0\ub9cc C==1 \uc774\uac70\ub098 H == 1 \u0026\u0026 W == 1 \uc778 NCHW 4D \ud150\uc11c\uc758 \ud2b9\uc218\ud55c \uacbd\uc6b0\uc5d0\ub294 to \ub9cc\uc774 Channel last \uba54\ubaa8\ub9ac \ud615\uc2dd\uc73c\ub85c \ud45c\ud604\ub41c \uc801\uc808\ud55c \ud3ed(stride)\uc744 \uc0dd\uc131\ud569\ub2c8\ub2e4. \uc774\ub294 \uc704\uc758 \ub450\uac00\uc9c0 \uacbd\uc6b0\uc5d0 \ud150\uc11c\uc758 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc774 \ubaa8\ud638\ud558\uae30 \ub54c\ubb38\uc785\ub2c8\ub2e4. \uc608\ub97c \ub4e4\uc5b4, \ud06c\uae30\uac00 N1HW \uc778 \uc5f0\uc18d\uc801\uc778 \ud150\uc11c(contiguous tensor)\ub294 \uc5f0\uc18d\uc801 \uc774\uba74\uc11c Channel last \ud615\uc2dd\uc73c\ub85c \uba54\ubaa8\ub9ac\uc5d0 \uc800\uc7a5\ub429\ub2c8\ub2e4. \ub530\ub77c\uc11c, \uc8fc\uc5b4\uc9c4 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc5d0 \ub300\ud574 \uc774\ubbf8 is_contiguous \ub85c \uac04\uc8fc\ub418\uc5b4 contiguous \ud638\ucd9c\uc740 \ub3d9\uc791\ud558\uc9c0 \uc54a\uac8c(no-op) \ub418\uc5b4, \ud3ed(stride)\uc744 \uac31\uc2e0\ud558\uc9c0 \uc54a\uac8c \ub429\ub2c8\ub2e4. \ubc18\uba74\uc5d0, to \ub294 \uc758\ub3c4\ud55c \uba54\ubaa8\ub9ac \ud615\uc2dd\uc73c\ub85c \uc801\uc808\ud558\uac8c \ud45c\ud604\ud558\uae30 \uc704\ud574 \ud06c\uae30\uac00 1\uc778 \ucc28\uc6d0\uc5d0\uc11c \uc758\ubbf8\uc788\ub294 \ud3ed(stride)\uc73c\ub85c \uc7ac\ubc30\uc5f4(restride)\ud569\ub2c8\ub2e4. special_x = torch.empty(4, 1, 4, 4) print(special_x.is_contiguous(memory_format=torch.channels_last)) # Ouputs: True print(special_x.is_contiguous(memory_format=torch.contiguous_format)) # Ouputs: True True True \uba85\uc2dc\uc801 \uce58\ud658(permutation) API\uc778 permute \uc5d0\uc11c\ub3c4 \ub3d9\uc77c\ud558\uac8c \uc801\uc6a9\ub429\ub2c8\ub2e4. \ubaa8\ud638\uc131\uc774 \ubc1c\uc0dd\ud560 \uc218 \uc788\ub294 \ud2b9\ubcc4\ud55c \uacbd\uc6b0\uc5d0, permute \ub294 \uc758\ub3c4\ud55c \uba54\ubaa8\ub9ac \ud615\uc2dd\uc73c\ub85c \uc804\ub2ec\ub418\ub294 \ud3ed(stride)\uc744 \uc0dd\uc131\ud558\ub294 \uac83\uc774 \ubcf4\uc7a5\ub418\uc9c0 \uc54a\uc2b5\ub2c8\ub2e4. to \ub85c \uba85\uc2dc\uc801\uc73c\ub85c \uba54\ubaa8\ub9ac \ud615\uc2dd\uc744 \uc9c0\uc815\ud558\uc5ec \uc758\ub3c4\uce58 \uc54a\uc740 \ub3d9\uc791\uc744 \ud53c\ud560 \uac83\uc744 \uad8c\uc7a5\ud569\ub2c8\ub2e4. \ub610\ud55c, 3\uac1c\uc758 \ube44-\ubc30\uce58(non-batch) \ucc28\uc6d0\uc774 \ubaa8\ub450 1 \uc778 \uadf9\ub2e8\uc801\uc778 \uacbd\uc6b0(C==1 \u0026\u0026 H==1 \u0026\u0026 W==1), \ud604\uc7ac \uad6c\ud604\uc740 \ud150\uc11c\ub97c Channels last \uba54\ubaa8\ub9ac \ud615\uc2dd\uc73c\ub85c \ud45c\uc2dc\ud560 \uc218 \uc5c6\uc74c\uc744 \uc54c\ub824\ub4dc\ub9bd\ub2c8\ub2e4. Channels last \ubc29\uc2dd\uc73c\ub85c \uc0dd\uc131\ud558\uae30 x = torch.empty(N, C, H, W, memory_format=torch.channels_last) print(x.stride()) # \uacb0\uacfc: (3072, 1, 96, 3) (3072, 1, 96, 3) clone \uc740 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc744 \ubcf4\uc874\ud569\ub2c8\ub2e4. y = x.clone() print(y.stride()) # \uacb0\uacfc: (3072, 1, 96, 3) (3072, 1, 96, 3) to, cuda, float \u2026 \ub4f1\ub3c4 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc744 \ubcf4\uc874\ud569\ub2c8\ub2e4. if torch.cuda.is_available(): y = x.cuda() print(y.stride()) # \uacb0\uacfc: (3072, 1, 96, 3) (3072, 1, 96, 3) empty_like, *_like \uc5f0\uc0b0\uc790\ub3c4 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc744 \ubcf4\uc874\ud569\ub2c8\ub2e4. y = torch.empty_like(x) print(y.stride()) # \uacb0\uacfc: (3072, 1, 96, 3) (3072, 1, 96, 3) Pointwise \uc5f0\uc0b0\uc790\ub3c4 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc744 \ubcf4\uc874\ud569\ub2c8\ub2e4. z = x + y print(z.stride()) # \uacb0\uacfc: (3072, 1, 96, 3) (3072, 1, 96, 3) `cudnn \ubc31\uc5d4\ub4dc\ub97c \uc0ac\uc6a9\ud558\ub294 Conv`, Batchnorm \ubaa8\ub4c8\uc740 Channels last\ub97c \uc9c0\uc6d0\ud569\ub2c8\ub2e4. (\ub2e8, CudNN \u003e=7.6 \uc5d0\uc11c\ub9cc \ub3d9\uc791) \ud569\uc131\uacf1(convolution) \ubaa8\ub4c8\uc740 \uc774\uc9c4 p-wise \uc5f0\uc0b0\uc790(binary p-wise operator)\uc640\ub294 \ub2e4\ub974\uac8c Channels last\uac00 \uc8fc\ub41c \uba54\ubaa8\ub9ac \ud615\uc2dd\uc785\ub2c8\ub2e4. \ubaa8\ub4e0 \uc785\ub825\uc740 \uc5f0\uc18d\uc801\uc778 \uba54\ubaa8\ub9ac \ud615\uc2dd\uc774\uba70, \uc5f0\uc0b0\uc790\ub294 \uc5f0\uc18d\ub41c \uba54\ubaa8\ub9ac \ud615\uc2dd\uc73c\ub85c \ucd9c\ub825\uc744 \uc0dd\uc131\ud569\ub2c8\ub2e4. \uadf8\ub807\uc9c0 \uc54a\uc73c\uba74, \ucd9c\ub825\uc740 channels last \uba54\ubaa8\ub9ac \ud615\uc2dd\uc785\ub2c8\ub2e4. if torch.backends.cudnn.is_available() and torch.backends.cudnn.version() \u003e= 7603: model = torch.nn.Conv2d(8, 4, 3).cuda().half() model = model.to(memory_format=torch.channels_last) # \ubaa8\ub4c8 \uc778\uc790\ub4e4\uc740 Channels last\ub85c \ubcc0\ud658\uc774 \ud544\uc694\ud569\ub2c8\ub2e4 input = torch.randint(1, 10, (2, 8, 4, 4), dtype=torch.float32, requires_grad=True) input = input.to(device=\"cuda\", memory_format=torch.channels_last, dtype=torch.float16) out = model(input) print(out.is_contiguous(memory_format=torch.channels_last)) # \uacb0\uacfc: True True \uc785\ub825 \ud150\uc11c\uac00 Channels last\ub97c \uc9c0\uc6d0\ud558\uc9c0 \uc54a\ub294 \uc5f0\uc0b0\uc790\ub97c \ub9cc\ub098\uba74 \uce58\ud658(permutation)\uc774 \ucee4\ub110\uc5d0 \uc790\ub3d9\uc73c\ub85c \uc801\uc6a9\ub418\uc5b4 \uc785\ub825 \ud150\uc11c\ub97c \uc5f0\uc18d\uc801\uc778 \ud615\uc2dd\uc73c\ub85c \ubcf5\uc6d0\ud569\ub2c8\ub2e4. \uc774 \uacbd\uc6b0 \uacfc\ubd80\ud558\uac00 \ubc1c\uc0dd\ud558\uc5ec channel last \uba54\ubaa8\ub9ac \ud615\uc2dd\uc758 \uc804\ud30c\uac00 \uc911\ub2e8\ub429\ub2c8\ub2e4. \uadf8\ub7fc\uc5d0\ub3c4 \ubd88\uad6c\ud558\uace0, \uc62c\ubc14\ub978 \ucd9c\ub825\uc740 \ubcf4\uc7a5\ub429\ub2c8\ub2e4. \uc131\ub2a5 \ud5a5\uc0c1# Channels last \uba54\ubaa8\ub9ac \ud615\uc2dd \ucd5c\uc801\ud654\ub294 GPU\uc640 CPU\uc5d0\uc11c \ubaa8\ub450 \uc0ac\uc6a9 \uac00\ub2a5\ud569\ub2c8\ub2e4. GPU\uc5d0\uc11c\ub294 \uc815\ubc00\ub3c4\ub97c \uc904\uc778(reduced precision torch.float16) \uc0c1\ud0dc\uc5d0\uc11c Tensor Cores\ub97c \uc9c0\uc6d0\ud558\ub294 NVIDIA\uc758 \ud558\ub4dc\uc6e8\uc5b4\uc5d0\uc11c \uac00\uc7a5 \uc758\ubbf8\uc2ec\uc7a5\ud55c \uc131\ub2a5 \ud5a5\uc0c1\uc744 \ubcf4\uc600\uc2b5\ub2c8\ub2e4. AMP (Automated Mixed Precision) \ud559\uc2b5 \uc2a4\ud06c\ub9bd\ud2b8\ub97c \ud65c\uc6a9\ud558\uc5ec \uc5f0\uc18d\uc801\uc778 \ud615\uc2dd\uc5d0 \ube44\ud574 Channels last \ubc29\uc2dd\uc774 22% \uc774\uc0c1\uc758 \uc131\ub2a5 \ud5a5\uc2b9\uc744 \ud655\uc778\ud560 \uc218 \uc788\uc5c8\uc2b5\ub2c8\ub2e4. \uc774 \ub54c, NVIDIA\uac00 \uc81c\uacf5\ud558\ub294 AMP\ub97c \uc0ac\uc6a9\ud588\uc2b5\ub2c8\ub2e4. NVIDIA/apex python main_amp.py -a resnet50 --b 200 --workers 16 --opt-level O2\u00a0 ./data # opt_level = O2 # keep_batchnorm_fp32 = None \u003cclass \u0027NoneType\u0027\u003e # loss_scale = None \u003cclass \u0027NoneType\u0027\u003e # CUDNN VERSION: 7603 # =\u003e creating model \u0027resnet50\u0027 # Selected optimization level O2: FP16 training with FP32 batchnorm and FP32 master weights. # Defaults for this optimization level are: # enabled : True # opt_level : O2 # cast_model_type : torch.float16 # patch_torch_functions : False # keep_batchnorm_fp32 : True # master_weights : True # loss_scale : dynamic # Processing user overrides (additional kwargs that are not None)... # After processing overrides, optimization options are: # enabled : True # opt_level : O2 # cast_model_type : torch.float16 # patch_torch_functions : False # keep_batchnorm_fp32 : True # master_weights : True # loss_scale : dynamic # Epoch: [0][10/125] Time 0.866 (0.866) Speed 230.949 (230.949) Loss 0.6735125184 (0.6735) Prec@1 61.000 (61.000) Prec@5 100.000 (100.000) # Epoch: [0][20/125] Time 0.259 (0.562) Speed 773.481 (355.693) Loss 0.6968704462 (0.6852) Prec@1 55.000 (58.000) Prec@5 100.000 (100.000) # Epoch: [0][30/125] Time 0.258 (0.461) Speed 775.089 (433.965) Loss 0.7877287269 (0.7194) Prec@1 51.500 (55.833) Prec@5 100.000 (100.000) # Epoch: [0][40/125] Time 0.259 (0.410) Speed 771.710 (487.281) Loss 0.8285319805 (0.7467) Prec@1 48.500 (54.000) Prec@5 100.000 (100.000) # Epoch: [0][50/125] Time 0.260 (0.380) Speed 770.090 (525.908) Loss 0.7370464802 (0.7447) Prec@1 56.500 (54.500) Prec@5 100.000 (100.000) # Epoch: [0][60/125] Time 0.258 (0.360) Speed 775.623 (555.728) Loss 0.7592862844 (0.7472) Prec@1 51.000 (53.917) Prec@5 100.000 (100.000) # Epoch: [0][70/125] Time 0.258 (0.345) Speed 774.746 (579.115) Loss 1.9698858261 (0.9218) Prec@1 49.500 (53.286) Prec@5 100.000 (100.000) # Epoch: [0][80/125] Time 0.260 (0.335) Speed 770.324 (597.659) Loss 2.2505953312 (1.0879) Prec@1 50.500 (52.938) Prec@5 100.000 (100.000) --channels-last true \uc778\uc790\ub97c \uc804\ub2ec\ud558\uc5ec Channels last \ud615\uc2dd\uc73c\ub85c \ubaa8\ub378\uc744 \uc2e4\ud589\ud558\uba74 22%\uc758 \uc131\ub2a5 \ud5a5\uc0c1\uc744 \ubcf4\uc785\ub2c8\ub2e4. python main_amp.py -a resnet50 --b 200 --workers 16 --opt-level O2 --channels-last true ./data # opt_level = O2 # keep_batchnorm_fp32 = None \u003cclass \u0027NoneType\u0027\u003e # loss_scale = None \u003cclass \u0027NoneType\u0027\u003e # # CUDNN VERSION: 7603 # # =\u003e creating model \u0027resnet50\u0027 # Selected optimization level O2: FP16 training with FP32 batchnorm and FP32 master weights. # # Defaults for this optimization level are: # enabled : True # opt_level : O2 # cast_model_type : torch.float16 # patch_torch_functions : False # keep_batchnorm_fp32 : True # master_weights : True # loss_scale : dynamic # Processing user overrides (additional kwargs that are not None)... # After processing overrides, optimization options are: # enabled : True # opt_level : O2 # cast_model_type : torch.float16 # patch_torch_functions : False # keep_batchnorm_fp32 : True # master_weights : True # loss_scale : dynamic # # Epoch: [0][10/125] Time 0.767 (0.767) Speed 260.785 (260.785) Loss 0.7579724789 (0.7580) Prec@1 53.500 (53.500) Prec@5 100.000 (100.000) # Epoch: [0][20/125] Time 0.198 (0.482) Speed 1012.135 (414.716) Loss 0.7007197738 (0.7293) Prec@1 49.000 (51.250) Prec@5 100.000 (100.000) # Epoch: [0][30/125] Time 0.198 (0.387) Speed 1010.977 (516.198) Loss 0.7113101482 (0.7233) Prec@1 55.500 (52.667) Prec@5 100.000 (100.000) # Epoch: [0][40/125] Time 0.197 (0.340) Speed 1013.023 (588.333) Loss 0.8943189979 (0.7661) Prec@1 54.000 (53.000) Prec@5 100.000 (100.000) # Epoch: [0][50/125] Time 0.198 (0.312) Speed 1010.541 (641.977) Loss 1.7113249302 (0.9551) Prec@1 51.000 (52.600) Prec@5 100.000 (100.000) # Epoch: [0][60/125] Time 0.198 (0.293) Speed 1011.163 (683.574) Loss 5.8537774086 (1.7716) Prec@1 50.500 (52.250) Prec@5 100.000 (100.000) # Epoch: [0][70/125] Time 0.198 (0.279) Speed 1011.453 (716.767) Loss 5.7595844269 (2.3413) Prec@1 46.500 (51.429) Prec@5 100.000 (100.000) # Epoch: [0][80/125] Time 0.198 (0.269) Speed 1011.827 (743.883) Loss 2.8196096420 (2.4011) Prec@1 47.500 (50.938) Prec@5 100.000 (100.000) \uc544\ub798 \ubaa9\ub85d\uc758 \ubaa8\ub378\ub4e4\uc740 Channels last \ud615\uc2dd\uc744 \uc804\uc801\uc73c\ub85c \uc9c0\uc6d0(full support)\ud558\uba70 Volta \uc7a5\ube44\uc5d0\uc11c 8%-35%\uc758 \uc131\ub2a5 \ud5a5\uc0c1\uc744 \ubcf4\uc785\ub2c8\ub2e4: alexnet, mnasnet0_5, mnasnet0_75, mnasnet1_0, mnasnet1_3, mobilenet_v2, resnet101, resnet152, resnet18, resnet34, resnet50, resnext50_32x4d, shufflenet_v2_x0_5, shufflenet_v2_x1_0, shufflenet_v2_x1_5, shufflenet_v2_x2_0, squeezenet1_0, squeezenet1_1, vgg11, vgg11_bn, vgg13, vgg13_bn, vgg16, vgg16_bn, vgg19, vgg19_bn, wide_resnet101_2, wide_resnet50_2 \uc544\ub798 \ubaa9\ub85d\uc758 \ubaa8\ub378\ub4e4\uc740 Channels last \ud615\uc2dd\uc744 \uc804\uc801\uc73c\ub85c \uc9c0\uc6d0\ud558\uba70 Intel(R) Xeon(R) Ice Lake (\ub610\ub294 \ucd5c\uc2e0) CPU\uc5d0\uc11c 26%-76% \uc131\ub2a5 \ud5a5\uc0c1\uc744 \ubcf4\uc5ec\uc90d\ub2c8\ub2e4: alexnet, densenet121, densenet161, densenet169, googlenet, inception_v3, mnasnet0_5, mnasnet1_0, resnet101, resnet152, resnet18, resnet34, resnet50, resnext101_32x8d, resnext50_32x4d, shufflenet_v2_x0_5, shufflenet_v2_x1_0, squeezenet1_0, squeezenet1_1, vgg11, vgg11_bn, vgg13, vgg13_bn, vgg16, vgg16_bn, vgg19, vgg19_bn, wide_resnet101_2, wide_resnet50_2 \uae30\uc874 \ubaa8\ub378\ub4e4 \ubcc0\ud658\ud558\uae30# Channels last \uc9c0\uc6d0\uc740 \uae30\uc874 \ubaa8\ub378\uc774 \ubb34\uc5c7\uc774\ub0d0\uc5d0 \ub530\ub77c \uc81c\ud55c\ub418\uc9c0 \uc54a\uc2b5\ub2c8\ub2e4. \uc5b4\ub5a0\ud55c \ubaa8\ub378\ub3c4 Channels last\ub85c \ubcc0\ud658\ud560 \uc218 \uc788\uc73c\uba70 \uc785\ub825(\ub610\ub294 \ud2b9\uc815 \uac00\uc911\uce58)\uc758 \ud615\uc2dd\ub9cc \ub9de\ucdb0\uc8fc\uba74 (\uc2e0\uacbd\ub9dd) \uadf8\ub798\ud504\ub97c \ud1b5\ud574 \ubc14\ub85c \uc804\ud30c(propagate)\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. # \ubaa8\ub378\uc744 \ucd08\uae30\ud654\ud55c(\ub610\ub294 \ubd88\ub7ec\uc628) \uc774\ud6c4, \ud55c \ubc88 \uc2e4\ud589\uc774 \ud544\uc694\ud569\ub2c8\ub2e4. model = model.to(memory_format=torch.channels_last) # \uc6d0\ud558\ub294 \ubaa8\ub378\ub85c \uad50\uccb4\ud558\uae30 # \ubaa8\ub4e0 \uc785\ub825\uc5d0 \ub300\ud574\uc11c \uc2e4\ud589\uc774 \ud544\uc694\ud569\ub2c8\ub2e4. input = input.to(memory_format=torch.channels_last) # \uc6d0\ud558\ub294 \uc785\ub825\uc73c\ub85c \uad50\uccb4\ud558\uae30 output = model(input) \uadf8\ub7ec\ub098, \ubaa8\ub4e0 \uc5f0\uc0b0\uc790\ub4e4\uc774 Channels last\ub97c \uc9c0\uc6d0\ud558\ub3c4\ub85d \uc644\uc804\ud788 \ubc14\ub010 \uac83\uc740 \uc544\ub2d9\ub2c8\ub2e4(\uc77c\ubc18\uc801\uc73c\ub85c\ub294 \uc5f0\uc18d\uc801\uc778 \ucd9c\ub825\uc744 \ub300\uc2e0 \ubc18\ud658\ud569\ub2c8\ub2e4). \uc704\uc758 \uc608\uc2dc\ub4e4\uc5d0\uc11c Channels last\ub97c \uc9c0\uc6d0\ud558\uc9c0 \uc54a\ub294 \uacc4\uce35(layer)\uc740 \uba54\ubaa8\ub9ac \ud615\uc2dd \uc804\ud30c\ub97c \uba48\ucd94\uac8c \ub429\ub2c8\ub2e4. \uadf8\ub7fc\uc5d0\ub3c4 \ubd88\uad6c\ud558\uace0, \ubaa8\ub378\uc744 channels last \ud615\uc2dd\uc73c\ub85c \ubcc0\ud658\ud588\uc73c\ubbc0\ub85c, Channels last \uba54\ubaa8\ub9ac \ud615\uc2dd\uc73c\ub85c 4\ucc28\uc6d0\uc758 \uac00\uc911\uce58\ub97c \uac16\ub294 \uac01 \ud569\uc131\uacf1 \uacc4\uce35(convolution layer)\uc5d0\uc11c\ub294 Channels last \ud615\uc2dd\uc73c\ub85c \ubcf5\uc6d0\ub418\uace0 \ub354 \ube60\ub978 \ucee4\ub110(faster kernel)\uc758 \uc774\uc810\uc744 \ub204\ub9b4 \uc218 \uc788\uac8c \ub429\ub2c8\ub2e4. \ud558\uc9c0\ub9cc Channels last\ub97c \uc9c0\uc6d0\ud558\uc9c0 \uc54a\ub294 \uc5f0\uc0b0\uc790\ub4e4\uc740 \uce58\ud658(permutation)\uc5d0 \uc758\ud574 \uacfc\ubd80\ud558\uac00 \ubc1c\uc0dd\ud558\uac8c \ub429\ub2c8\ub2e4. \uc120\ud0dd\uc801\uc73c\ub85c, \ubcc0\ud658\ub41c \ubaa8\ub378\uc758 \uc131\ub2a5\uc744 \ud5a5\uc0c1\uc2dc\ud0a4\uace0 \uc2f6\uc740 \uacbd\uc6b0 \ubaa8\ub378\uc758 \uc5f0\uc0b0\uc790\ub4e4 \uc911 channel last\ub97c \uc9c0\uc6d0\ud558\uc9c0 \uc54a\ub294 \uc5f0\uc0b0\uc790\ub97c \uc870\uc0ac\ud558\uace0 \uc2dd\ubcc4\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uc774\ub294 Channel Last \uc9c0\uc6d0 \uc5f0\uc0b0\uc790 \ubaa9\ub85d pytorch/pytorch \uc5d0\uc11c \uc0ac\uc6a9\ud55c \uc5f0\uc0b0\uc790\ub4e4\uc774 \uc874\uc7ac\ud558\ub294\uc9c0 \ud655\uc778\ud558\uac70\ub098, eager \uc2e4\ud589 \ubaa8\ub4dc\uc5d0\uc11c \uba54\ubaa8\ub9ac \ud615\uc2dd \uac80\uc0ac\ub97c \ub3c4\uc785\ud558\uace0 \ubaa8\ub378\uc744 \uc2e4\ud589\ud574\uc57c \ud569\ub2c8\ub2e4. \uc544\ub798 \ucf54\ub4dc\uc5d0\uc11c, \uc5f0\uc0b0\uc790\ub4e4\uc758 \ucd9c\ub825\uc774 \uc785\ub825\uc758 \uba54\ubaa8\ub9ac \ud615\uc2dd\uacfc \uc77c\uce58\ud558\uc9c0 \uc54a\uc73c\uba74 \uc608\uc678(exception)\ub97c \ubc1c\uc0dd\uc2dc\ud0b5\ub2c8\ub2e4. def contains_cl(args): for t in args: if isinstance(t, torch.Tensor): if t.is_contiguous(memory_format=torch.channels_last) and not t.is_contiguous(): return True elif isinstance(t, list) or isinstance(t, tuple): if contains_cl(list(t)): return True return False def print_inputs(args, indent=\"\"): for t in args: if isinstance(t, torch.Tensor): print(indent, t.stride(), t.shape, t.device, t.dtype) elif isinstance(t, list) or isinstance(t, tuple): print(indent, type(t)) print_inputs(list(t), indent=indent + \" \") else: print(indent, t) def check_wrapper(fn): name = fn.__name__ def check_cl(*args, **kwargs): was_cl = contains_cl(args) try: result = fn(*args, **kwargs) except Exception as e: print(\"`{}` inputs are:\".format(name)) print_inputs(args) print(\"-------------------\") raise e failed = False if was_cl: if isinstance(result, torch.Tensor): if result.dim() == 4 and not result.is_contiguous(memory_format=torch.channels_last): print( \"`{}` got channels_last input, but output is not channels_last:\".format(name), result.shape, result.stride(), result.device, result.dtype, ) failed = True if failed and True: print(\"`{}` inputs are:\".format(name)) print_inputs(args) raise Exception(\"Operator `{}` lost channels_last property\".format(name)) return result return check_cl old_attrs = dict() def attribute(m): old_attrs[m] = dict() for i in dir(m): e = getattr(m, i) exclude_functions = [\"is_cuda\", \"has_names\", \"numel\", \"stride\", \"Tensor\", \"is_contiguous\", \"__class__\"] if i not in exclude_functions and not i.startswith(\"_\") and \"__call__\" in dir(e): try: old_attrs[m][i] = e setattr(m, i, check_wrapper(e)) except Exception as e: print(i) print(e) attribute(torch.Tensor) attribute(torch.nn.functional) attribute(torch) \ub9cc\uc57d Channels last \ud150\uc11c\ub97c \uc9c0\uc6d0\ud558\uc9c0 \uc54a\ub294 \uc5f0\uc0b0\uc790\ub97c \ubc1c\uacac\ud558\uc600\uace0, \uae30\uc5ec\ud558\uae30\ub97c \uc6d0\ud55c\ub2e4\uba74 \ub2e4\uc74c \uac1c\ubc1c \ubb38\uc11c\ub97c \ucc38\uace0\ud574\uc8fc\uc138\uc694. pytorch/pytorch \uc544\ub798 \ucf54\ub4dc\ub294 torch\uc758 \uc18d\uc131(attributes)\ub97c \ubcf5\uc6d0\ud569\ub2c8\ub2e4. for (m, attrs) in old_attrs.items(): for (k,v) in attrs.items(): setattr(m, k, v) \ud574\uc57c\ud560 \uc77c# \ub2e4\uc74c\uacfc \uac19\uc774 \uc5ec\uc804\ud788 \ud574\uc57c \ud560 \uc77c\uc774 \ub9ce\uc774 \ub0a8\uc544\uc788\uc2b5\ub2c8\ub2e4: N1HW \uc640 NC11 Tensors\uc758 \ubaa8\ud638\uc131 \ud574\uacb0\ud558\uae30; \ubd84\uc0b0 \ud559\uc2b5\uc744 \uc9c0\uc6d0\ud558\ub294\uc9c0 \ud655\uc778\ud558\uae30; \uc5f0\uc0b0\uc790 \ubc94\uc704(operators coverage) \uac1c\uc120(improve)\ud558\uae30 \uac1c\uc120\ud560 \ubd80\ubd84\uc5d0 \ub300\ud55c \ud53c\ub4dc\ubc31 \ub610\ub294 \uc81c\uc548\uc774 \uc788\ub2e4\uba74 \uc774\uc288\ub97c \ub9cc\ub4e4\uc5b4 \uc54c\ub824\uc8fc\uc138\uc694. \uacb0\ub860# \uc774 \ud29c\ud1a0\ub9ac\uc5bc\uc5d0\uc11c\ub294 \u201cchannels last\u201d \uba54\ubaa8\ub9ac \ud615\uc2dd\uc744 \uc18c\uac1c\ud558\uace0 \uc131\ub2a5 \ud5a5\uc0c1\uc744 \uc704\ud574 \uc5b4\ub5bb\uac8c \uc0ac\uc6a9\ud560 \uc218 \uc788\ub294\uc9c0 \uc0b4\ud3b4\ubcf4\uc558\uc2b5\ub2c8\ub2e4. Channels last\ub97c \uc0ac\uc6a9\ud558\uc5ec \ube44\uc804 \ubaa8\ub378\uc744 \uac00\uc18d\ud654\ud558\ub294 \uc2e4\uc9c8\uc801\uc778 \uc608\uc2dc\ub294 \uc5ec\uae30 \ub97c \ucc38\uace0\ud558\uc138\uc694. Total running time of the script: (0 minutes 5.257 seconds) Download Jupyter notebook: memory_format_tutorial.ipynb Download Python source code: memory_format_tutorial.py Download zipped: memory_format_tutorial.zip",
       "author": {
         "@type": "Organization",
         "name": "PyTorch Contributors",
         "url": "https://pytorch.org"
       },
       "image": "../_static/img/pytorch_seo.png",
       "mainEntityOfPage": {
         "@type": "WebPage",
         "@id": "/intermediate/memory_format_tutorial.html"
       },
       "datePublished": "2023-01-01T00:00:00Z",
       "dateModified": "2023-01-01T00:00:00Z"
     }
 

article:modified_time2022-11-30T07:09:41+00:00
og:typearticle
og:site_namePyTorch Tutorials KR
og:image../_static/img/pytorch_seo.png
og:image:altPyTorch Tutorials KR
og:ignore_canonicaltrue
docsearch:languageko
docbuild:last-update2022년 11월 30일
None2
pytorch_projecttutorials

Links:

https://pytorch.kr/
PyTorch 시작하기 https://pytorch.kr/get-started/locally/
기본 익히기 https://tutorials.pytorch.kr/beginner/basics/intro.html
한국어 튜토리얼 https://tutorials.pytorch.kr/
한국어 모델 허브 https://pytorch.kr/hub/
Official Tutorials https://docs.pytorch.org/tutorials/
블로그 https://pytorch.kr/blog/
PyTorch API https://docs.pytorch.org/docs/
Domain API 소개 https://pytorch.kr/domains/
한국어 튜토리얼 https://tutorials.pytorch.kr/
Official Tutorials https://docs.pytorch.org/tutorials/
한국어 커뮤니티 https://discuss.pytorch.kr/
개발자 정보 https://pytorch.kr/resources/
Landscape https://landscape.pytorch.org/
https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html
https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html
PyTorch 시작하기https://pytorch.kr/get-started/locally/
기본 익히기https://tutorials.pytorch.kr/beginner/basics/intro.html
한국어 튜토리얼https://tutorials.pytorch.kr/
한국어 모델 허브https://pytorch.kr/hub/
Official Tutorialshttps://docs.pytorch.org/tutorials/
블로그https://pytorch.kr/blog/
PyTorch APIhttps://docs.pytorch.org/docs/
Domain API 소개https://pytorch.kr/domains/
한국어 튜토리얼https://tutorials.pytorch.kr/
Official Tutorialshttps://docs.pytorch.org/tutorials/
한국어 커뮤니티https://discuss.pytorch.kr/
개발자 정보https://pytorch.kr/resources/
Landscapehttps://landscape.pytorch.org/
Skip to main contenthttps://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#main-content
v2.8.0+cu128https://tutorials.pytorch.kr/index.html
Intro https://tutorials.pytorch.kr/intro.html
Compilers https://tutorials.pytorch.kr/compilers_index.html
Domains https://tutorials.pytorch.kr/domains.html
Distributed https://tutorials.pytorch.kr/distributed.html
Deep Dive https://tutorials.pytorch.kr/deep-dive.html
Extension https://tutorials.pytorch.kr/extension.html
Ecosystem https://tutorials.pytorch.kr/ecosystem.html
Recipes https://tutorials.pytorch.kr/recipes_index.html
한국어 튜토리얼 GitHub 저장소https://github.com/PyTorchKorea/tutorials-kr
파이토치 한국어 커뮤니티https://discuss.pytorch.kr/
Intro https://tutorials.pytorch.kr/intro.html
Compilers https://tutorials.pytorch.kr/compilers_index.html
Domains https://tutorials.pytorch.kr/domains.html
Distributed https://tutorials.pytorch.kr/distributed.html
Deep Dive https://tutorials.pytorch.kr/deep-dive.html
Extension https://tutorials.pytorch.kr/extension.html
Ecosystem https://tutorials.pytorch.kr/ecosystem.html
Recipes https://tutorials.pytorch.kr/recipes_index.html
한국어 튜토리얼 GitHub 저장소https://github.com/PyTorchKorea/tutorials-kr
파이토치 한국어 커뮤니티https://discuss.pytorch.kr/
PyTorch 모듈 프로파일링하기https://tutorials.pytorch.kr/beginner/profiler.html
Parametrizations Tutorialhttps://tutorials.pytorch.kr/intermediate/parametrizations.html
가지치기 기법(Pruning) 튜토리얼https://tutorials.pytorch.kr/intermediate/pruning_tutorial.html
Inductor CPU backend debugging and profilinghttps://tutorials.pytorch.kr/intermediate/inductor_debug_cpu.html
(Beta) Scaled Dot Product Attention (SDPA)로 고성능 트랜스포머(Transformers) 구현하기https://tutorials.pytorch.kr/intermediate/scaled_dot_product_attention_tutorial.html
Knowledge Distillation Tutorialhttps://tutorials.pytorch.kr/beginner/knowledge_distillation_tutorial.html
(베타) PyTorch를 사용한 Channels Last 메모리 형식https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html
Forward-mode Automatic Differentiation (Beta)https://tutorials.pytorch.kr/intermediate/forward_ad_usage.html
Jacobians, Hessians, hvp, vhp, and more: composing function transformshttps://tutorials.pytorch.kr/intermediate/jacobians_hessians.html
모델 앙상블https://tutorials.pytorch.kr/intermediate/ensembling.html
Per-sample-gradientshttps://tutorials.pytorch.kr/intermediate/per_sample_grads.html
PyTorch C++ 프론트엔드 사용하기https://tutorials.pytorch.kr/advanced/cpp_frontend.html
C++ 프론트엔드의 자동 미분 (autograd)https://tutorials.pytorch.kr/advanced/cpp_autograd.html
https://tutorials.pytorch.kr/index.html
Deep Divehttps://tutorials.pytorch.kr/deep-dive.html
Go to the endhttps://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#sphx-glr-download-intermediate-memory-format-tutorial-py
#https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#pytorch-channels-last
Vitaly Fedyuninhttps://github.com/VitalyFedyunin
Choi Yoonjeonghttps://github.com/potatochips178
#https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#memory-format-api
#https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#id1
NVIDIA/apexhttps://github.com/NVIDIA/apex
#https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#id2
pytorch/pytorchhttps://github.com/pytorch/pytorch/wiki/Operators-with-Channels-Last-support
pytorch/pytorchhttps://github.com/pytorch/pytorch/wiki/Writing-memory-format-aware-operators
#https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#id3
이슈를 만들어https://github.com/pytorch/pytorch/issues
#https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#id5
여기https://pytorch.org/blog/accelerating-pytorch-vision-models-with-channels-last-on-cpu/
Download Jupyter notebook: memory_format_tutorial.ipynbhttps://tutorials.pytorch.kr/_downloads/f11c58c36c9b8a5daf09d3f9a792ef84/memory_format_tutorial.ipynb
Download Python source code: memory_format_tutorial.pyhttps://tutorials.pytorch.kr/_downloads/591028d309d0401740cd71eb6b14bf93/memory_format_tutorial.py
Download zipped: memory_format_tutorial.ziphttps://tutorials.pytorch.kr/_downloads/5597f45e0443e4b44273c796119bef90/memory_format_tutorial.zip
이전 Knowledge Distillation Tutorial https://tutorials.pytorch.kr/beginner/knowledge_distillation_tutorial.html
다음 Forward-mode Automatic Differentiation (Beta) https://tutorials.pytorch.kr/intermediate/forward_ad_usage.html
PyData Sphinx Themehttps://pydata-sphinx-theme.readthedocs.io/en/stable/index.html
이전 Knowledge Distillation Tutorial https://tutorials.pytorch.kr/beginner/knowledge_distillation_tutorial.html
다음 Forward-mode Automatic Differentiation (Beta) https://tutorials.pytorch.kr/intermediate/forward_ad_usage.html
메모리 형식(Memory Format) APIhttps://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#memory-format-api
성능 향상https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#id1
기존 모델들 변환하기https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#id2
해야할 일https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#id3
결론https://tutorials.pytorch.kr/intermediate/memory_format_tutorial.html#id5
torchaohttps://docs.pytorch.org/ao
torchrechttps://docs.pytorch.org/torchrec
torchfthttps://docs.pytorch.org/torchft
TorchCodechttps://docs.pytorch.org/torchcodec
torchvisionhttps://docs.pytorch.org/vision
ExecuTorchhttps://docs.pytorch.org/executorch
PyTorch on XLA Deviceshttps://docs.pytorch.org/xla
GitHub로 이동https://github.com/PyTorchKorea
튜토리얼로 이동https://tutorials.pytorch.kr/
커뮤니티로 이동https://discuss.pytorch.kr/
https://pytorch.kr/
파이토치 한국 사용자 모임https://pytorch.kr/
사용자 모임 소개https://pytorch.kr/about
기여해주신 분들https://pytorch.kr/contributors
리소스https://pytorch.kr/resources/
행동 강령https://pytorch.kr/coc
행동 강령https://pytorch.kr/coc
Linux Foundation의 정책https://www.linuxfoundation.org/policies/
our code of conducthttps://pytorch.kr/coc
Linux Foundation's policieshttps://www.linuxfoundation.org/policies/
Cookies Policyhttps://www.facebook.com/policies/cookies/
Sphinxhttps://www.sphinx-doc.org/
PyData Sphinx Themehttps://pydata-sphinx-theme.readthedocs.io/en/stable/index.html

Viewport: width=device-width, initial-scale=1


URLs of crawlers that visited me.