Title: 기초부터 시작하는 NLP: 문자-단위 RNN으로 이름 분류하기 — 파이토치 한국어 튜토리얼 (PyTorch tutorials in Korean)
Open Graph Title: 기초부터 시작하는 NLP: 문자-단위 RNN으로 이름 분류하기
Description: Author: Sean Robertson, 번역: 황성수, 김제필,. 이 튜토리얼은 3부로 구성된 시리즈의 일부입니다: 기초부터 시작하는 NLP: 문자-단위 RNN으로 이름 분류하기, 기초부터 시작하는 NLP: 문자-단위 RNN으로 이름 생성하기, 기초부터 시작하는 NLP: Sequence to Sequence 네트워크와 Attention을 이용한 번역. 여기에서는 단어를 분류하기 위해 기초적인 문자-단위의 순환 신경망(RNN, Recurrent Nueral Network)을 구축하고 학습할 예정입니다. 이 튜토리얼 및 이후 ...
Open Graph Description: Author: Sean Robertson, 번역: 황성수, 김제필,. 이 튜토리얼은 3부로 구성된 시리즈의 일부입니다: 기초부터 시작하는 NLP: 문자-단위 RNN으로 이름 분류하기, 기초부터 시작하는 NLP: 문자-단위 RNN으로 이름 생성하기, 기초부터 시작하는 NLP: Sequence to Sequence 네트워크와 Attention을 이용한 번역. 여기에서는 단어를 분류하기 위해 기초적인 문자-단위의 순환 신경망(RNN, Recurrent Nueral Network)을 구축하고 학습할 예정입니다. 이 튜토리얼 및 이후 ...
Opengraph URL: https://tutorials.pytorch.kr/intermediate/char_rnn_classification_tutorial.html
Domain: tutorials.pytorch.kr
{
"@context": "https://schema.org",
"@type": "Article",
"name": "\uae30\ucd08\ubd80\ud130 \uc2dc\uc791\ud558\ub294 NLP: \ubb38\uc790-\ub2e8\uc704 RNN\uc73c\ub85c \uc774\ub984 \ubd84\ub958\ud558\uae30",
"headline": "\uae30\ucd08\ubd80\ud130 \uc2dc\uc791\ud558\ub294 NLP: \ubb38\uc790-\ub2e8\uc704 RNN\uc73c\ub85c \uc774\ub984 \ubd84\ub958\ud558\uae30",
"description": "PyTorch Documentation. Explore PyTorch, an open-source machine learning library that accelerates the path from research prototyping to production deployment. Discover tutorials, API references, and guides to help you build and deploy deep learning models efficiently.",
"url": "/intermediate/char_rnn_classification_tutorial.html",
"articleBody": "\ucc38\uace0 Go to the end to download the full example code. \uae30\ucd08\ubd80\ud130 \uc2dc\uc791\ud558\ub294 NLP: \ubb38\uc790-\ub2e8\uc704 RNN\uc73c\ub85c \uc774\ub984 \ubd84\ub958\ud558\uae30# Author: Sean Robertson\ubc88\uc5ed: \ud669\uc131\uc218, \uae40\uc81c\ud544 \uc774 \ud29c\ud1a0\ub9ac\uc5bc\uc740 3\ubd80\ub85c \uad6c\uc131\ub41c \uc2dc\ub9ac\uc988\uc758 \uc77c\ubd80\uc785\ub2c8\ub2e4: \uae30\ucd08\ubd80\ud130 \uc2dc\uc791\ud558\ub294 NLP: \ubb38\uc790-\ub2e8\uc704 RNN\uc73c\ub85c \uc774\ub984 \ubd84\ub958\ud558\uae30 \uae30\ucd08\ubd80\ud130 \uc2dc\uc791\ud558\ub294 NLP: \ubb38\uc790-\ub2e8\uc704 RNN\uc73c\ub85c \uc774\ub984 \uc0dd\uc131\ud558\uae30 \uae30\ucd08\ubd80\ud130 \uc2dc\uc791\ud558\ub294 NLP: Sequence to Sequence \ub124\ud2b8\uc6cc\ud06c\uc640 Attention\uc744 \uc774\uc6a9\ud55c \ubc88\uc5ed \uc5ec\uae30\uc5d0\uc11c\ub294 \ub2e8\uc5b4\ub97c \ubd84\ub958\ud558\uae30 \uc704\ud574 \uae30\ucd08\uc801\uc778 \ubb38\uc790-\ub2e8\uc704\uc758 \uc21c\ud658 \uc2e0\uacbd\ub9dd(RNN, Recurrent Nueral Network)\uc744 \uad6c\ucd95\ud558\uace0 \ud559\uc2b5\ud560 \uc608\uc815\uc785\ub2c8\ub2e4. \uc774 \ud29c\ud1a0\ub9ac\uc5bc \ubc0f \uc774\ud6c4 2\uac1c \ud29c\ud1a0\ub9ac\uc5bc\uc778 \uae30\ucd08\ubd80\ud130 \uc2dc\uc791\ud558\ub294 NLP: \ubb38\uc790-\ub2e8\uc704 RNN\uc73c\ub85c \uc774\ub984 \uc0dd\uc131\ud558\uae30 \ubc0f \uae30\ucd08\ubd80\ud130 \uc2dc\uc791\ud558\ub294 NLP: Sequence to Sequence \ub124\ud2b8\uc6cc\ud06c\uc640 Attention\uc744 \uc774\uc6a9\ud55c \ubc88\uc5ed \uc5d0\uc11c\ub294 \uc790\uc5f0\uc5b4 \ucc98\ub9ac(NLP, Natural Language Processing) \ubd84\uc57c\uc5d0\uc11c \uc5b4\ub5bb\uac8c \ub370\uc774\ud130\ub97c \uc804\ucc98\ub9ac\ud558\uace0 NLP \ubaa8\ub378\uc744 \uad6c\ucd95\ud558\ub294\uc9c0\ub97c \ubc11\ubc14\ub2e5\ubd80\ud130(from scratch) \uc124\uba85\ud569\ub2c8\ub2e4. \uc774\ub97c \uc704\ud574 \uc774 \ud29c\ud1a0\ub9ac\uc5bc \uc2dc\ub9ac\uc988\uc5d0\uc11c\ub294 NLP \ubaa8\ub378\ub9c1\uc744 \uc704\ud55c \ub370\uc774\ud130 \uc804\ucc98\ub9ac\uac00 \ubc11\ubc14\ub2e5(low-level)\uc5d0\uc11c \uc5b4\ub5bb\uac8c \uc9c4\ud589\ub418\ub294\uc9c0 \uc54c \uc218 \uc788\uc2b5\ub2c8\ub2e4. \ubb38\uc790-\ub2e8\uc704 RNN\uc740 \ub2e8\uc5b4\ub97c \ubb38\uc790\uc758 \uc5f0\uc18d\uc73c\ub85c \uc77d\uc5b4 \ub4e4\uc5ec\uc11c \uac01 \ub2e8\uacc4\uc758 \uc608\uce21\uacfc \u201c\uc740\ub2c9 \uc0c1\ud0dc(Hidden State)\u201d\ub97c \ucd9c\ub825\ud558\uace0, \ub2e4\uc74c \ub2e8\uacc4\uc5d0 \uc774\uc804 \ub2e8\uacc4\uc758 \uc740\ub2c9 \uc0c1\ud0dc\ub97c \uc804\ub2ec\ud569\ub2c8\ub2e4. \ub2e8\uc5b4\uac00 \uc18d\ud55c \ud074\ub798\uc2a4\ub85c \ucd9c\ub825\ub418\ub3c4\ub85d \ucd5c\uc885 \uc608\uce21\uc73c\ub85c \uc120\ud0dd\ud569\ub2c8\ub2e4. \uad6c\uccb4\uc801\uc73c\ub85c, 18\uac1c \uc5b8\uc5b4\ub85c \ub41c \uc218\ucc9c \uac1c\uc758 \uc131(\u59d3)\uc744 \ud6c8\ub828\uc2dc\ud0a4\uace0, \ucca0\uc790\uc5d0 \ub530\ub77c \uc774\ub984\uc774 \uc5b4\ub5a4 \uc5b8\uc5b4\uc778\uc9c0 \uc608\uce21\ud569\ub2c8\ub2e4. Torch \uc900\ube44# \ud558\ub4dc\uc6e8\uc5b4(CPU \ub610\ub294 CUDA)\uc5d0 \ub9de\ucdb0 GPU \uac00\uc18d\uc744 \uc0ac\uc6a9\ud560 \uc218 \uc788\ub3c4\ub85d \uc801\uc808\ud55c \uc7a5\uce58\ub97c \uae30\ubcf8 \uc7a5\uce58\ub85c \uc124\uc815\ud569\ub2c8\ub2e4. import torch # Check if CUDA is available device = torch.device(\u0027cpu\u0027) if torch.cuda.is_available(): device = torch.device(\u0027cuda\u0027) torch.set_default_device(device) print(f\"Using device = {torch.get_default_device()}\") Using device = cuda:0 \ub370\uc774\ud130 \uc900\ube44# \uc5ec\uae30 \uc5d0\uc11c \ub370\uc774\ud130\ub97c \ub2e4\uc6b4\ub85c\ub4dc \ubc1b\uace0 \ud604\uc7ac \ub514\ub809\ud1a0\ub9ac\uc5d0 \uc555\ucd95\uc744 \ud489\ub2c8\ub2e4. data/names \ub514\ub809\ud1a0\ub9ac\uc5d0\ub294 [Language].txt \ub77c\ub294 18 \uac1c\uc758 \ud14d\uc2a4\ud2b8 \ud30c\uc77c\uc774 \uc788\uc2b5\ub2c8\ub2e4. \uac01 \ud30c\uc77c\uc5d0\ub294 \ud55c \uc904\uc5d0 \ud558\ub098\uc758 \uc774\ub984\uc774 \ud3ec\ud568\ub418\uc5b4 \uc788\uc73c\uba70 \ub300\ubd80\ubd84 \ub85c\ub9c8\uc790\ub85c \ub418\uc5b4 \uc788\uc2b5\ub2c8\ub2e4. (\ud558\uc9c0\ub9cc \uc720\ub2c8\ucf54\ub4dc\uc5d0\uc11c ASCII\ub85c \ubcc0\ud658\uc740 \ud574\uc57c \ud569\ub2c8\ub2e4) \uccab\ubc88\uc9f8 \ub2e8\uacc4\ub294 \ub370\uc774\ud130\ub97c \uc815\uc758\ud558\uace0 \uc815\ub9ac\ud558\ub294 \uac83\uc785\ub2c8\ub2e4. \ucd08\uae30\uc5d0\ub294 \uc720\ub2c8\ucf54\ub4dc\ub97c \uc77c\ubc18 ASCII\ub85c \ubcc0\ud658\ud558\uc5ec RNN \uc785\ub825 \ub808\uc774\uc5b4\ub97c \uc81c\ud55c\ud574\uc57c \ud569\ub2c8\ub2e4. \uc774\ub294 \uc720\ub2c8\ucf54\ub4dc \ubb38\uc790\uc5f4\uc744 ASCII\ub85c \ubcc0\ud658\ud558\uace0 \ud5c8\uc6a9\ub41c \ubb38\uc790\uc758 \uc791\uc740 \uc9d1\ud569\ub9cc\uc744 \ud5c8\uc6a9\ud558\uc5ec \uc774\ub8e8\uc5b4\uc9d1\ub2c8\ub2e4. import string import unicodedata # \"_\" \ub97c \uc0ac\uc6a9\ud558\uc5ec \uc5b4\ud718\uc9d1(Vocabulary)\uc5d0 \uc5c6\ub294 \ubb38\uc790\ub97c \ud45c\ud604\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uc989, \ubaa8\ub378\uc5d0\uc11c \ucc98\ub9ac\ud558\uc9c0 \uc54a\ub294 \ubaa8\ub4e0 \ubb38\uc790\ub97c \ud45c\ud604\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. allowed_characters = string.ascii_letters + \" .,;\u0027\" + \"_\" n_letters = len(allowed_characters) # \uc720\ub2c8\ucf54\ub4dc \ubb38\uc790\uc5f4\uc744 \uc77c\ubc18 ASCII\ub85c \ubcc0\ud658\ud558\uae30: https://stackoverflow.com/a/518232/2809427 def unicodeToAscii(s): return \u0027\u0027.join( c for c in unicodedata.normalize(\u0027NFD\u0027, s) if unicodedata.category(c) != \u0027Mn\u0027 and c in allowed_characters ) \uc720\ub2c8\ucf54\ub4dc \uc54c\ud30c\ubcb3 \uc774\ub984\uc744 \uc77c\ubc18 ASCII\ub85c \ubcc0\ud658\ud558\ub294 \uc608\uc2dc\uc785\ub2c8\ub2e4. \uc774\ub807\uac8c \ud558\uba74 \uc785\ub825 \ub808\uc774\uc5b4\ub97c \ub2e8\uc21c\ud654\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. print (f\"converting \u0027\u015alus\u00e0rski\u0027 to {unicodeToAscii(\u0027\u015alus\u00e0rski\u0027)}\") converting \u0027\u015alus\u00e0rski\u0027 to Slusarski \uc774\ub984\uc744 Tensor\ub85c \ubcc0\uacbd# \uc774\uc81c \ubaa8\ub4e0 \uc774\ub984\uc744 \uccb4\uacc4\ud654\ud588\uc73c\ubbc0\ub85c, \uc774\ub97c \ud65c\uc6a9\ud558\uae30 \uc704\ud574 Tensor\ub85c \ubcc0\ud658\ud574\uc57c \ud569\ub2c8\ub2e4. \ud558\ub098\uc758 \ubb38\uc790\ub97c \ud45c\ud604\ud558\uae30 \uc704\ud574 \ud06c\uae30\uac00 \u003c1 x n_letters\u003e \uc778 \u201cOne-Hot \ubca1\ud130\u201d\ub97c \uc0ac\uc6a9\ud569\ub2c8\ub2e4. One-Hot \ubca1\ud130\ub294 \ud604\uc7ac \ubb38\uc790\uc758 \uc8fc\uc18c\uc5d0\ub294 1\uc774, \uadf8 \uc678 \ub098\uba38\uc9c0 \uc8fc\uc18c\uc5d0\ub294 0\uc774 \ucc44\uc6cc\uc9c4 \ubca1\ud130\uc785\ub2c8\ub2e4. \uc608\uc2dc \"b\" = \u003c0 1 0 0 0 ...\u003e . \ub2e8\uc5b4\ub97c \ub9cc\ub4e4\uae30 \uc704\ud574 One-Hot \ubca1\ud130\ub4e4\uc744 2\ucc28\uc6d0 \ud589\ub82c \u003cline_length x 1 x n_letters\u003e \uc5d0 \uacb0\ud569\uc2dc\ud0b5\ub2c8\ub2e4. \uc704\uc5d0\uc11c \ubcf4\uc774\ub294 \ucd94\uac00\uc801\uc778 1\ucc28\uc6d0\uc740 PyTorch\uc5d0\uc11c \ubaa8\ub4e0 \uac83\uc774 \ubc30\uce58(batch)\uc5d0 \uc788\ub2e4\uace0 \uac00\uc815\ud558\uae30 \ub54c\ubb38\uc5d0 \ubc1c\uc0dd\ud569\ub2c8\ub2e4. \uc5ec\uae30\uc11c\ub294 \ubc30\uce58 \ud06c\uae30 1\uc744 \uc0ac\uc6a9\ud558\uace0 \uc788\uc2b5\ub2c8\ub2e4. # .. note:: # \uc5ed\uc790 \uc8fc: One-Hot \ubca1\ud130\ub294 \uc5b8\uc5b4 \ubc0f \ubc94\uc8fc\ud615 \ub370\uc774\ud130\ub97c \ub2e4\ub8f0 \ub54c \uc8fc\ub85c \uc0ac\uc6a9\ud558\uba70, # \ub2e8\uc5b4, \uae00\uc790 \ub4f1\uc744 \ubca1\ud130\ub85c \ud45c\ud604\ud560 \ub54c \ub2e8\uc5b4, \uae00\uc790 \uc0ac\uc774\uc758 \uc0c1\uad00 \uad00\uacc4\ub97c \ubbf8\ub9ac \uc54c \uc218 \uc5c6\uc744 \uacbd\uc6b0, # One-Hot\uc73c\ub85c \ud45c\ud604\ud558\uc5ec \uc11c\ub85c \uc9c1\uad50\ud55c\ub2e4\uace0 \uac00\uc815\ud558\uace0 \ud559\uc2b5\uc744 \uc2dc\uc791\ud569\ub2c8\ub2e4. # \uc774\uc640 \ub3d9\uc77c\ud558\uac8c, \uc0c1\uad00 \uad00\uacc4\ub97c \uc54c \uc218 \uc5c6\ub294 \ub2e4\ub978 \ub370\uc774\ud130\uc758 \uacbd\uc6b0\uc5d0\ub3c4 One-Hot \ubca1\ud130\ub97c \ud65c\uc6a9\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. # import torch # all_letters \ub85c \ubb38\uc790\uc758 \uc8fc\uc18c \ucc3e\uae30, \uc608\uc2dc \"a\" = 0 def letterToIndex(letter): # \ubaa8\ub378\uc774 \ubaa8\ub974\ub294 \uae00\uc790\ub97c \ub9cc\ub098\uba74, \uc5b4\ud718\uc9d1\uc5d0 \uc874\uc7ac\ud558\uc9c0 \uc54a\ub294 \ubb38\uc790(\"_\")\ub97c \ubc18\ud658\ud569\ub2c8\ub2e4. if letter not in allowed_characters: return allowed_characters.find(\"_\") else: return allowed_characters.find(letter) # \uac80\uc99d\uc744 \uc704\ud574\uc11c \ud55c \uac1c\uc758 \ubb38\uc790\ub97c \u003c1 x n_letters\u003e Tensor\ub85c \ubcc0\ud658 def letterToTensor(letter): tensor = torch.zeros(1, n_letters) tensor[0][letterToIndex(letter)] = 1 return tensor # \ud55c \uc904(\uc774\ub984)\uc744 \u003cline_length x 1 x n_letters\u003e, # \ub610\ub294 One-Hot \ubb38\uc790 \ubca1\ud130\uc758 Array\ub85c \ubcc0\uacbd def lineToTensor(line): tensor = torch.zeros(len(line), 1, n_letters) for li, letter in enumerate(line): tensor[li][0][letterToIndex(letter)] = 1 return tensor print(letterToTensor(\u0027J\u0027)) print(lineToTensor(\u0027Jones\u0027).size()) tensor([[0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]], device=\u0027cuda:0\u0027) torch.Size([5, 1, 58]) Here are some examples of how to use lineToTensor() for a single and multiple character string. print (f\"The letter \u0027a\u0027 becomes {lineToTensor(\u0027a\u0027)}\") #notice that the first position in the tensor = 1 print (f\"The name \u0027Ahn\u0027 becomes {lineToTensor(\u0027Ahn\u0027)}\") #notice \u0027A\u0027 sets the 27th index to 1 The letter \u0027a\u0027 becomes tensor([[[1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]]], device=\u0027cuda:0\u0027) The name \u0027Ahn\u0027 becomes tensor([[[0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]], [[0., 0., 0., 0., 0., 0., 0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]], [[0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]]], device=\u0027cuda:0\u0027) Congratulations, you have built the foundational tensor objects for this learning task! You can use a similar approach for other RNN tasks with text. Next, we need to combine all our examples into a dataset so we can train, test and validate our models. For this, we will use the Dataset and DataLoader classes to hold our dataset. Each Dataset needs to implement three functions: __init__, __len__, and __getitem__. from io import open import glob import os import time import torch from torch.utils.data import Dataset class NamesDataset(Dataset): def __init__(self, data_dir): self.data_dir = data_dir #for provenance of the dataset self.load_time = time.localtime #for provenance of the dataset labels_set = set() #set of all classes self.data = [] self.data_tensors = [] self.labels = [] self.labels_tensors = [] #read all the ``.txt`` files in the specified directory text_files = glob.glob(os.path.join(data_dir, \u0027*.txt\u0027)) for filename in text_files: label = os.path.splitext(os.path.basename(filename))[0] labels_set.add(label) lines = open(filename, encoding=\u0027utf-8\u0027).read().strip().split(\u0027\\n\u0027) for name in lines: self.data.append(name) self.data_tensors.append(lineToTensor(name)) self.labels.append(label) #Cache the tensor representation of the labels self.labels_uniq = list(labels_set) for idx in range(len(self.labels)): temp_tensor = torch.tensor([self.labels_uniq.index(self.labels[idx])], dtype=torch.long) self.labels_tensors.append(temp_tensor) def __len__(self): return len(self.data) def __getitem__(self, idx): data_item = self.data[idx] data_label = self.labels[idx] data_tensor = self.data_tensors[idx] label_tensor = self.labels_tensors[idx] return label_tensor, data_tensor, data_label, data_item Here we can load our example data into the NamesDataset alldata = NamesDataset(\"data/names\") print(f\"loaded {len(alldata)} items of data\") print(f\"example = {alldata[0]}\") loaded 20074 items of data example = (tensor([1], device=\u0027cuda:0\u0027), tensor([[[0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]], [[0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]], [[0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]], [[1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]], [[0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 1., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0., 0.]]], device=\u0027cuda:0\u0027), \u0027English\u0027, \u0027Abbas\u0027) Using the dataset object allows us to easily split the data into train and test sets. Here we create a 80/20 split but the torch.utils.data has more useful utilities. Here we specify a generator since we need to use the same device as PyTorch defaults to above. train_set, test_set = torch.utils.data.random_split(alldata, [.85, .15], generator=torch.Generator(device=device).manual_seed(2024)) print(f\"train examples = {len(train_set)}, validation examples = {len(test_set)}\") train examples = 17063, validation examples = 3011 Now we have a basic dataset containing 20074 examples where each example is a pairing of label and name. We have also split the dataset into training and testing so we can validate the model that we build. \ub124\ud2b8\uc6cc\ud06c \uc0dd\uc131# Autograd \uc804\uc5d0, Torch\uc5d0\uc11c RNN(recurrent neural network) \uc0dd\uc131\uc740 \uc5ec\ub7ec \uc2dc\uac04 \ub2e8\uacc4 \uac78\uccd0\uc11c \uacc4\uce35\uc758 \ub9e4\uac1c\ubcc0\uc218\ub97c \ubcf5\uc81c\ud558\ub294 \uc791\uc5c5\uc744 \ud3ec\ud568\ud569\ub2c8\ub2e4. \uacc4\uce35\uc740 \uc740\ub2c9 \uc0c1\ud0dc\uc640 \ubcc0\ud654\ub3c4(Gradient)\ub97c \uac00\uc9c0\uba70, \uc774\uc81c \uc774\uac83\ub4e4\uc740 \uadf8\ub798\ud504 \uc790\uccb4\uc5d0\uc11c \uc644\uc804\ud788 \ucc98\ub9ac\ub429\ub2c8\ub2e4. \uc774\ub294 feed-forward \uacc4\uce35\uacfc \uac19\uc740 \ub9e4\uc6b0 \u201c\uc21c\uc218\ud55c\u201d \ubc29\ubc95\uc73c\ub85c RNN\uc744 \uad6c\ud604\ud560 \uc218 \uc788\uc74c\uc744 \uc758\ubbf8\ud569\ub2c8\ub2e4. \ucc38\uace0 \uc5ed\uc790 \uc8fc: \uc5ec\uae30\uc11c\ub294 \ud559\uc2b5 \ubaa9\uc801\uc73c\ub85c nn.RNN \ub300\uc2e0 \uc9c1\uc811 RNN\uc744 \uc0ac\uc6a9\ud569\ub2c8\ub2e4. \uc774 RNN \ubaa8\ub4c8\uc740 \u201c\uae30\ubcf8(vanilla)\uc801\uc778 RNN\u201d\uc744 \uad6c\ud604\ud558\uba70, \uc785\ub825\uacfc \uc740\ub2c9 \uc0c1\ud0dc(hidden state), \uadf8\ub9ac\uace0 \ucd9c\ub825 \ub4a4 \ub3d9\uc791\ud558\ub294 LogSoftmax \uacc4\uce35\uc774 \uc788\ub294 3\uac1c\uc758 \uc120\ud615 \uacc4\uce35\ub9cc\uc744 \uac00\uc9d1\ub2c8\ub2e4. This CharRNN class implements an RNN with three components. First, we use the nn.RNN implementation. Next, we define a layer that maps the RNN hidden layers to our output. And finally, we apply a softmax function. Using nn.RNN leads to a significant improvement in performance, such as cuDNN-accelerated kernels, versus implementing each layer as a nn.Linear. It also simplifies the implementation in forward(). import torch.nn as nn import torch.nn.functional as F class CharRNN(nn.Module): def __init__(self, input_size, hidden_size, output_size): super(CharRNN, self).__init__() self.rnn = nn.RNN(input_size, hidden_size) self.h2o = nn.Linear(hidden_size, output_size) self.softmax = nn.LogSoftmax(dim=1) def forward(self, line_tensor): rnn_out, hidden = self.rnn(line_tensor) output = self.h2o(hidden[0]) output = self.softmax(output) return output We can then create an RNN with 58 input nodes, 128 hidden nodes, and 18 outputs: n_hidden = 128 rnn = CharRNN(n_letters, n_hidden, len(alldata.labels_uniq)) print(rnn) CharRNN( (rnn): RNN(58, 128) (h2o): Linear(in_features=128, out_features=18, bias=True) (softmax): LogSoftmax(dim=1) ) After that we can pass our Tensor to the RNN to obtain a predicted output. Subsequently, we use a helper function, label_from_output, to derive a text label for the class. def label_from_output(output, output_labels): top_n, top_i = output.topk(1) label_i = top_i[0].item() return output_labels[label_i], label_i input = lineToTensor(\u0027Albert\u0027) output = rnn(input) #this is equivalent to ``output = rnn.forward(input)`` print(output) print(label_from_output(output, alldata.labels_uniq)) tensor([[-2.9395, -3.0940, -2.8990, -2.9367, -2.8716, -2.7607, -2.9501, -2.8651, -2.8425, -2.8339, -2.8015, -2.9372, -2.9002, -2.9046, -2.8947, -2.8643, -2.9327, -2.8415]], device=\u0027cuda:0\u0027, grad_fn=\u003cLogSoftmaxBackward0\u003e) (\u0027Greek\u0027, 5) \ud559\uc2b5# \uc2e0\uacbd\ub9dd \ud559\uc2b5# \uc774\uc81c \uc774 \ub124\ud2b8\uc6cc\ud06c\ub97c \ud559\uc2b5\ud558\ub294 \ub370 \ud544\uc694\ud55c \uc608\uc2dc(\ud559\uc2b5 \ub370\uc774\ud130)\ub97c \ubcf4\uc5ec\uc8fc\uace0 \ucd94\uc815\ud569\ub2c8\ub2e4. \ub9cc\uc77c \ud2c0\ub838\ub2e4\uba74 \uc54c\ub824 \uc90d\ub2c8\ub2e4. We do this by defining a train() function which trains the model on a given dataset using minibatches. RNNs RNNs are trained similarly to other networks; therefore, for completeness, we include a batched training method here. The loop (for i in batch) computes the losses for each of the items in the batch before adjusting the weights. This operation is repeated until the number of epochs is reached. import random import numpy as np def train(rnn, training_data, n_epoch = 10, n_batch_size = 64, report_every = 50, learning_rate = 0.2, criterion = nn.NLLLoss()): \"\"\" Learn on a batch of training_data for a specified number of iterations and reporting thresholds \"\"\" # Keep track of losses for plotting current_loss = 0 all_losses = [] rnn.train() optimizer = torch.optim.SGD(rnn.parameters(), lr=learning_rate) start = time.time() print(f\"training on data set with n = {len(training_data)}\") for iter in range(1, n_epoch + 1): rnn.zero_grad() # clear the gradients # create some minibatches # we cannot use dataloaders because each of our names is a different length batches = list(range(len(training_data))) random.shuffle(batches) batches = np.array_split(batches, len(batches) //n_batch_size ) for idx, batch in enumerate(batches): batch_loss = 0 for i in batch: #for each example in this batch (label_tensor, text_tensor, label, text) = training_data[i] output = rnn.forward(text_tensor) loss = criterion(output, label_tensor) batch_loss += loss # optimize parameters batch_loss.backward() nn.utils.clip_grad_norm_(rnn.parameters(), 3) optimizer.step() optimizer.zero_grad() current_loss += batch_loss.item() / len(batch) all_losses.append(current_loss / len(batches) ) if iter % report_every == 0: print(f\"{iter} ({iter / n_epoch:.0%}): \\t average batch loss = {all_losses[-1]}\") current_loss = 0 return all_losses We can now train a dataset with minibatches for a specified number of epochs. The number of epochs for this example is reduced to speed up the build. You can get better results with different parameters. start = time.time() all_losses = train(rnn, train_set, n_epoch=27, learning_rate=0.15, report_every=5) end = time.time() print(f\"training took {end-start}s\") training on data set with n = 17063 5 (19%): average batch loss = 0.8897688550849814 10 (37%): average batch loss = 0.6987682116383276 15 (56%): average batch loss = 0.5792883489287404 20 (74%): average batch loss = 0.4965955759470279 25 (93%): average batch loss = 0.43885368771157285 training took 270.788941860199s \uacb0\uacfc \ub3c4\uc2dd\ud654# all_losses \ub97c \uc774\uc6a9\ud55c \uc190\uc2e4 \ub3c4\uc2dd\ud654\ub294 \ub124\ud2b8\uc6cc\ud06c\uc758 \ud559\uc2b5\uc744 \ubcf4\uc5ec\uc90d\ub2c8\ub2e4: import matplotlib.pyplot as plt import matplotlib.ticker as ticker plt.figure() plt.plot(all_losses) plt.show() \uacb0\uacfc \ud3c9\uac00# \ub124\ud2b8\uc6cc\ud06c\uac00 \ub2e4\ub978 \uce74\ud14c\uace0\ub9ac\uc5d0\uc11c \uc5bc\ub9c8\ub098 \uc798 \uc791\ub3d9\ud558\ub294\uc9c0 \ubcf4\uae30 \uc704\ud574 \ubaa8\ub4e0 \uc2e4\uc81c \uc5b8\uc5b4(\ud589)\uac00 \ub124\ud2b8\uc6cc\ud06c\uc5d0\uc11c \uc5b4\ub5a4 \uc5b8\uc5b4\ub85c \ucd94\uce21(\uc5f4)\ub418\ub294\uc9c0 \ub098\ud0c0\ub0b4\ub294 \ud63c\ub780 \ud589\ub82c(confusion matrix)\uc744 \ub9cc\ub4ed\ub2c8\ub2e4. \ud63c\ub780 \ud589\ub82c\uc744 \uacc4\uc0b0\ud558\uae30 \uc704\ud574 evaluate() \ub85c \ub9ce\uc740 \uc218\uc758 \uc0d8\ud50c\uc744 \ub124\ud2b8\uc6cc\ud06c\uc5d0 \uc2e4\ud589\ud569\ub2c8\ub2e4. evaluate() \uc740 train () \uacfc \uc5ed\uc804\ud30c\ub97c \ube7c\uba74 \ub3d9\uc77c\ud569\ub2c8\ub2e4. def evaluate(rnn, testing_data, classes): confusion = torch.zeros(len(classes), len(classes)) rnn.eval() #set to eval mode with torch.no_grad(): # do not record the gradients during eval phase for i in range(len(testing_data)): (label_tensor, text_tensor, label, text) = testing_data[i] output = rnn(text_tensor) guess, guess_i = label_from_output(output, classes) label_i = classes.index(label) confusion[label_i][guess_i] += 1 # Normalize by dividing every row by its sum for i in range(len(classes)): denom = confusion[i].sum() if denom \u003e 0: confusion[i] = confusion[i] / denom # Set up plot fig = plt.figure() ax = fig.add_subplot(111) cax = ax.matshow(confusion.cpu().numpy()) #numpy uses cpu here so we need to use a cpu version fig.colorbar(cax) # Set up axes ax.set_xticks(np.arange(len(classes)), labels=classes, rotation=90) ax.set_yticks(np.arange(len(classes)), labels=classes) # Force label at every tick ax.xaxis.set_major_locator(ticker.MultipleLocator(1)) ax.yaxis.set_major_locator(ticker.MultipleLocator(1)) # sphinx_gallery_thumbnail_number = 2 plt.show() evaluate(rnn, test_set, classes=alldata.labels_uniq) \uc8fc\ucd95\uc5d0\uc11c \ubc97\uc5b4\ub09c \ubc1d\uc740 \uc810\uc744 \uc120\ud0dd\ud558\uc5ec \uc798\ubabb \ucd94\uce21\ud55c \uc5b8\uc5b4\ub97c \ud45c\uc2dc\ud560 \uc218 \uc788\uc2b5\ub2c8\ub2e4. \uc608\ub97c \ub4e4\uc5b4 \ud55c\uad6d\uc5b4\ub294 \uc911\uad6d\uc5b4\ub85c \uc774\ud0c8\ub9ac\uc544\uc5b4\ub85c \uc2a4\ud398\uc778\uc5b4\ub85c. \uadf8\ub9ac\uc2a4\uc5b4\ub294 \ub9e4\uc6b0 \uc798\ub418\ub294 \uac83\uc73c\ub85c \uc601\uc5b4\ub294 \ub9e4\uc6b0 \ub098\uc05c \uac83\uc73c\ub85c \ubcf4\uc785\ub2c8\ub2e4. (\ub2e4\ub978 \uc5b8\uc5b4\ub4e4\uacfc\uc758 \uc911\ucca9 \ub54c\ubb38\uc73c\ub85c \ucd94\uc815) \uc5f0\uc2b5# Get better results with a bigger and/or better shaped network Adjust the hyperparameters to enhance performance, such as changing the number of epochs, batch size, and learning rate Try the nn.LSTM and nn.GRU layers Modify the size of the layers, such as increasing or decreasing the number of hidden nodes or adding additional linear layers Combine multiple of these RNNs as a higher level network \u201cline -\u003e label\u201d \uc758 \ub2e4\ub978 \ub370\uc774\ud130 \uc9d1\ud569\uc73c\ub85c \uc2dc\ub3c4\ud574 \ubcf4\uc2ed\uc2dc\uc624, \uc608\ub97c \ub4e4\uc5b4: \ub2e8\uc5b4 -\u003e \uc5b8\uc5b4 \uc774\ub984 -\u003e \uc131\ubcc4 \uce90\ub9ad\ud130 \uc774\ub984 -\u003e \uc791\uac00 \ud398\uc774\uc9c0 \uc81c\ubaa9 -\u003e \ube14\ub85c\uadf8 \ub610\ub294 \uc11c\ube0c\ub808\ub527 \ub354 \ud06c\uace0 \ub354 \ub098\uc740 \ubaa8\uc591\uc758 \ub124\ud2b8\uc6cc\ud06c\ub85c \ub354 \ub098\uc740 \uacb0\uacfc\ub97c \uc5bb\uc73c\uc2ed\uc2dc\uc624. \ub354 \ub9ce\uc740 \uc120\ud615 \uacc4\uce35\uc744 \ucd94\uac00\ud574 \ubcf4\uc2ed\uc2dc\uc624. nn.LSTM \uacfc nn.GRU \uacc4\uce35\uc744 \ucd94\uac00\ud574 \ubcf4\uc2ed\uc2dc\uc624. \uc704\uc640 \uac19\uc740 RNN \uc5ec\ub7ec \uac1c\ub97c \uc0c1\uc704 \uc218\uc900 \ub124\ud2b8\uc6cc\ud06c\ub85c \uacb0\ud569\ud574 \ubcf4\uc2ed\uc2dc\uc624. Total running time of the script: (4 minutes 41.329 seconds) Download Jupyter notebook: char_rnn_classification_tutorial.ipynb Download Python source code: char_rnn_classification_tutorial.py Download zipped: char_rnn_classification_tutorial.zip",
"author": {
"@type": "Organization",
"name": "PyTorch Contributors",
"url": "https://pytorch.org"
},
"image": "../_static/img/pytorch_seo.png",
"mainEntityOfPage": {
"@type": "WebPage",
"@id": "/intermediate/char_rnn_classification_tutorial.html"
},
"datePublished": "2023-01-01T00:00:00Z",
"dateModified": "2023-01-01T00:00:00Z"
}
| article:modified_time | 2022-11-30T07:09:41+00:00 |
| og:type | article |
| og:site_name | PyTorch Tutorials KR |
| og:image | ../_static/img/pytorch_seo.png |
| og:image:alt | PyTorch Tutorials KR |
| og:ignore_canonical | true |
| docsearch:language | ko |
| docbuild:last-update | 2022년 11월 30일 |
| None | 1 |
| pytorch_project | tutorials |
Links:
Viewport: width=device-width, initial-scale=1