Pytorch--模型加载

最新推荐文章于 2024-06-21 14:35:06 发布

原创最新推荐文章于 2024-06-21 14:35:06 发布 · 1.5k 阅读

1 ·

本内容遵循CC 4.0 BY-SA版权协议

收录于

函数

Pytorch

本文详细介绍了PyTorch中模型加载与保存的相关操作，包括使用`load_state_dict()`加载模型参数，利用`nn.DataParallel`进行模型并行化，以及如何保存和加载整个模型状态。同时，还展示了如何只加载特定层的参数，以及设置模型层参数的训练状态。通过实例代码演示了整个流程。

一、 Pytorch–模型加载

1、load_state_dict()函数使用

model = VGG()# 实例化自己的模型；
checkpoint = torch.load('checkpoint.pt', map_location='cpu')  # 加载模型文件，pt, pth 文件都可以；
if torch.cuda.device_count() > 1:
    # 如果有多个GPU，将模型并行化，用DataParallel来操作。这个过程会将key值加一个"module. ***"。
    model = nn.DataParallel(model) 
model.load_state_dict(checkpoint) # 接着就可以将模型参数load进模型。

如果有多个GPU，将模型并行化，用DataParallel来操作。
2、state_dict()一个简单的python的字典对象,将每一层与它的对应参数建立映射关系，
是在定义了model或optimizer之后pytorch自动生成的,可以直接调用.常用的保存state_dict的格式是".pt"或’.pth’的文件,即下面命令的 PATH="./***.pt"

torch.save(model.state_dict(), PATH)

注：只有那些参数可以训练的layer才会被保存到模型的state_dict中,如卷积层,线性层等等
3、仅保存学习到的参数，并加载

torch.save(model.state_dict(), PATH)

model = TheModelClass(*args, **kwargs)
model.load_state_dict(torch.load(PATH))

for param_tensor in model.state_dict():
    print(param_tensor,'\t',model.state_dict()[param_tensor].size())
    
model.eval()

4、保存整个model的状态，与加载

torch.save(model,PATH)

model = torch.load(PATH)
model.eval()

5、仅加载某一层的训练的到的参数，设置某层某参数的"是否需要训练"(param.requires_grad)

conv1_weight_state = torch.load('./model_state_dict.pt')['conv1.weight']

for param in list(model.pretrained.parameters()):
 param.requires_grad = False

6、全部代码

#-*-coding:utf-8-*-
import torch
import torch.nn as nn
import torch.nn.functional as F
import torch.optim as optim
 
 
 
# define model
class TheModelClass(nn.Module):
    def __init__(self):
        super(TheModelClass,self).__init__()
        self.conv1 = nn.Conv2d(3,6,5)
        self.pool = nn.MaxPool2d(2,2)
        self.conv2 = nn.Conv2d(6,16,5)
        self.fc1 = nn.Linear(16*5*5,120)
        self.fc2 = nn.Linear(120,84)
        self.fc3 = nn.Linear(84,10)
 
    def forward(self,x):
        x = self.pool(F.relu(self.conv1(x)))
        x = self.pool(F.relu(self.conv2(x)))
        x = x.view(-1,16*5*5)
        x = F.relu(self.fc1(x))
        x = F.relu(self.fc2(x))
        x = self.fc3(x)
        return x
 
# initial model
model = TheModelClass()
 
#initialize the optimizer
optimizer = optim.SGD(model.parameters(),lr=0.001,momentum=0.9)
 
# print the model's state_dict
print("model's state_dict:")
for param_tensor in model.state_dict():
    print(param_tensor,'\t',model.state_dict()[param_tensor].size())
 
print("\noptimizer's state_dict")
for var_name in optimizer.state_dict():
    print(var_name,'\t',optimizer.state_dict()[var_name])
 
print("\nprint particular param")
print('\n',model.conv1.weight.size())
print('\n',model.conv1.weight)
 
print("------------------------------------")
torch.save(model.state_dict(),'./model_state_dict.pt')
# model_2 = TheModelClass()
# model_2.load_state_dict(torch.load('./model_state_dict'))
# model.eval()
# print('\n',model_2.conv1.weight)
# print((model_2.conv1.weight == model.conv1.weight).size())
## 仅仅加载某一层的参数
conv1_weight_state = torch.load('./model_state_dict.pt')['conv1.weight']
print(conv1_weight_state==model.conv1.weight)
 
model_2 = TheModelClass()
model_2.load_state_dict(torch.load('./model_state_dict.pt'))
model_2.conv1.requires_grad=False
print(model_2.conv1.requires_grad)
print(model_2.conv1.bias.requires_grad)``

>https://blog.csdn.net/Strive_For_Future/article/details/83240081

标签

#python