当前位置：首页 > news >正文

U-Net for Image Segmentation

news 2025/7/6 11:43:21

1.Unet for Image Segmentation

笔记来源：使用Pytorch搭建U-Net网络并基于DRIVE数据集训练(语义分割)

1.1 DoubleConv (Conv2d+BatchNorm2d+ReLU)

import torch
import torch.nn as nn
import torch.nn.functional as F# nn.Sequential 按照类定义的顺序去执行模型，并且不需要写forward函数
class DoubleConv(nn.Sequential): # class 子类(父类)
# __init__() 是一个类的构造函数，用于初始化对象的属性。它会在创建对象时自动调用，而且通常在这里完成对象所需的所有初始化操作
# forward方法是实现模型的功能，实现各个层之间的连接关系的核心# def __init__(self) 只有一个self，指的是实例的本身。它允许定义一个空的类对象,需要实例化之后，再进行赋值# def __init__(self,args) 属性值不允许为空,实例化时，直接传入参数def __init__(self,in_channels,out_channels,mid_channels=None) # DoubleConv输入的通道数，输出的通道数，DoubleConv中第一次conv后输出的通道数if mid_channels is None: # mid_channels为none时执行语句mid_channels = out_channelssuper(DoubleConv,self).__init__(  #继承父类并且调用父类的初始化方法nn.Conv2d(in_channels,mid_channels,kernel_size=3,padding=1,bias=False),nn.BatchNorm2d(mid_channels), # 进行数据的归一化处理，这使得数据在进行Relu之前不会因为数据过大而导致网络性能的不稳定nn.ReLU(inplace=True), # 指原地进行操作，操作完成后覆盖原来的变量nn.Conv2d(in_channels,out_channels,kernel_size=3,padding=1,bias=False),nn.BatchNorm2d(out_channels),nn.ReLU(inplace=True))

1.1.1 Conv2d

直观了解kernel size、stride、padding、bias
原始图像通过与卷积核的数学运算，可以提取出图像的某些指定特征，卷积核相当于“滤镜”可以着重提取它感兴趣的特征

下图来自：Expression recognition based on residual rectification convolution neural network

1.1.2 BatchNorm2d

To understand what happens without normalization, let’s look at an example with just two features that are on drastically different scales. Since the network output is a linear combination of each feature vector, this means that the network learns weights for each feature that are also on different scales. Otherwise, the large feature will simply drown out the small feature.
Then during gradient descent, in order to “move the needle” for the Loss, the network would have to make a large update to one weight compared to the other weight. This can cause the gradient descent trajectory to oscillate back and forth along one dimension, thus taking more steps to reach the minimum.—Batch Norm Explained Visually — How it works, and why neural networks need it

if the features are on the same scale, the loss landscape is more uniform like a bowl. Gradient descent can then proceed smoothly down to the minimum. —Batch Norm Explained Visually — How it works, and why neural networks need it

Batch Norm is just another network layer that gets inserted between a hidden layer and the next hidden layer. Its job is to take the outputs from the first hidden layer and normalize them before passing them on as the input of the next hidden layer.

without batch norm
include batch norm

下图来自：BatchNorm2d原理、作用及其pytorch中BatchNorm2d函数的参数讲解

1.1.3 ReLU

How an Activation function works?

We know, the neural network has neurons that work in correspondence with weight, bias, and their respective activation function. In a neural network, we would update the weights and biases of the neurons on the basis of the error at the output. This process is known as Back-propagation. Activation functions make the back-propagation possible since the gradients are supplied along with the error to update the weights and biases. —Role of Activation functions in Neural Networks

Why do we need Non-linear Activation function?

Activation functions introduce non-linearity into the model, allowing it to learn and perform complex tasks. Without them, no matter how many layers we stack in the network, it would still behave as a single-layer perceptron because the composition of linear functions is a linear function. —Convolutional Neural Network — Lesson 9: Activation Functions in CNNs

Doesn’t matter how many hidden layers we attach in neural net, all layers will behave same way because the composition of two linear function is a linear function itself. Neuron cannot learn with just a linear function attached to it. A non-linear activation function will let it learn as per the difference w.r.t error. Hence, we need an activation function.—Role of Activation functions in Neural Networks

下图来自：Activation Functions 101: Sigmoid, Tanh, ReLU, Softmax and more

1.2 Down (MaxPool+DoubleConv)

class Down(nn.Sequential):def __init__(self,in_channels,out_channels):super(Down,self).__init__(nn.MaxPool2d(2,stride=2), # kernel_size核大小，stride核的移动步长DoubleConv(in_channels,out_channels))

1.2.1 MaxPool

Two reasons for applying Max Pooling :

Downscaling Image by extracting most important feature
Removing Invariances like shift, rotational and scale

下图来自：Max Pooling, Why use it and its advantages.

Max Pooling is advantageous because it adds translation invariance. There are following types of it

Shift Invariance(Invariance in Position)
Rotational Invariance(Invariance in Rotation)
Scale Invariance(Invariance in Scale(small or big))

1.3 Up (Upsample/ConvTranspose+DoubleConv)

双线性插值进行上采样

转置卷积进行上采样

class Up(nn.Module):
# __init__() 是一个类的构造函数，用于初始化对象的属性。它会在创建对象时自动调用，而且通常在这里完成对象所需的所有初始化操作def __init__(self,in_channels,out_channels,bilinear=True): #默认通过双线性插值进行上采样super(Up,self).__init__()if bilinear: #通过双线性插值进行上采样self.up = nn.Upsample(scale_factor=2,mode='bilinear',align_corners=True) #输出为输入的多少倍数、上采样算法、输入的角像素将与输出张量对齐self.conv = DoubleConv(in_channels,out_channels,in_channels//2)else #通过转置卷积进行上采样self.up = nn.ConvTranspose2d(in_channels,in_channels//2,kernel_size=2,stride=2)self.conv = DoubleConv(in_channels,out_channels) # mid_channels = none时 mid_channels = mid_channels
# forward方法是实现模型的功能，实现各个层之间的连接关系的核心def forward(self,x1,x2) # 对x1进行上采样,将上采样后的x1与x2进行cat# 对x1进行上采样x1 = self.up(x1)# 对上采样后的x1进行padding，使得上采样后的x1与x2的高宽一致# [N,C,H,W] Number of data samples、Image channels、Image height、Image widthdiff_y = x2.size()[2] - x1.size()[2] # x2的高度与上采样后的x1的高度之差diff_x = x2.size()[3] - x2.size()[3] # x2的宽度与上采样后的x1的宽度之差# padding_left、padding_right、padding_top、padding_bottomx1 = F.pad(x1,[diff_x//2, diff_x-diff_x//2, diff_y//2, diff_y-diff_y//2])# 上采样后的x1与x2进行catx = torch.cat([x2,x1],dim=1)# x1和x2进行cat之后进行doubleconvx = self.conv(x)return x

1.3.1 Upsample

上采样率scale_factor

双线性插值bilinear

角像素对齐align_corners
下图来自：[PyTorch]Upsample

1.4 OutConv

class OutConv(nn.Sequential):def __init__(self, in_channels, num_classes):super(OutConv,self).__init__(nn.Conv2d(in_channels,num_classes,kernel_size=1) #输出分类类别的数量num_classes)

1.5 Unet

class UNet(nn.Module):def __init__(self,in_channels: int = 1, # 网络输入图片的通道数num_classes: int = 2, # 网络输出分类类别数bilinear: bool = True, # 上采样时使用双线性插值base_c: int = 64): # 基础通道数，网络中其他层的通道数均为该基础通道数的倍数super(UNet,self).__init__()self.in_channels = in_channelsself.num_classes = num_classesself.bilinear = bilinearself.in_conv = DoubleConv(in_channels,base_c) # in_channels、out_channels、mid_channels=noneself.down1 = Down(base_c,base_c*2) # in_channels、out_channels self.down2 = Down(base_c*2,base_c*4)self.down3 = Down(base_c*4,base_c*8)factor = 2 if bilinear else 1 #如果上采样使用双线性插值则factor=2，若使用转置卷积则factor=1self.down4 = Down(base_c*8,base_c*16//factor)self.up1 = Up(base_c*16, base_c*8//factor, bilinear) # in_channels、out_channels、bilinear=Trueself.up2 = Up(base_c*8, base_c*4//factor, bilinear)self.up3 = Up(base_c*4, base_c*2//factor, bilinear)self.up4 = Up(base_c*2, base_c, bilinear)self.out_conv = OutConv(base_c, num_classes) # in_channels, num_classesdef forward(self,x):x1 = self.in_conv(x)x2 = self.down1(x1)x3 = self.down2(x2)x4 = self.down3(x3)x5 = self.down4(x4)x = self.up1(x5,x4) # 对x5进行上采样,将上采样后的x5与x4进行catx = self.up2(x,x3)x = self.up3(x,x2)x = self.up4(x,x1)logits = self.out_conv(x)return {"out": logits} # 以字典形式返回

为何没有调用forward方法却能直接使用？
nn.Module类是所有神经网络模块的基类，需要重载__init__和forward函数

self.up1 = Up(base_c*16, base_c*8//factor, bilinear) //实例化Up
....
x = self.up1(x5,x4) //调用 Up() 中的 forward(self,x1,x2) 函数

up1调用父类nn.Module中的__call__方法，__call__方法又调用UP()中的forward方法

U-Net for Image Segmentation

1.Unet for Image Segmentation 笔记来源：使用Pytorch搭建U-Net网络并基于DRIVE数据集训练(语义分割) 1.1 DoubleConv (Conv2dBatchNorm2dReLU) import torch import torch.nn as nn import torch.nn.functional as F# nn.Sequential 按照类定义的顺序去执行模型&…...

编程日记 2024/6/23 4:19:42

POI导入带有合并单元格的excel，demo实例，直接可以运行

直接可以运行 import org.apache.poi.hssf.usermodel.HSSFWorkbook; import org.apache.poi.ss.usermodel.Cell; import org.apache.poi.ss.usermodel.Row; import org.apache.poi.ss.usermodel.Sheet; import org.apache.poi.ss.usermodel.Workbook; import org.apache.poi.s…...

编程日记 2024/6/23 4:18:40

【C语言】解决C语言报错：Use-After-Free

文章目录简介什么是Use-After-FreeUse-After-Free的常见原因如何检测和调试Use-After-Free解决Use-After-Free的最佳实践详细实例解析示例1：释放内存后未将指针置为NULL示例2：多次释放同一指针示例3：全局或静态指针被释放后继续使用示例4&am…...

编程日记 2024/6/23 4:17:38

C语言经典例题-19

1.字符串左旋结果题目内容：写一个函数，判断一个字符串是否为另外一个字符串旋转之后的字符串。例：给定s1 AABCD和s2 BCDAA,返回1 给定s1 abcd和s2 ACBD,返回0 AABCD左旋一个字符得到ABCDA AABCD左旋两个字符得到BCDAA AABCD右旋一…...

编程日记 2024/6/23 4:14:35

AlmaLinux 更换CN镜像地址

官方镜像列表官方列表：https://mirrors.almalinux.org/CN 开头的站点，不同区域查询即可一键更改镜像地址脚本以下是更改从默认更改到阿里云地址 cat <<EOF>>/AlmaLinux_Update_repo.sh #!/bin/bash # -*- coding: utf-8 -*- # Author:…...

编程日记 2024/6/23 4:13:33

【笔记】【矩阵的二分】668. 乘法表中第k小的数

力扣链接：题目参考地址：参考思路：二分查找把矩阵想象成一维的已排好序的数组，用二分法找第k小的数字。假设m行n列，则对应一维下标范围是从1到mn，初始： l1; rmn; mid(lr)/2 设mid在第i行&a…...

编程日记 2024/6/23 4:11:31

红米手机RedNot11无法使用谷歌框架，打开游戏闪退的问题，红米手机如何开启谷歌框架

红米手机RedNot11无法使用谷歌框架，打开游戏闪退的问题， 1.问题描述2.问题原因3.解决方案3.1配置谷歌框架：3.1软件优化 4.附图 1.问题描述红米手机打开安卓APP没有广告，直接闪退，无法使用谷歌框架异常关键词中包含&…...

编程日记 2024/6/23 4:08:27

emqx5.6.1 数据、配置备份与迁移

EMQX 支持导入和导出的数据包括： EMQX 配置重写的内容： 认证与授权配置规则、连接器与 Sink/Source监听器、网关配置其他 EMQX 配置内置数据库 (Mnesia) 的数据 Dashboard 用户和 REST API 密钥客户端认证凭证（内置数据库密码认证、增强认证…...

编程日记 2024/6/23 4:07:26

VUE3脚手架工具cli配置搭建及创建VUE工程

1、VUE的脚手架工具(CLI） 开发大型vue的时候，不能通过html编写一个大型的项目，这个时候需要用到vue的脚手架工具通过vue的脚手架，可以快速的生成vue工程 1.1、安装nodejs和npm 【下载nodejs】 https://nodejs.org/en 【安装…...

编程日记 2024/6/23 4:06:24

前端开发之DNS协议

上一篇👉: 前端开发之计算机网络模型认识文章目录 DNS协议详介绍1. DNS 协议概述2. DNS协议与TCP/UDP3. DNS查询过程4. 迭代与递归查询5. DNS记录与报文结构资源记录类型对比 6. 总结 DNS协议详介绍 1. DNS 协议概述 DNS（Domain Name System&#xf…...

编程日记 2024/6/23 4:05:23

如何在 Tailwind CSS 中实现居中对齐

如何在 Tailwind CSS 中实现居中对齐： 1. 使用 text-center 类（针对行内元素或行内块元素） 这个类用于将文本或行内块元素水平居中对齐。 <div class"text-center"><span>这是一个行内元素</span> </div&g…...

编程日记 2024/6/23 4:04:21

【iOS】编译二进制文件说明

编译二进制文件说明如何生成文件路径文件说明第一部分：.o文件第二部分：link第三部分：Segment第四部分：Symbol 如何生成使用Xcode进行编译 ，会生成二进制相关文件，可以更详细看产物的布局项目Target -&…...

编程日记 2024/6/23 4:03:20

红队内网攻防渗透：内网渗透之内网对抗：隧道技术篇防火墙组策略FRPNPSChiselSocks代理端口映射C2上线

红队内网攻防渗透 1. 内网隧道技术1.1 Frp内网穿透C2上线1.1.1 双网内网穿透C2上线1.1.1.1 服务端配置1.1.1.2 客户端配置1.1.2 内网穿透信息收集1.1.2.1、建立Socks节点（入站没限制采用）1.1.2.2 主动转发数据（出站没限制采用）1.2 Nps内网穿透工具1.2.1 NPS内网穿透C2上线1…...

编程日记 2024/6/23 4:01:18

qt+halcon实战

注意建QT工程项目用的是MSVC，如果选成MinGW,则会报错 INCLUDEPATH $$PWD/include INCLUDEPATH $$PWD/include/halconcppLIBS $$PWD/lib/x64-win64/halconcpp.lib LIBS $$PWD/lib/x64-win64/halcon.lib#include "halconcpp/HalconCpp.h" #include &quo…...

编程日记 2024/6/23 4:00:17

Java_POJO

概念 POJO即简单的Java对象，区别于JavaBean JavaBean：特殊的Java类，容易被重用或插入到其他应用程序中去，通过封装属性和方法成为具有某种功能或者处理某个业务的对象这个类必须有public的无参构造器所有属性都是private的所有属…...

编程日记 2024/6/23 3:59:16

24年安克创新社招入职自适应能力cata测评真题分享北森测评高频题库

第一部分：安克创新自适应能力cata测评感谢您关注安克创新社会招聘，期待与您一起弘扬中国智造之美。为对您做出全面的评估，现诚邀您参加我们的在线测评。测评名称：社招-安克创新自适应能力cata测评第二部分：安克…...

编程日记 2024/6/23 3:57:13

OpenCV中的圆形标靶检测——findCirclesGrid()（三）

前面说到cv::findCirclesGrid2()内部先使用SimpleBlobDetector进行圆斑检测，然后使用CirclesGridClusterFinder算法类执行基于层次聚类的标靶检测。如下图所示，由于噪声的影响，SimpleBlobDetector检出的标靶可能包含噪声。而CirclesGridClusterFinder算法类会执行基…...

编程日记 2024/6/23 3:54:08

C++拷贝构造函数、运算符重载函数、赋值运算符重载函数、前置++和后置++重载等的介绍

文章目录前言一、拷贝构造函数1. 概念2. 特征3. 编译器生成默认拷贝构造函数4. 拷贝构造函数典型使用场景二、运算符重载函数三、赋值运算符重载函数1. 赋值运算符重载格式2. 赋值运算符只能重载成类的成员函数不能重载成全局函数3.编译器生成一个默认赋值运算符重载4. 运算符…...

编程日记 2024/6/23 3:52:06

视频智能分析平台智能边缘分析一体机视频监控业务平台区域人数不足检测算法

智能边缘分析一体机区域人数不足检测算法是一种集成了先进图像处理、目标检测、跟踪和计数等功能的算法，专门用于实时监测和统计指定区域内的人数，并在人数不足时发出警报。以下是对该算法的详细介绍： 一、算法概述智能边缘分析一体机区域…...

编程日记 2024/6/23 3:51:05

揭秘MMAdapt：如何利用AI跨领域战胜新兴健康谣言？

MMAdapt: A Knowledge-Guided Multi-Source Multi-Class Domain Adaptive Framework for Early Health Misinformation Detection 论文地址： MMAdapt: A Knowledge-guided Multi-source Multi-class Domain Adaptive Framework for Early Health Misinformation Detection …...

编程日记 2024/6/23 3:50:03