容器化进阶Kubernetes核心技术
容器化进阶Kubernetes核心技术
1Pod详解
Pod是Kubernetes的最重要概念,每一个Pod都有一个特殊的被称为”根容器“的Pause容器。Pause容器对应的镜 像属于Kubernetes平台的一部分,除了Pause容器,每个Pod还包含一个或多个紧密相关的使用者业务容器。

Pod vs 应用
每个Pod都是应用的一个例项,有专用的IP Pod vs 容器
一个Pod可以有多个容器,彼此间共享网络和储存资源,每个Pod 中有一个Pause容器储存所有的容器状态, 通过管理pause容器,达到管理pod中所有容器的效果
Pod vs 节点
同一个Pod中的容器总会被排程到相同Node节点,不同节点间Pod的通讯基于虚拟二层网络技术实现
Pod vs Pod
普通的Pod和静态Pod
1.1Pod的定义
下面是yaml档案定义的Pod的完整内容
apiVersion: v1
kind: Pod
metadata: //元资料
name: string
namespace: string
labels:
-name: string
annotations:
-name: string
spec:
containers: //pod中的容器列表,可以有多个容器
- name: string //容器的名称
image: string //容器中的映象
imagesPullPolicy: [Always|Never|IfNotPresent]//获取映象的策略,预设值为Always,每次都尝试重新下载映象
command: [string] //容器的启动命令列表(不配置的话使用映象内部的命令) args: [string] //启动引数列表
workingDir: string //容器的工作目录volumeMounts: //挂载到到容器内部的储存卷设定
-name: string
mountPath: string //储存卷在容器内部Mount的绝对路径readOnly: boolean //预设值为读写
ports: //容器需要暴露的埠号列表
-name: string
containerPort: int //容器要暴露的埠
hostPort: int //容器所在主机监听的埠(容器暴露埠对映到宿主机的埠,设定hostPort时同一 台宿主机将不能再启动该容器的第2份副本)
protocol: string //TCP和UDP,预设值为TCP env: //容器执行前要设定的环境列表
-name: string value: string
resources:
limits: //资源限制,容器的最大可用资源数量cpu: Srting
memory: string
requeste: //资源限制,容器启动的初始可用资源数量cpu: string
memory: string
livenessProbe: //pod内容器健康检查的设定exec:
command: [string] //exec方式需要指定的命令或指令码httpGet: //通过httpget检查健康
path: string port: number host: string scheme: Srtring httpHeaders:
- name: Stirng value: string
tcpSocket: //通过tcpSocket检查健康
port: number initialDelaySeconds: 0//首次检查时间timeoutSeconds: 0 //检查超时时间
periodSeconds: 0 //检查间隔时间
successThreshold: 0
failureThreshold: 0 securityContext: //安全配置
privileged: falae
restartPolicy: [Always|Never|OnFailure]//重启策略,预设值为Always
nodeSelector: object //节点选择,表示将该Pod排程到包含这些label的Node上,以key:value格式指定
imagePullSecrets:
-name: string
hostNetwork: false //是否使用主机网络模式,弃用Docker网桥,预设否
volumes: //在该pod上定义共享储存卷列表
-name: string emptyDir: {} hostPath:
path: string secret:
secretName: string item:
-key: string path: string
configMap: name: string items:
-key: string
path: string
1.2Pod的基本用法
在kubernetes中对执行容器的要求为:容器的主程式需要一直在前台执行,而不是后台执行。应用需要改造成前 台执行的方式。如果我们建立的Docker映象的启动命令是后台执行程式,则在kubelet建立包含这个容器的pod之 后执行完该命令,即认为Pod已经结束,将立刻销毁该Pod。如果为该Pod定义了RC,则建立、销毁会陷入一个无 限循环的过程中。
Pod可以由1个或多个容器组合而成。由一个容器组成的Pod示例
# 一个容器组成的Pod apiVersion: v1 kind: Pod metadata:
name: mytomcat labels:
name: mytomcat spec:
containers:
- name: mytomcat image: tomcat ports:
- containerPort: 8000
由两个为紧耦合的容器组成的Pod示例
#两个紧密耦合的容器
apiVersion: v1 kind: Pod metadata:
name: myweb labels:
name: tomcat-redis
spec:
containers:
-name: tomcat image: tomcat ports:
-containerPort: 8080
-name: redis image: redis ports:
-containerPort: 6379
建立
kubectl create -f xxx.yaml
检视
kubectl get pod/po
kubectl get pod/po -o wide
kubectl describe pod/po
删除
kubectl delete -f pod pod_name.yaml
kubectl delete pod --all/[pod_name]
1.3Pod的分类
Pod有两种型别
普通Pod
普通Pod一旦被建立,就会被放入到etcd中储存,随后会被Kubernetes Master排程到某个具体的Node上并进行系结,随后该Pod对应的Node上的kubelet程序例项化成一组相关的Docker容器并启动起来。在预设情 况下,当Pod里某个容器停止时,Kubernetes会自动检测到这个问题并且重新启动这个Pod里某所有容器, 如果Pod所在的Node宕机,则会将这个Node上的所有Pod重新排程到其它节点上。
静态Pod
静态Pod是由kubelet进行管理的仅存在于特定Node上的Pod,它们不能通过 API Server进行管理,无法与ReplicationController、Deployment或DaemonSet进行关联,并且kubelet也无法对它们进行健康检查。
1.4Pod生命周期和重启策略
Pod的状态

Pod重启策略
Pod的重启策略包括Always、OnFailure和Never,预设值是Always

常见状态转换

1.5Pod资源配置
每个Pod都可以对其能使用的服务器上的计算资源设定限额,Kubernetes中可以设定限额的计算资源有CPU与Memory两种,其中CPU的资源单位为CPU数量,是一个绝对值而非相对值。Memory配额也是一个绝对值,它的单 位是内存字节数。
Kubernetes里,一个计算资源进行配额限定需要设定以下两个引数: Requests 该资源最小申请数量,系统必须满足要求
Limits 该资源最大允许使用的量,不能突破,当容器试图使用超过这个量的资源时,可能会被Kubernetes Kill并重启
sepc
containers:
- name: db
image: mysql
resources:
requests:
memory: "64Mi"
cpu: "250m"
limits:
memory: "128Mi"
cpu: "500m"
上述程式码表明MySQL容器申请最少0.25个CPU以及64MiB内存,在执行过程中容器所能使用的资源配额为0.5个
CPU以及128MiB内存。
2Label详解
Label是Kubernetes系统中另一个核心概念。一个Label是一个key=value的键值对,其中key与value由使用者自己指 定。Label可以附加到各种资源物件上,如Node、Pod、Service、RC,一个资源物件可以定义任意数量的Label, 同一个Label也可以被新增到任意数量的资源物件上,Label通常在资源物件定义时确定,也可以在物件建立后动态 新增或删除。Label的最常见的用法是使用metadata.labels字段,来为物件新增Label,通过spec.selector来引用物件
apiVersion: v1
kind: ReplicationController metadata:
name: nginx spec:
replicas: 3 selector:
app: nginx template:
metadata:
labels:
app: nginx spec:
containers:
- name: nginx image: nginx ports:
- containerPort: 80
-------------------------------------
apiVersion: v1 kind: Service metadata: name: nginx
spec:
type: NodePort ports:
- port: 80
nodePort: 3333 selector:
app: nginx
Label附加到Kubernetes丛集中的各种资源物件上,目的就是对这些资源物件进行分组管理,而分组管理的核心就 是Label Selector。Label与Label Selector都是不能单独定义,必须附加在一些资源物件的定义档案上,一般附加 在RC和Service的资源定义档案中。
3Replication Controller详解
Replication Controller(RC)是Kubernetes系统中核心概念之一,当我们定义了一个RC并提交到Kubernetes丛集中以后,Master节点上的Controller Manager元件就得到通知,定期检查系统中存活的Pod,并确保目标Pod例项的数量刚好等于RC的预期值,如果有过多或过少的Pod执行,系统就会停掉或建立一些Pod.此外我们也可以通过修改 RC的副本数量,来实现Pod的动态缩放功能。kubectl scale rc nginx --replicas=5
由于Replication Controller与Kubernetes程式码中的模组Replication Controller同名,所以在Kubernetes v1.2时, 它就升级成了另外一个新的概念Replica Sets,官方解释为下一代的RC,它与RC区别是:Replica Sets支援基于集合的Label selector,而RC只支援基于等式的Label Selector。我们很少单独使用Replica Set,它主要被Deployment这个更高层面的资源物件所使用,从而形成一整套Pod建立、删除、更新的编排机制。最好不要越过RC直接建立Pod, 因为Replication Controller会通过RC管理Pod副本,实现自动建立、补足、替换、删除Pod副本,这样就能提高应用的容灾能力,减少由于节点崩溃等意外状况造成的损失。即使应用程序只有一个Pod副本,也强烈建议使用RC来 定义Pod
4Replica Set详解
ReplicaSet 跟 ReplicationController 没有本质的不同,只是名字不一样,并且 ReplicaSet 支援集合式的selector(ReplicationController 仅支援等式)。Kubernetes官方强烈建议避免直接使用ReplicaSet,而应该通过Deployment来建立RS和Pod。由于ReplicaSet是ReplicationController的代替物,因此用法基本相同,唯一的区别在于ReplicaSet支援集合式的selector。5Deployment详解
Deployment是Kubenetes v1.2引入的新概念,引入的目的是为了更好的解决Pod的编排问题,Deployment内部使用了Replica Set来实现。Deployment的定义与Replica Set的定义很类似,除了API宣告与Kind型别有所区别:apiVersion: extensions/v1beta1 kind: Deployment
metadata:
name: frontend spec:
replicas: 1 selector:
matchLabels:
tier: frontend matchExpressions:
- {key: tier, operator: In, values: [frontend]} template:
metadata:
labels:
app: app-demo tier: frontend
spec:
containers:
- name: tomcat-demo image: tomcat ports:
- containerPort: 8080
6Horizontal Pod Autoscaler
Horizontal Pod Autoscal(Pod横向扩容 简称HPA)与RC、Deployment一样,也属于一种Kubernetes资源物件。通过追踪分析RC控制的所有目标Pod的负载变化情况,来确定是否需要针对性地调整目标Pod的副本数,这是HPA的 实现原理。Kubernetes对Pod扩容与缩容提供了手动和自动两种模式,手动模式通过kubectl scale命令对一个Deployment/RC进行Pod副本数量的设定。自动模式则需要使用者根据某个效能指标或者自定义业务指标,并指定Pod副本数量的范围,系统将自动在这个范围内根据效能指标的变化进行调整。
手动扩容和缩容
kubectl scale deployment frontend --replicas 1
自动扩容和缩容
HPA控制器基本Master的kube-controller-manager服务启动引数 --horizontal-pod-autoscaler-sync-period 定义的时长(预设值为30s),周期性地监测Pod的CPU使用率,并在满足条件时对RC或Deployment中的Pod副 本数量进行调整,以符合使用者定义的平均Pod CPU使用率。
apiVersion: extensions/v1beta1 kind: Deployment
metadata:
name: nginx-deployment spec:
replicas: 1 template:
metadata: name: nginx labels:
app: nginx spec:
containers:
- name: nginx image: nginx
resources:
requests:
cpu: 50m ports:
- containerPort: 80
-------------------------------
apiVersion: v1 kind: Service metadata:
name: nginx-svc spec:
ports:
- port: 80 selector:
app: nginx
-----------------------------------
apiVersion: autoscaling/v1 kind: HorizontalPodAutoscaler metadata:
name: nginx-hpa spec:
scaleTargetRef:
apiVersion: app/v1beta1 kind: Deployment
name: nginx-deployment minReplicas: 1
maxReplicas: 10
targetCPUUtilizationPercentage: 50
7Volume详解
Volume是Pod中能够被多个容器访问的共享目录。Kubernetes的Volume定义在Pod上,它被一个Pod中的多个容 器挂载到具体的档案目录下。Volume与Pod的生命周期相同,但与容器的生命周期不相关,当容器终止或重启时,Volume中的资料也不会丢失。要使用volume,pod需要指定volume的型别和内容(字段),
和对映到容器的位置(字段)。 Kubernetes支援多种型别的Volume,包括:
emptyDir、hostPath、gcePersistentDisk、awsElasticBlockStore、nfs、iscsi、flocker、glusterfs、rbd、cephfs、gitRepo、secret、persistentVolumeClaim、downwardAPI、azureFileVolume、azureDisk、vsphereVolume、Quobyte、PortworxVolume、ScaleIO。
emptyDir
EmptyDir型别的volume创建于pod被排程到某个宿主机上的时候,而同一个pod内的容器都能读写EmptyDir 中的同一个档案。一旦这个pod离开了这个宿主机,EmptyDir中的资料就会被永久删除。所以目前EmptyDir 型别的volume主要用作临时空间,比如Web服务器写日志或者tmp档案需要的临时目录。yaml示例如下
apiVersion: v1 kind: Pod metadata:
name: test-pd spec:
containers:
- image: docker.io/nazarpc/webserver
name: test-container
volumeMounts:
- mountPath: /cache name: cache-volume
volumes:
- name: cache-volume emptyDir: {}
hostPath
HostPath属性的volume使得对应的容器能够访问当前宿主机上的指定目录。例如,需要执行一个访问Docker系统目录的容器,那么就使用/var/lib/docker目录作为一个HostDir型别的volume;或者要在一个容器内部执行CAdvisor,那么就使用/dev/cgroups目录作为一个HostDir型别的volume。一旦这个pod离开了这个宿主机,HostDir中的资料虽然不会被永久删除,但资料也不会随pod迁移到其他宿主机上。因此,需要 注意的是,由于各个宿主机上的档案系统结构和内容并不一定完全相同,所以相同pod的HostDir可能会在不 同的宿主机上表现出不同的行为。yaml示例如下:
apiVersion: v1 kind: Pod metadata:
name: test-pd spec:
containers:
-image: docker.io/nazarpc/webserver name: test-container
# 指定在容器中挂接路径
volumeMounts:
- mountPath: /test-pd name: test-volume
# 指定所提供的储存卷
volumes:
-name: test-volume # 宿主机上的目录hostPath:
# directory location on host path: /data
nfs
NFS型别的volume。允许一块现有的网络硬盘在同一个pod内的容器间共享。yaml示例如下:
apiVersion: apps/v1 # for versions before 1.9.0 use apps/v1beta2 kind: Deployment
metadata:
name: redis spec:
selector: matchLabels:
app: redis revisionHistoryLimit: 2 template:
metadata:
labels:
app: redis spec:
containers:
# 应用的映象
-image: redis name: redis
imagePullPolicy: IfNotPresent # 应用的内部埠
ports:
-containerPort: 6379 name: redis6379
env:
-name: ALLOW_EMPTY_PASSWORD
value: "yes"
-name: REDIS_PASSWORD
value: "redis"
# 持久化挂接位置,在docker中
volumeMounts:
-name: redis-persistent-storage mountPath: /data
volumes:
# 宿主机上的目录
-name: redis-persistent-storage nfs:
path: /k8s-nfs/redis/data server: 192.168.126.112
8. Namespace详解
Namespace在很多情况下用于实现多使用者的资源隔离,通过将丛集内部的资源物件分配到不同的Namespace中, 形成逻辑上的分组,便于不同的分组在共享使用整个丛集的资源同时还能被分别管理。Kubernetes丛集在启动后,会建立一个名为"default"的Namespace,如果不特别指明Namespace,则使用者建立的Pod,RC,Service都将 被系统 建立到这个预设的名为default的Namespace中。
Namespace建立
apiVersion: v1 kind: Namespace metadata:
name: development
---------------------
apiVersion: v1 kind: Pod metadata:
name: busybox namespace: development
spec:
containers:
- image: busybox command:
- sleep
- -"3600"
name: busybox
Namespace检视
kubectl get pods --namespace=development
9Service 详解
Service是Kubernetes最核心概念,通过建立Service,可以为一组具有相同功能的容器应用提供一个统一的入口地 址,并且将请求负载分发到后端的各个容器应用上。9.1Service的定义
yaml格式的Service定义档案
apiVersion: v1 kind: Service matadata:
name: string namespace: string labels:
-name: string annotations:
-name: string spec:
selector: [] type: string clusterIP: string
sessionAffinity: string ports:
-name: string protocol: string port: int targetPort: int nodePort: int
status: loadBalancer:
ingress:
ip: string hostname: string


9.2Service的基本用法
一般来说,对外提供服务的应用程序需要通过某种机制来实现,对于容器应用最简便的方式就是通过TCP/IP机制及 监听IP和埠号来实现。建立一个基本功能的Service
apiVersion: v1
kind: ReplicationController metadata:
name: mywebapp spec:
replicas: 2 template:
metadata:
name: mywebapp labels:
app: mywebapp spec:
containers:
-name: mywebapp image: tomcat ports:
-containerPort: 8080
我们可以通过kubectl get pods -l app=mywebapp -o yaml | grep podIP来获取Pod的IP地址和埠号来访问Tomcat服务,但是直接通过Pod的IP地址和埠访问应用服务是不可靠的,因为当Pod所在的Node发生故障时, Pod将被kubernetes重新排程到另一台Node,Pod的地址会发生改变。我们可以通过配置档案来定义Service,再 通过kubectl create来建立,这样可以通过Service地址来访问后端的Pod.
apiVersion: v1 kind: Service metadata:
name: mywebAppService spec:
ports:
- port: 8081
targetPort: 8080 selector:
app: mywebapp
9.2.1多埠Service
有时一个容器应用也可能需要提供多个埠的服务,那么在Service的定义中也可以相应地设定为将多个埠对应 到多个应用服务。
apiVersion: v1 kind: Service metadata:
name: mywebAppService spec:
ports:
- port: 8080
targetPort: 8080 name: web
- port: 8005
targetPort: 8005 name: management
selector:
app: mywebapp
9.2.2外部服务Service
在某些特殊环境中,应用系统需要将一个外部数据库作为后端服务进行连线,或将另一个丛集或Namespace中的 服务作为服务的后端,这时可以通过建立一个无Label Selector的Service来实现。
apiVersion: v1 kind: Service metadata:
name: my-service spec:
ports:
- protocol: TCP port: 80
targetPort: 80
--------------------------
apiVersion: v1
kind: Endpoints metadata:
name: my-service subsets:
- addresses:
- IP: 10.254.74.3
ports:
- port: 8080