torque-6.1.2 安装问题,子节点down状态如何启动
torque-6.1.2 安装问题,节点down状态如何启动
qterm -t quick
pbs_server
pbsnodes -a
发现子节点是 state = down
已关防火墙,配置正确,可ssh切换,节点服务都启动,还是出问题
主节点:
[root@calserver calserver]# for i in pbs_server pbs_sched pbs_mom trqauthd; do service $i start; done
Starting pbs_server (via systemctl): [ OK ]
Starting pbs_sched (via systemctl): [ OK ]
Starting pbs_mom (via systemctl): [ OK ]
Starting trqauthd (via systemctl): [ OK ]
[root@calserver calserver]# ps -ef | grep pbs
root 1160 1 0 01:18 ? 00:00:00 /usr/local/torque/sbin/pbs_server -F -d /var/spool/torque
root 3566 1 0 01:20 ? 00:00:00 /usr/local/torque/sbin/pbs_sched -d /var/spool/torque
root 3593 1 0 01:20 ? 00:00:00 /usr/local/torque/sbin/pbs_mom -F -d /var/spool/torque
root 3659 3428 0 01:21 pts/0 00:00:00 grep --color=auto pbs
[root@calserver calserver]# qnodes
calserver
state = free
power_state = Running
np = 16
ntype = cluster
status = opsys=linux,uname=Linux calserver 3.10.0-862.14.4.el7.x86_64 #1 SMP Wed Sep 26 15:12:11 UTC 2018 x86_64,sessions=1593 2113 2237 2247 2501 3135 3185 3240,nsessions=8,nusers=2,idletime=256,totmem=5960692kb,availmem=4875732kb,physmem=3863544kb,ncpus=16,loadave=0.18,gres=,netload=89393,state=free,varattr= ,cpuclock=Fixed,macaddr=00:0c:29:a0:9b:d2,version=6.1.2,rectime=1540660913,jobs=
mom_service_port = 15002
mom_manager_port = 15003
calnode02
state = down
power_state = Running
np = 4
ntype = cluster
mom_service_port = 15002
mom_manager_port = 15003
calnode03
state = down
power_state = Running
np = 12
ntype = cluster
mom_service_port = 15002
mom_manager_port = 15003
计算节点:
[root@calnode02 ~]# systemctl status pbs_mom.service -l
● pbs_mom.service - TORQUE pbs_mom daemon
Loaded: loaded (/usr/lib/systemd/system/pbs_mom.service; enabled; vendor preset: disabled)
Active: active (running) since Sun 2018-10-28 01:18:50 CST; 10min ago
Main PID: 1041 (pbs_mom)
Tasks: 11
Memory: 101.8M
CGroup: /system.slice/pbs_mom.service
└─1041 /usr/local/torque/sbin/pbs_mom -F -d /var/spool/torque
Oct 28 01:29:05 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Could not contact any of the servers to send an update
Oct 28 01:29:05 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Status not successfully updated for 154 MOM status update intervals
Oct 28 01:29:09 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Could not contact any of the servers to send an update
Oct 28 01:29:09 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Status not successfully updated for 155 MOM status update intervals
Oct 28 01:29:14 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Could not contact any of the servers to send an update
Oct 28 01:29:14 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Status not successfully updated for 156 MOM status update intervals
Oct 28 01:29:18 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Could not contact any of the servers to send an update
Oct 28 01:29:18 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Status not successfully updated for 157 MOM status update intervals
Oct 28 01:29:22 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Could not contact any of the servers to send an update
Oct 28 01:29:22 calnode02 pbs_mom[1041]: LOG_ERROR::send_update_to_a_server, Status not successfully updated for 158 MOM status update intervals
参考安装方法
MS7、Torque在CentOS6.5上的安装-即MS计算集群搭建(原创) - 第一性原理 - MS - 小木虫论坛-学术科研互动平台 http://muchong.com/t-9836836-1-authorid-1192095
Centos7安装-多节点Torque - u012460749的博客 - CSDN博客 https://blog.csdn.net/u012460749/article/details/78583026
上面小木虫的安装方法里面
nfs分享ms 的目录为什么提示找不到
将 Accelrys 目录共享给其他计算节点:
# echo ‘/home/xxx/Accelrys *(rw,no_root_squash)’ >> /etc/exports
重启 nfs 服务:
$ sudo service nfs restart (centos7 : systemctl restart nfs.service)
b) 计算节点配置
创建共享文件夹 Accelrys,并挂载服务节点共享的 Accelrys:
$ cd
$ mkdir Accelrys
$ sudo mount –t nfs calserver:/home/xxx/Accelrys/ /home/xxx/Accelrys/
这一步找不到地址
配置开机自动挂载 Accelrys:
# echo ‘mount –t nfs calserver:/home/xxx/Accelrys/ /home/xxx/Accelrys/’ >> /etc/rc.d/rc.local 返回小木虫查看更多
主节点上pbs_server的log
[root@calserver calserver]# systemctl status pbs_server.service -l
● pbs_server.service - TORQUE pbs_server daemon
Loaded: loaded (/usr/lib/systemd/system/pbs_server.service; enabled; vendor preset: disabled)
Active: active (running) since Sun 2018-10-28 01:18:08 CST; 35min ago
Main PID: 1160 (pbs_server)
Tasks: 12
Memory: 1.6M
CGroup: /system.slice/pbs_server.service
└─1160 /usr/local/torque/sbin/pbs_server -F -d /var/spool/torque
Oct 28 01:18:08 calserver systemd[1]: Starting TORQUE pbs_server daemon...
Oct 28 01:18:08 calserver PBS_Server[1160]: LOG_ERROR::tcp_connect_sockaddr, Failed when trying to open tcp connection - connect() failed [rc = -2] [addr = 127.0.0.1:15003]
Oct 28 01:18:08 calserver PBS_Server[1160]: LOG_ERROR::sendHierarchyToNode, Could not send mom hierarchy to host calserver:15003
Oct 28 01:18:08 calserver PBS_Server[1160]: LOG_ERROR::tcp_connect_sockaddr, Failed when trying to open tcp connection - connect() failed [rc = 15096] [addr = 192.168.10.102:15003]
Oct 28 01:18:08 calserver PBS_Server[1160]: LOG_ERROR::sendHierarchyToNode, Could not send mom hierarchy to host calnode02:15003
Oct 28 01:18:08 calserver PBS_Server[1160]: LOG_ERROR::tcp_connect_sockaddr, Failed when trying to open tcp connection - connect() failed [rc = 15096] [addr = 192.168.10.103:15003]
Oct 28 01:18:08 calserver PBS_Server[1160]: LOG_ERROR::sendHierarchyToNode, Could not send mom hierarchy to host calnode03:15003
Oct 28 01:28:09 calserver pbs_server[1160]: Assertion failed, bad pointer in link: file "req_select.c", line 401
Oct 28 01:38:09 calserver pbs_server[1160]: Assertion failed, bad pointer in link: file "req_select.c", line 401
Oct 28 01:48:09 calserver pbs_server[1160]: Assertion failed, bad pointer in link: file "req_select.c", line 401,
意思是共享calserver节点的ms目录那一步没完成?那肯定是不行的,问题就处在这里了!为什么会找不到地址?检查下地址有没有写错呀,那个地址就是就是ms的安装地址,写对了应该能找到的呀
不是 ms布置都没问题,是torque服务节点和计算节点安装完成后状态是down,不是free
ms里面torque设置没问题,torque调用不了计算节点
楼主解决了么,我的是 4.2.10的torque,就直接一个节点挂了所有的核,然后state一直是down,重启,关防火墙,qterm一系列的都试过了怎么都开不了,系统是centos7
提供几个思路:
1、防火墙不能关闭,需要彻底卸载;
2、所有服务节点、计算节点节点间配置无密互访;
3、路由器中绑定死所有节点的MAC和ip地址,并在hosts文件中写死;
4、服务节点的几个配置文件写死;
5、计算节点的几个配置文件写死;
6、qmgr建立恰当的默认计算队列。
检查以上几点,一般不会出问题。经验之谈。