24小时热门版块排行榜    

CyRhmU.jpeg
查看: 1779  |  回复: 4
当前只显示满足指定条件的回帖,点击这里查看本话题的所有回帖

04nylxb

木虫 (正式写手)

[求助] 集群mpich2调试出问题mpdboot -n 无法启动

在集群搭建的时候,用的是mpich2-1.4.1p1,ssh nfs nis都已经OK,现在卡在mpi的调试上,一直无法启动跨节点的mpi,总是出现以下的错误,请问有解决方法不?都按照集群mpi进行配置了(mpd.hosts .mpd.hosts .mpd.conf  mpd.conf,然后都是600的权限),还是不行,是否需要重装mpich?
[root@node-1 ~]# mpdboot -n 4 -f mpd.hosts
Traceback (most recent call last):
  File "/usr/local/bin/mpdboot", line 482, in ?
    mpdboot()
  File "/usr/local/bin/mpdboot", line 234, in mpdboot
    (k,v) = kv.split('=',1)
ValueError: need more than 1 value to unpack

或者是这样的错误
[lixb@node-1 ~]$ mpdboot -n 2 -f mpd.hosts
unable to open (or read) hostsfile mpd.hosts
回复此楼

» 猜你喜欢

» 本主题相关价值贴推荐,对您同样有帮助:

集中精力发文章
已阅   回复此楼   关注TA 给TA发消息 送TA红花 TA的回帖

04nylxb

木虫 (正式写手)

引用回帖:
2楼: Originally posted by bluewhale at 2012-01-06 20:02:48:
我记得mpich2 1.4 根本不需要boot和exit daemon这二步了。我们天天在集群上运行,好像没有任何问题。Version 1.2是需要的。

你好,非常感谢啊。
我直接运行mpirun的时候出现了这样的问题:请问有遇到过吗?谢谢啊
[lixb@node-1 ~]$ mpirun -machinefile /usr/local/mpich2-1.4.1p1/bin/nodes.LINUX -np 16 ./hellocluster >out1
--------------------------------------------------------------------------
Open RTE detected a parse error in the hostfile:
    /usr/local/mpich2-1.4.1p1/bin/nodes.LINUX
It occured on line number 2 on token 5:
    node-2
--------------------------------------------------------------------------
[node-1:25170] [[64214,0],0] ORTE_ERROR_LOG: Error in file base/ras_base_allocate.c at line 236
[node-1:25170] [[64214,0],0] ORTE_ERROR_LOG: Error in file base/plm_base_launch_support.c at line 72
[node-1:25170] [[64214,0],0] ORTE_ERROR_LOG: Error in file plm_rsh_module.c at line 990
--------------------------------------------------------------------------
A daemon (pid unknown) died unexpectedly on signal 1  while attempting to
launch so we are aborting.

There may be more information reported by the environment (see above).

This may be because the daemon was unable to find all the needed shared
libraries on the remote node. You may set your LD_LIBRARY_PATH to have the
location of the shared libraries on the remote nodes and this will
automatically be forwarded to the remote nodes.
--------------------------------------------------------------------------
--------------------------------------------------------------------------
mpirun noticed that the job aborted, but has no info as to the process
that caused that situation.
--------------------------------------------------------------------------
mpirun: clean termination accomplished
集中精力发文章
3楼2012-01-06 22:42:44
已阅   回复此楼   关注TA 给TA发消息 送TA红花 TA的回帖
相关版块跳转 我要订阅楼主 04nylxb 的主题更新
信息提示
请填处理意见