레이블이 ppc64le인 게시물을 표시합니다. 모든 게시물 표시
레이블이 ppc64le인 게시물을 표시합니다. 모든 게시물 표시

2022년 12월 7일 수요일

Microsoft .NET 7, 이제 IBM Power에서도 지원

 


.NET 7이 이제 Linux on Power (ppc64le)에서도 지원됩니다.  즉 ppc64le 아키텍처의 Power 서버에서는 C# 언어도 지원하게 되었습니다.  이로써 microservices app 개발에 사용되는 5개 프로그래밍 언어가 모두 ppc64le에서 지원됩니다.  





즉, 과거에는 Oracle 등 AIX에서 수행되는 기간계 시스템과 연계하여 C# 언어로 작성된 업무 프로그램을 사용하기 위해서는 Windows 또는 Linux를 설치한 x86 시스템이 필요했으나, 이제는 x86이 필요하지 않고 같은 Power 시스템에 설치된 Redhat 또는 Openshift container에서 C# 업무 프로그램을 운용할 수 있게 된 것입니다.  





2021년 Stack Overflow에서의 개발자 설문에 따르면 .NET framework은 전업 개발자들에게 있어서 가장 중요한 프레임웍 및 라이브러리입니다. 







관련 발표는 아래 URL에서 좀 더 자세히 보실 수 있습니다.


https://community.ibm.com/community/user/powerdeveloper/blogs/janani-janakiraman/2022/11/07/net7-support-linux-on-power


또한 .NET 7 ppc64le는 이미 고객 reference가 있습니다.  독일의 컴퓨터 설계 관련 기업인 SKM Informatik GmbH인데, 이 기업의 요청으로 .NET 7의 ppc64le 지원이 시작되었다고 합니다.


https://www.ibm.com/downloads/cas/29RYARBY


.NET 7은 github에 그 source code가 공개되니 직접 build 하셔도 됩니다만, .NET 7 for ppc64le는 Openshift container 이미지로 제공되고 또 RHEL 8.7 또는 RHEL 9.1에서 YUM repository를 통해서도 제공됩니다.  


.NET 7 관련된 패키지들은 yum 명령어를 통해 패키지로 아래 예와 같이 설치하실 수 있습니다.  


[cecuser@p1274-pvm1 ~]$ hostnamectl

   Static hostname: p1274-pvm1

         Icon name: computer-vm

           Chassis: vm

        Machine ID: e5e70adcd98c4b56b00cc49cbf093412

           Boot ID: 36e5448112b5432c91f40abe413cb05a

    Virtualization: powervm

  Operating System: Red Hat Enterprise Linux 8.7 (Ootpa)

       CPE OS Name: cpe:/o:redhat:enterprise_linux:8::baseos

            Kernel: Linux 4.18.0-372.9.1.el8.ppc64le

      Architecture: ppc64-le



[cecuser@p1274-pvm1 ~]$ yum list | grep dotnet

dotnet.ppc64le                                          7.0.100-1.el8_7                                             rhel-8-for-ppc64le-appstream-rpms

dotnet-apphost-pack-7.0.ppc64le                         7.0.0-1.el8_7                                               rhel-8-for-ppc64le-appstream-rpms

dotnet-host.ppc64le                                     7.0.0-1.el8_7                                               rhel-8-for-ppc64le-appstream-rpms

dotnet-hostfxr-7.0.ppc64le                              7.0.0-1.el8_7                                               rhel-8-for-ppc64le-appstream-rpms

dotnet-runtime-7.0.ppc64le                              7.0.0-1.el8_7                                               rhel-8-for-ppc64le-appstream-rpms

dotnet-sdk-7.0.ppc64le                                  7.0.100-1.el8_7                                             rhel-8-for-ppc64le-appstream-rpms

dotnet-sdk-7.0-source-built-artifacts.ppc64le           7.0.100-1.el8_7                                             codeready-builder-for-rhel-8-ppc64le-rpms

dotnet-targeting-pack-7.0.ppc64le                       7.0.0-1.el8_7                                               rhel-8-for-ppc64le-appstream-rpms

dotnet-templates-7.0.ppc64le                            7.0.100-1.el8_7                                             rhel-8-for-ppc64le-appstream-rpms




[cecuser@p1274-pvm1 ~]$ sudo yum install dotnet-sdk-7.0

Updating Subscription Management repositories.

Last metadata expiration check: 23:10:11 ago on Tue 06 Dec 2022 03:10:01 AM EST.

Dependencies resolved.

===================================================================================================

 Package                        Arch    Version            Repository                         Size

===================================================================================================

Installing:

 dotnet-sdk-7.0                 ppc64le 7.0.100-1.el8_7    rhel-8-for-ppc64le-appstream-rpms  53 M

Installing dependencies:

 aspnetcore-runtime-7.0         ppc64le 7.0.0-1.el8_7      rhel-8-for-ppc64le-appstream-rpms 2.8 M

 aspnetcore-targeting-pack-7.0  ppc64le 7.0.0-1.el8_7      rhel-8-for-ppc64le-appstream-rpms 1.6 M

 dotnet-apphost-pack-7.0        ppc64le 7.0.0-1.el8_7      rhel-8-for-ppc64le-appstream-rpms 100 k

 dotnet-host                    ppc64le 7.0.0-1.el8_7      rhel-8-for-ppc64le-appstream-rpms 190 k

 dotnet-hostfxr-7.0             ppc64le 7.0.0-1.el8_7      rhel-8-for-ppc64le-appstream-rpms 175 k

 dotnet-runtime-7.0             ppc64le 7.0.0-1.el8_7      rhel-8-for-ppc64le-appstream-rpms 7.8 M

 dotnet-targeting-pack-7.0      ppc64le 7.0.0-1.el8_7      rhel-8-for-ppc64le-appstream-rpms 2.9 M

 dotnet-templates-7.0           ppc64le 7.0.100-1.el8_7    rhel-8-for-ppc64le-appstream-rpms 3.1 M

 netstandard-targeting-pack-2.1 ppc64le 7.0.100-1.el8_7    rhel-8-for-ppc64le-appstream-rpms 1.5 M


Transaction Summary

===================================================================================================

Install  10 Packages


Total download size: 73 M

Installed size: 311 M

Is this ok [y/N]: y


...


Installed:

  aspnetcore-runtime-7.0-7.0.0-1.el8_7.ppc64le

  aspnetcore-targeting-pack-7.0-7.0.0-1.el8_7.ppc64le

  dotnet-apphost-pack-7.0-7.0.0-1.el8_7.ppc64le

  dotnet-host-7.0.0-1.el8_7.ppc64le

  dotnet-hostfxr-7.0-7.0.0-1.el8_7.ppc64le

  dotnet-runtime-7.0-7.0.0-1.el8_7.ppc64le

  dotnet-sdk-7.0-7.0.100-1.el8_7.ppc64le

  dotnet-targeting-pack-7.0-7.0.0-1.el8_7.ppc64le

  dotnet-templates-7.0-7.0.100-1.el8_7.ppc64le

  netstandard-targeting-pack-2.1-7.0.100-1.el8_7.ppc64le


Complete!


설치는 위와 같이 매우 간단합니다.  이제 "Hello, World!"를 프린트하는 샘플 프로그램을 돌려보도록 하겠습니다.


[cecuser@p1274-pvm1 ~]$ dotnet


Usage: dotnet [options]

Usage: dotnet [path-to-application]


Options:

  -h|--help         Display help.

  --info            Display .NET information.

  --list-sdks       Display the installed SDKs.

  --list-runtimes   Display the installed runtimes.


path-to-application:

  The path to an application .dll file to execute.



[cecuser@p1274-pvm1 ~]$ dotnet new console -o MyApp -f net7.0


Welcome to .NET 7.0!

---------------------

SDK Version: 7.0.100


----------------

Installed an ASP.NET Core HTTPS development certificate.

To trust the certificate run 'dotnet dev-certs https --trust' (Windows and macOS only).

Learn about HTTPS: https://aka.ms/dotnet-https

----------------

Write your first app: https://aka.ms/dotnet-hello-world

Find out what's new: https://aka.ms/dotnet-whats-new

Explore documentation: https://aka.ms/dotnet-docs

Report issues and find source on GitHub: https://github.com/dotnet/core

Use 'dotnet --help' to see available commands or visit: https://aka.ms/dotnet-cli

--------------------------------------------------------------------------------------

The template "Console App" was created successfully.


Processing post-creation actions...

Restoring /home/cecuser/MyApp/MyApp.csproj:

  Determining projects to restore...

  Restored /home/cecuser/MyApp/MyApp.csproj (in 200 ms).

Restore succeeded.



[cecuser@p1274-pvm1 ~]$ cd MyApp


[cecuser@p1274-pvm1 MyApp]$  ls -la

total 12

drwxrwxr-x.  3 cecuser cecuser   55 Dec  7 02:24 .

drwx------. 23 cecuser cecuser 4096 Dec  7 02:24 ..

-rw-rw-r--.  1 cecuser cecuser  239 Dec  7 02:24 MyApp.csproj

drwxrwxr-x.  2 cecuser cecuser  168 Dec  7 02:24 obj

-rw-rw-r--.  1 cecuser cecuser  103 Dec  7 02:24 Program.cs


[cecuser@p1274-pvm1 MyApp]$ cat Program.cs

// See https://aka.ms/new-console-template for more information

Console.WriteLine("Hello, World!");


[cecuser@p1274-pvm1 MyApp]$ dotnet run

Hello, World!


[cecuser@p1274-pvm1 MyApp]$ ls -ltr

total 8

-rw-rw-r--. 1 cecuser cecuser 239 Dec  7 02:24 MyApp.csproj

-rw-rw-r--. 1 cecuser cecuser 103 Dec  7 02:24 Program.cs

drwxrwxr-x. 3 cecuser cecuser 181 Dec  7 02:26 obj

drwxrwxr-x. 3 cecuser cecuser  19 Dec  7 02:26 bin


[cecuser@p1274-pvm1 MyApp]$ ls -ltr bin/Debug/net7.0

total 168

-rw-rw-r--. 1 cecuser cecuser  10828 Dec  7 02:26 MyApp.pdb

-rw-rw-r--. 1 cecuser cecuser   4608 Dec  7 02:26 MyApp.dll

-rwxr-xr-x. 1 cecuser cecuser 139488 Dec  7 02:26 MyApp

-rw-rw-r--. 1 cecuser cecuser    385 Dec  7 02:26 MyApp.deps.json

-rw-rw-r--. 1 cecuser cecuser    139 Dec  7 02:26 MyApp.runtimeconfig.json


[cecuser@p1274-pvm1 MyApp]$ ./bin/Debug/net7.0/MyApp

Hello, World!




다만 아래의 MS 홈페이지의 .NET download 사이트에서는 ppc64le 관련 패키지는 별도로 제공되지 않네요.


https://devblogs.microsoft.com/dotnet/announcing-dotnet-7/







2020년 12월 8일 화요일

IBM PowerAI (Watson ML Community Edition)이 설치된 Ubuntu ppc64le 기반의 docker image 만들기

 


먼저 다음 링크를 참조하여 ppc64le (IBM POWER9) nvidia-docker2 환경에서 Ubuntu 기반의 docker image를 만듭니다.  


http://hwengineer.blogspot.com/2019/05/ppc64le-ibm-power9-nvidia-docker2.html


참고로 ppc64le (IBM POWER9)에서의 CUDA 설치는 NVIDIA CUDA download page의 안내와 같이 아래처럼 진행하시면 됩니다.


# wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/ppc64el/cuda-repo-ubuntu1804_10.1.105-1_ppc64el.deb

# dpkg -i cuda-repo-ubuntu1804_10.1.105-1_ppc64el.deb

# apt-key adv --fetch-keys http://developer.download.nvidia.com/compute/cuda/repos/ubuntu1804/ppc64el/7fa2af80.pub

# apt-get update

# apt-get install cuda



또는 이미 만들어둔 docker image를 다음과 같이 pull 해와도 됩니다.


# docker pull bsyu/ubuntu18.04_cuda10-1_ppc64le:v0.1


이렇게 pull 받아온 docker image를 확인합니다.


root@unigpu:/files/docker# docker images

REPOSITORY                          TAG                 IMAGE ID            CREATED             SIZE

bsyu/ubuntu18.04_cuda10-1_ppc64le   v0.1                ef8dd4d654e7        2 hours ago         6.33GB

ubuntu                              18.04               ecc8dc2e4170        4 weeks ago         106MB


이 image를 다음과 같이 구동합니다.  


root@unigpu:/files/docker# docker run --runtime=nvidia -ti --rm bsyu/ubuntu18.04_cuda10-1_ppc64lel:v0.1 bash


이제 그 image 속에서 다음과 같이 IBM PowerAI (IBM Watson ML Community Edition)을 설치합니다.  이는 IBM이 마련한 conda channel을 등록하고 거기에서 conda install 명령을 수행하는 방식으로 설치됩니다.


# conda config --prepend channels https://public.dhe.ibm.com/ibmdl/export/pub/software/server/ibm-ai/conda/


# conda create --name wmlce_env python=3.6


# conda activate wmlce_env


# apt-get install openssh-server


# conda install powerai     


위와 같이 conda install powerai 명령을 내리면 tensorflow 뿐만 아니라 caffe, pytorch 등이 모두 설치됩니다.  가령 Tensorflow 1.14만 설치하고자 한다면 위 명령 대신 conda install tensorflow=1.14를 수행하시면 됩니다.


powerai 전체 package 설치는 network 사정에 따라 1~2시간이 걸리기도 합니다.  설치가 완료되면 다음과 같이 docker commit하여 docker image를 저장합니다.


# docker ps -a

CONTAINER ID        IMAGE                                   COMMAND             CREATED             STATUS              PORTS               NAMES

45fb663f025c        bsyu/ubuntu18.04_cuda10-1_ppc64le:v0.1   "bash"              42 seconds ago      Up 39 seconds                           elastic_gates


이어서 v0.2 등의 새로운 tag로 commit 하시면 됩니다.


[root@ac922 docker]# docker commit 45fb663f025c bsyu/ubuntu18.04_cuda10-1_ppc64le:v0.2



아래는 그렇게 만들어진 docker image들의 사용예입니다.   제가 만든 그런 image들은 https://hub.docker.com/u/bsyu 에 올려져 있습니다.



root@unigpu:~# docker run --runtime=nvidia -ti --rm bsyu/ubuntu18.04_cuda10-1_tf1.15_pytorch1.2_ppc64le:latest 


(wmlce_env) root@bdbf11e90094:/# python

Python 3.6.10 |Anaconda, Inc.| (default, Mar 26 2020, 00:22:27)

[GCC 7.3.0] on linux

Type "help", "copyright", "credits" or "license" for more information.


>>> import torch


>>> import tensorflow as tf

2020-05-29 04:53:09.637141: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcudart.so.10.1


>>> sess=tf.Session()

2020-05-29 04:53:20.795027: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcuda.so.1

2020-05-29 04:53:20.825918: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1639] Found device 0 with properties:

name: Tesla V100-SXM2-16GB major: 7 minor: 0 memoryClockRate(GHz): 1.53

pciBusID: 0004:04:00.0

2020-05-29 04:53:20.827162: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1639] Found device 1 with properties:

name: Tesla V100-SXM2-16GB major: 7 minor: 0 memoryClockRate(GHz): 1.53

pciBusID: 0004:05:00.0

2020-05-29 04:53:20.828421: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1639] Found device 2 with properties:

name: Tesla V100-SXM2-16GB major: 7 minor: 0 memoryClockRate(GHz): 1.53

pciBusID: 0035:03:00.0

2020-05-29 04:53:20.829655: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1639] Found device 3 with properties:

name: Tesla V100-SXM2-16GB major: 7 minor: 0 memoryClockRate(GHz): 1.53

pciBusID: 0035:04:00.0

2020년 10월 27일 화요일

IBM POWER9 (ppc64le) Redhat 7에서 Spectrum Scale (GPFS) v5 구성하기

 


여기서의 설정은 gw(2.1.1.1) 서버를 1대 뿐인 GPFS 서버로, 그리고 2대의 서버 tac1과 tac2 (각각 2.1.1.3, 2.1.1.4)를 GPFS client 노드로 등록하는 것입니다.  즉 GPFS의 물리적 disk가 직접 연결되는 것은 gw 서버이고, tac1과 tac2 서버는 gw 서버가 보내주는 GPFS filesystem을 NSD (network storage device) 형태로 network을 통해서 받게 됩니다.


먼저 모든 서버에서 firewalld를 disable합니다.  이와 함께 각 서버 간에 passwd 없이 ssh가 가능하도록 미리 설정해둡니다.


[root@gw ~]# systemctl stop firewalld


[root@gw ~]# systemctl disable firewalld


여기서는 GPFS (새이름 SpectrumScale) installer를 이용하여 설치하겠습니다.  GPFS v5부터는 ansible을 이용하여 1대에서만 설치하면 다른 cluster node들에게도 자동으로 설치가 되어 편합니다.  먼저 install 파일을 수행하면 self-extraction이 시작되며 파일들이 생성됩니다.


[root@gw SW]# ./Spectrum_Scale_Advanced-5.0.4.0-ppc64LE-Linux-install

Extracting Product RPMs to /usr/lpp/mmfs/5.0.4.0 ...

tail -n +641 ./Spectrum_Scale_Advanced-5.0.4.0-ppc64LE-Linux-install | tar -C /usr/lpp/mmfs/5.0.4.0 --wildcards -xvz  installer gpfs_rpms/rhel/rhel7 hdfs_debs/ubuntu16/hdfs_3.1.0.x hdfs_rpms/rhel7/hdfs_2.7.3.x hdfs_rpms/rhel7/hdfs_3.0.0.x hdfs_rpms/rhel7/hdfs_3.1.0.x zimon_debs/ubuntu/ubuntu16 ganesha_rpms/rhel7 ganesha_rpms/rhel8 gpfs_debs/ubuntu16 gpfs_rpms/sles12 object_rpms/rhel7 smb_rpms/rhel7 smb_rpms/rhel8 tools/repo zimon_debs/ubuntu16 zimon_rpms/rhel7 zimon_rpms/rhel8 zimon_rpms/sles12 zimon_rpms/sles15 gpfs_debs gpfs_rpms manifest 1> /dev/null

   - installer

   - gpfs_rpms/rhel/rhel7

   - hdfs_debs/ubuntu16/hdfs_3.1.0.x

   - hdfs_rpms/rhel7/hdfs_2.7.3.x

...

   - gpfs_debs

   - gpfs_rpms

   - manifest


Removing License Acceptance Process Tool from /usr/lpp/mmfs/5.0.4.0 ...

rm -rf  /usr/lpp/mmfs/5.0.4.0/LAP_HOME /usr/lpp/mmfs/5.0.4.0/LA_HOME


Removing JRE from /usr/lpp/mmfs/5.0.4.0 ...

rm -rf /usr/lpp/mmfs/5.0.4.0/ibm-java*tgz


==================================================================

Product packages successfully extracted to /usr/lpp/mmfs/5.0.4.0


   Cluster installation and protocol deployment

      To install a cluster or deploy protocols with the Spectrum Scale Install Toolkit:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale -h

      To install a cluster manually:  Use the gpfs packages located within /usr/lpp/mmfs/5.0.4.0/gpfs_<rpms/debs>


      To upgrade an existing cluster using the Spectrum Scale Install Toolkit:

      1) Copy your old clusterdefinition.txt file to the new /usr/lpp/mmfs/5.0.4.0/installer/configuration/ location

      2) Review and update the config:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale config update

      3) (Optional) Update the toolkit to reflect the current cluster config:

         /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale config populate -N <node>

      4) Run the upgrade:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale upgrade -h


      To add nodes to an existing cluster using the Spectrum Scale Install Toolkit:

      1) Add nodes to the clusterdefinition.txt file:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale node add -h

      2) Install GPFS on the new nodes:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale install -h

      3) Deploy protocols on the new nodes:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale deploy -h


      To add NSDs or file systems to an existing cluster using the Spectrum Scale Install Toolkit:

      1) Add nsds and/or filesystems with:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale nsd add -h

      2) Install the NSDs:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale install -h

      3) Deploy the new file system:  /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale deploy -h


      To update the toolkit to reflect the current cluster config examples:

         /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale config populate -N <node>

      1) Manual updates outside of the install toolkit

      2) Sync the current cluster state to the install toolkit prior to upgrade

      3) Switching from a manually managed cluster to the install toolkit


==================================================================================

To get up and running quickly, visit our wiki for an IBM Spectrum Scale Protocols Quick Overview:

https://www.ibm.com/developerworks/community/wikis/home?lang=en#!/wiki/General%20Parallel%20File%20System%20%28GPFS%29/page/Protocols%20Quick%20Overview%20for%20IBM%20Spectrum%20Scale

===================================================================================



먼저 아래와 같이 spectrumscale 명령으로 gw 서버, 즉 2.1.1.5를 installer 서버로 지정합니다.  


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale setup -s 2.1.1.5


이어서 gw 서버를 manager node이자 admin node로 지정합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale node add 2.1.1.5 -m -n -a

[ INFO  ] Adding node gw as a GPFS node.

[ INFO  ] Adding node gw as a manager node.

[ INFO  ] Adding node gw as an NSD server.

[ INFO  ] Configuration updated.

[ INFO  ] Tip :If all node designations are complete, add NSDs to your cluster definition and define required filessytems:./spectrumscale nsd add <device> -p <primary node> -s <secondary node> -fs <file system>

[ INFO  ] Setting gw as an admin node.

[ INFO  ] Configuration updated.

[ INFO  ] Tip : Designate protocol or nsd nodes in your environment to use during install:./spectrumscale node add <node> -p -n



각 node들을 quorum node로 등록합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale node add 2.1.1.3 -q


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale node add 2.1.1.4 -q


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale node add 2.1.1.5 -q

[ INFO  ] Adding node gwp as a quorum node.



node list를 확인합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale node list

[ INFO  ] List of nodes in current configuration:

[ INFO  ] [Installer Node]

[ INFO  ] 2.1.1.5

[ INFO  ]

[ INFO  ] [Cluster Details]

[ INFO  ] No cluster name configured

[ INFO  ] Setup Type: Spectrum Scale

[ INFO  ]

[ INFO  ] [Extended Features]

[ INFO  ] File Audit logging     : Disabled

[ INFO  ] Watch folder           : Disabled

[ INFO  ] Management GUI         : Disabled

[ INFO  ] Performance Monitoring : Enabled

[ INFO  ] Callhome               : Enabled

[ INFO  ]

[ INFO  ] GPFS  Admin  Quorum  Manager   NSD   Protocol  Callhome   OS   Arch

[ INFO  ] Node   Node   Node     Node   Server   Node     Server

[ INFO  ] gw      X       X       X       X                       rhel7  ppc64le

[ INFO  ] tac1p           X                                       rhel7  ppc64le

[ INFO  ] tac2p           X                                       rhel7  ppc64le

[ INFO  ]

[ INFO  ] [Export IP address]

[ INFO  ] No export IP addresses configured



sdc와 sdd를 nsd로 등록합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale nsd add /dev/sdc -p 2.1.1.5 --name data_nsd -fs data


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale nsd add /dev/sdd -p 2.1.1.5 --name backup_nsd -fs backup


nsd를 확인합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale nsd list

[ INFO  ] Name       FS     Size(GB) Usage   FG Pool    Device   Servers

[ INFO  ] data_nsd   data   400      Default 1  Default /dev/sdc [gwp]

[ INFO  ] backup_nsd backup 400      Default 1  Default /dev/sdd [gwp]


filesystem을 확인합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale filesystem list

[ INFO  ] Name   BlockSize   Mountpoint   NSDs Assigned  Default Data Replicas     Max Data Replicas     Default Metadata Replicas     Max Metadata Replicas

[ INFO  ] data   Default (4M)/ibm/data    1              1                         2                     1                             2

[ INFO  ] backup Default (4M)/ibm/backup  1              1                         2                     1                             2

[ INFO  ]



GPFS cluster를 정의합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale config gpfs -c tac_gpfs

[ INFO  ] Setting GPFS cluster name to tac_gpfs


다른 node들에게의 통신은 ssh와 scp를 이용하는 것으로 지정합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale config gpfs -r /usr/bin/ssh

[ INFO  ] Setting Remote shell command to /usr/bin/ssh


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale config gpfs -rc /usr/bin/scp

[ INFO  ] Setting Remote file copy command to /usr/bin/scp


확인합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale config gpfs --list

[ INFO  ] Current settings are as follows:

[ INFO  ] GPFS cluster name is tac_gpfs.

[ INFO  ] GPFS profile is default.

[ INFO  ] Remote shell command is /usr/bin/ssh.

[ INFO  ] Remote file copy command is /usr/bin/scp.

[ WARN  ] No value for GPFS Daemon communication port range in clusterdefinition file.



기본으로 GPFS 서버는 장애 발생시 IBM으로 연락하는 callhome 기능이 있습니다.  Internet에 연결된 노드가 아니므로 disable합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale callhome disable

[ INFO  ] Disabling the callhome.

[ INFO  ] Configuration updated.


이제 install 준비가 되었습니다.  Install 하기 전에 precheck을 수행합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale install -pr

[ INFO  ] Logging to file: /usr/lpp/mmfs/5.0.4.0/installer/logs/INSTALL-PRECHECK-23-10-2020_21:13:23.log

[ INFO  ] Validating configuration

...

[ INFO  ] The install toolkit will not configure call home as it is disabled. To enable call home, use the following CLI command: ./spectrumscale callhome enable

[ INFO  ] Pre-check successful for install.

[ INFO  ] Tip : ./spectrumscale install


이상 없으면 install을 수행합니다.  이때 gw 뿐만 아니라 tac1과 tac2에도 GPFS가 설치됩니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale install

...

[ INFO  ] GPFS active on all nodes

[ INFO  ] GPFS ACTIVE

[ INFO  ] Checking state of NSDs

[ INFO  ] NSDs ACTIVE

[ INFO  ] Checking state of Performance Monitoring

[ INFO  ] Running Performance Monitoring post-install checks

[ INFO  ] pmcollector running on all nodes

[ INFO  ] pmsensors running on all nodes

[ INFO  ] Performance Monitoring ACTIVE

[ INFO  ] SUCCESS

[ INFO  ] All services running

[ INFO  ] StanzaFile and NodeDesc file for NSD, filesystem, and cluster setup have been saved to /usr/lpp/mmfs folder on node: gwp

[ INFO  ] Installation successful. 3 GPFS nodes active in cluster tac_gpfs.tac1p. Completed in 2 minutes 52 seconds.

[ INFO  ] Tip :If all node designations and any required protocol configurations are complete, proceed to check the deploy configuration:./spectrumscale deploy --precheck



참고로 여기서 아래와 같은 error가 나는 경우는 전에 이미 GPFS NSD로 사용된 disk이기 때문입니다.  


[ FATAL ] gwp failure whilst: Creating NSDs  (SS16)

[ WARN  ] SUGGESTED ACTION(S):

[ WARN  ] Review your NSD device configuration in configuration/clusterdefinition.txt

[ WARN  ] Ensure all disks are not damaged and can be written to.

[ FATAL ] FAILURE REASON(s) for gwp:

[ FATAL ] gwp ---- Begin output of /usr/lpp/mmfs/bin/mmcrnsd -F /usr/lpp/mmfs/StanzaFile  ----

[ FATAL ] gwp STDOUT: mmcrnsd: Processing disk sdc

[ FATAL ] gwp mmcrnsd: Processing disk sdd

[ FATAL ] gwp STDERR: mmcrnsd: Disk device sdc refers to an existing NSD

[ FATAL ] gwp mmcrnsd: Disk device sdd refers to an existing NSD

[ FATAL ] gwp mmcrnsd: Command failed. Examine previous error messages to determine cause.

[ FATAL ] gwp ---- End output of /usr/lpp/mmfs/bin/mmcrnsd -F /usr/lpp/mmfs/StanzaFile  ----

[ INFO  ] Detailed error log: /usr/lpp/mmfs/5.0.4.0/installer/logs/INSTALL-23-10-2020_21:20:05.log

[ FATAL ] Installation failed on one or more nodes. Check the log for more details.


이건 다음과 같이 disk 앞부분 약간을 덮어쓰면 됩니다.


[root@gw SW]# dd if=/dev/zero of=/dev/sdc bs=1M count=100

100+0 records in

100+0 records out

104857600 bytes (105 MB) copied, 0.0736579 s, 1.4 GB/s


[root@gw SW]# dd if=/dev/zero of=/dev/sdd bs=1M count=100

100+0 records in

100+0 records out

104857600 bytes (105 MB) copied, 0.0737598 s, 1.4 GB/s



이제 각 node 상태를 check 합니다.  


[root@gw SW]# mmgetstate -a


 Node number  Node name        GPFS state

-------------------------------------------

       1      tac1p            active

       2      tac2p            active

       3      gwp              active


nsd 상태를 check 합니다.  그런데 GPFS filesystem 정의가 (free disk)로 빠져 있는 것을 보실 수 있습니다.


[root@gw SW]# mmlsnsd


 File system   Disk name    NSD servers

---------------------------------------------------------------------------

 (free disk)   backup_nsd   gwp

 (free disk)   data_nsd     gwp


spectrumscale filesystem list 명령으로 다시 GPFS filesystem 상태를 보면 거기엔 정보가 들어가 있습니다.  다만 mount point가 /ibm/data 이런 식으로 잘못 되어 있네요.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale filesystem list

[ INFO  ] Name   BlockSize   Mountpoint   NSDs Assigned  Default Data Replicas     Max Data Replicas     Default Metadata Replicas     Max Metadata Replicas

[ INFO  ] data   Default (4M)/ibm/data    1              1                         2                     1                             2

[ INFO  ] backup Default (4M)/ibm/backup  1              1                         2                     1                             2

[ INFO  ]


잘못된 mount point들을 제대로 수정합니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale filesystem modify data -m /data

[ INFO  ] The data filesystem will be mounted at /data on all nodes.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale filesystem modify backup -m /backup

[ INFO  ] The backup filesystem will be mounted at /backup on all nodes.


확인합니다.  그러나 여전히 mount는 되지 않습니다.


[root@gw SW]# /usr/lpp/mmfs/5.0.4.0/installer/spectrumscale filesystem list

[ INFO  ] Name   BlockSize   Mountpoint   NSDs Assigned  Default Data Replicas     Max Data Replicas     Default Metadata Replicas     Max Metadata Replicas

[ INFO  ] data   Default (4M)/data        1              1                         2                     1                             2

[ INFO  ] backup Default (4M)/backup      1              1                         2                     1                             2

[ INFO  ]


이를 수정하기 위해 GPFS fileystem 설정은 예전 방식, 즉 mmcrnsd와 mmcrfs 명령을 쓰겠습니다.  먼저 disk description 파일을 아래와 같이 만듭니다.


[root@gw ~]# vi /home/SW/gpfs/disk.desc1

/dev/sdc:gwp::dataAndMetadata:1:nsd_data


[root@gw ~]# vi /home/SW/gpfs/disk.desc2

/dev/sdd:gwp::dataAndMetadata:1:nsd_backup


그리고 예전의 NSD 포맷을 지우기 위해 sdc와 sdd에 아래와 같이 dd로 overwrite를 합니다.


[root@gw ~]# dd if=/dev/zero of=/dev/sdc bs=1M count=100

100+0 records in

100+0 records out

104857600 bytes (105 MB) copied, 0.0130229 s, 8.1 GB/s


[root@gw ~]# dd if=/dev/zero of=/dev/sdd bs=1M count=100

100+0 records in

100+0 records out

104857600 bytes (105 MB) copied, 0.0128207 s, 8.2 GB/s


mmcrnsd 명령과 mmcrfs 명령을 수행하여 NSD와 GPFS filesystem을 만듭니다.


[root@gw ~]# mmcrnsd -F /home/SW/gpfs/disk.desc1


[root@gw ~]# mmcrnsd -F /home/SW/gpfs/disk.desc2


[root@gw ~]# mmcrfs /data /dev/nsd_data -F /home/SW/gpfs/disk.desc1


The following disks of nsd_data will be formatted on node gw:

    nsd_data: size 409600 MB

Formatting file system ...

Disks up to size 3.18 TB can be added to storage pool system.

Creating Inode File

Creating Allocation Maps

Creating Log Files

Clearing Inode Allocation Map

Clearing Block Allocation Map

Formatting Allocation Map for storage pool system

Completed creation of file system /dev/nsd_data.

mmcrfs: Propagating the cluster configuration data to all

  affected nodes.  This is an asynchronous process.



[root@gw ~]# mmcrfs /backup /dev/nsd_backup -F /home/SW/gpfs/disk.desc2


The following disks of nsd_backup will be formatted on node gw:

    nsd_backup: size 409600 MB

Formatting file system ...

Disks up to size 3.18 TB can be added to storage pool system.

Creating Inode File

Creating Allocation Maps

Creating Log Files

Clearing Inode Allocation Map

Clearing Block Allocation Map

Formatting Allocation Map for storage pool system

Completed creation of file system /dev/nsd_backup.

mmcrfs: Propagating the cluster configuration data to all

  affected nodes.  This is an asynchronous process.


이제 모든 node에서 mount 해봅니다.


[root@gw ~]# mmmount all -a

Sat Oct 24 09:45:43 KST 2020: mmmount: Mounting file systems ...


[root@gw ~]# df -h

Filesystem      Size  Used Avail Use% Mounted on

devtmpfs        1.7G     0  1.7G   0% /dev

tmpfs           1.8G   18M  1.8G   1% /dev/shm

tmpfs           1.8G   96M  1.7G   6% /run

tmpfs           1.8G     0  1.8G   0% /sys/fs/cgroup

/dev/sda5        50G  5.4G   45G  11% /

/dev/sda6       345G  8.6G  337G   3% /home

/dev/sda2      1014M  178M  837M  18% /boot

tmpfs           355M     0  355M   0% /run/user/0

/dev/sr0        3.4G  3.4G     0 100% /home/cdrom

nsd_backup      400G  2.8G  398G   1% /backup

nsd_data        400G  2.8G  398G   1% /data


테스트를 위해 /data 밑에 /etc/hosts 파일을 copy해 둡니다.


[root@gw ~]# cp /etc/hosts /data


[root@gw ~]# ls -l /data

total 1

-rw-r--r--. 1 root root 298 Oct 24 09:49 hosts



Client node들에서도 잘 mount 되었는지 확인합니다.  그리고 아까 copy해둔 hosts 파일이 있는 확인합니다.


[root@gw ~]# ssh tac1

Last login: Sat Oct 24 09:33:46 2020 from gwp


[root@tac1 ~]# df -h

Filesystem      Size  Used Avail Use% Mounted on

devtmpfs         28G     0   28G   0% /dev

tmpfs            28G     0   28G   0% /dev/shm

tmpfs            28G   15M   28G   1% /run

tmpfs            28G     0   28G   0% /sys/fs/cgroup

/dev/sda5        50G  3.3G   47G   7% /

/dev/sda6       321G  2.8G  319G   1% /home

/dev/sda2      1014M  178M  837M  18% /boot

tmpfs           5.5G     0  5.5G   0% /run/user/0

nsd_data        400G  2.8G  398G   1% /data

nsd_backup      400G  2.8G  398G   1% /backup


[root@tac1 ~]# ls -l /data

total 1

-rw-r--r--. 1 root root 298 Oct 24 09:49 hosts




[root@gw ~]# ssh tac2

Last login: Sat Oct 24 09:33:46 2020 from gwp


[root@tac2 ~]# df -h

Filesystem      Size  Used Avail Use% Mounted on

devtmpfs         28G     0   28G   0% /dev

tmpfs            28G     0   28G   0% /dev/shm

tmpfs            28G   15M   28G   1% /run

tmpfs            28G     0   28G   0% /sys/fs/cgroup

/dev/sda5        50G  3.2G   47G   7% /

/dev/sda6       321G  3.3G  318G   2% /home

/dev/sda2      1014M  178M  837M  18% /boot

tmpfs           5.5G     0  5.5G   0% /run/user/0

nsd_backup      400G  2.8G  398G   1% /backup

nsd_data        400G  2.8G  398G   1% /data



[root@tac2 ~]# ls -l /data

total 1

-rw-r--r--. 1 root root 298 Oct 24 09:49 hosts





참고로 어떤 disk가 GPFS nsd인지는 fdisk 명령으로 아래와 같이 확인하실 수 있습니다.  fdisk -l 로 볼 때, 아래와 같이 IBM General Par GPFS라고 나오는 것이 GPFS nsd 입니다.



[root@tac1 ~]# fdisk -l | grep sd

WARNING: fdisk GPT support is currently new, and therefore in an experimental phase. Use at your own discretion.

WARNING: fdisk GPT support is currently new, and therefore in an experimental phase. Use at your own discretion.

Disk /dev/sda: 429.5 GB, 429496729600 bytes, 838860800 sectors

/dev/sda1   *        2048       10239        4096   41  PPC PReP Boot

/dev/sda2           10240     2107391     1048576   83  Linux

/dev/sda3         2107392    60829695    29361152   82  Linux swap / Solaris

/dev/sda4        60829696   838860799   389015552    5  Extended

/dev/sda5        60831744   165689343    52428800   83  Linux

/dev/sda6       165691392   838860799   336584704   83  Linux

Disk /dev/sdb: 429.5 GB, 429496729600 bytes, 838860800 sectors

Disk /dev/sdc: 429.5 GB, 429496729600 bytes, 838860800 sectors

Disk /dev/sdd: 429.5 GB, 429496729600 bytes, 838860800 sectors

Disk /dev/sde: 429.5 GB, 429496729600 bytes, 838860800 sectors

Disk /dev/sdf: 429.5 GB, 429496729600 bytes, 838860800 sectors

Disk /dev/sdg: 429.5 GB, 429496729600 bytes, 838860800 sectors



[root@tac1 ~]# fdisk -l /dev/sdb

WARNING: fdisk GPT support is currently new, and therefore in an experimental phase. Use at your own discretion.


Disk /dev/sdb: 429.5 GB, 429496729600 bytes, 838860800 sectors

Units = sectors of 1 * 512 = 512 bytes

Sector size (logical/physical): 512 bytes / 512 bytes

I/O size (minimum/optimal): 512 bytes / 512 bytes

Disk label type: gpt

Disk identifier: 236CE033-C570-41CC-8D2E-E20E6F494C38



#         Start          End    Size  Type            Name

 1           48    838860751    400G  IBM General Par GPFS:



[root@tac1 ~]# fdisk -l /dev/sdc

WARNING: fdisk GPT support is currently new, and therefore in an experimental phase. Use at your own discretion.


Disk /dev/sdc: 429.5 GB, 429496729600 bytes, 838860800 sectors

Units = sectors of 1 * 512 = 512 bytes

Sector size (logical/physical): 512 bytes / 512 bytes

I/O size (minimum/optimal): 512 bytes / 512 bytes

Disk label type: gpt

Disk identifier: 507A299C-8E96-49E2-8C25-9D051BC9B935



#         Start          End    Size  Type            Name

 1           48    838860751    400G  IBM General Par GPFS:



일반 disk는 아래와 같이 평범하게 나옵니다.


[root@tac1 ~]# fdisk -l /dev/sdd


Disk /dev/sdd: 429.5 GB, 429496729600 bytes, 838860800 sectors

Units = sectors of 1 * 512 = 512 bytes

Sector size (logical/physical): 512 bytes / 512 bytes

I/O size (minimum/optimal): 512 bytes / 512 bytes



IBM POWER9 Redhat 7에서 Redhat HA Cluster 구성하는 방법

 


Redhat HA cluster를 IBM POWER9 (ppc64le) 기반의 Redhat 7에서 설치하는 방법입니다.


먼저 firewalld를 stop 시킵니다.


[root@ha1 ~]# systemctl stop firewalld


[root@ha1 ~]# systemctl disable firewalld


아래의 package들을 설치합니다.  이건 Redhat OS DVD에는 없고 별도의 yum repository에 들어있습니다.  ppc64le의 경우엔 rhel-ha-for-rhel-7-server-for-power-le-rpms 라는 yum repo에 있습니다.


[root@ha1 ~]# yum install pcs fence-agents-all


설치하면 hacluster라는 user가 자동 생성되는데 여기에 passwd를지정해줘야 합니다.


[root@ha1 ~]# passwd hacluster


그리고 pcsd daemon을 start 합니다.  Reboot 후에도 자동 start 되도록 enable도 합니다.


[root@ha1 ~]# systemctl start pcsd.service


[root@ha1 ~]# systemctl enable pcsd.service


참여할 node에 아래와 같이 인증 작업을 합니다.


[root@ha1 ~]# pcs cluster auth ha1 ha2

Username: hacluster

Password:

ha1: Authorized

ha2: Authorized


간단히 아래와 같이 corosysnc.conf 파일을 만듭니다.  ha1, ha2 노드는 물론 /etc/hosts에 등록된 IP 주소입니다.


[root@ha1 ~]# vi /etc/corosync/corosync.conf

totem {

version: 2

secauth: off

cluster_name: tibero_cluster

transport: udpu

}


nodelist {

  node {

        ring0_addr: ha1

        nodeid: 1

       }

  node {

        ring0_addr: ha2

        nodeid: 2

       }

}


quorum {

provider: corosync_votequorum

two_node: 1

}


logging {

to_syslog: yes

}


Cluster를 전체 node에서 enable합니다.


[root@ha1 ~]# pcs cluster enable --all

ha1: Cluster Enabled

ha2: Cluster Enabled


다음과 같이 cluster start 합니다.


[root@ha1 ~]# pcs cluster start --all

ha1: Starting Cluster (corosync)...

ha2: Starting Cluster (corosync)...

ha1: Starting Cluster (pacemaker)...

ha2: Starting Cluster (pacemaker)...



상태 확인해봅니다.


[root@ha1 ~]# pcs cluster status

Cluster Status:

 Stack: unknown

 Current DC: NONE

 Last updated: Mon Oct 26 10:14:59 2020

 Last change: Mon Oct 26 10:14:55 2020 by hacluster via crmd on ha1

 2 nodes configured

 0 resource instances configured


PCSD Status:

  ha1: Online

  ha2: Online


이때 ha2 노드에 가보면 corosysnc.conf 파일은 없습니다.   


[root@ha2 ~]# ls -l /etc/corosync/

total 12

-rw-r--r--. 1 root root 2881 Jun  5 23:10 corosync.conf.example

-rw-r--r--. 1 root root  767 Jun  5 23:10 corosync.conf.example.udpu

-rw-r--r--. 1 root root 3278 Jun  5 23:10 corosync.xml.example

drwxr-xr-x. 2 root root    6 Jun  5 23:10 uidgid.d


이걸 ha2에서 생성시키려면 cluster를 sync하면 됩니다.


[root@ha1 ~]# pcs cluster sync

ha1: Succeeded

ha2: Succeeded


생성된 것을 확인하실 수 있습니다.


[root@ha2 ~]# ls -l /etc/corosync/

total 16

-rw-r--r--. 1 root root  295 Oct 26 10:17 corosync.conf

-rw-r--r--. 1 root root 2881 Jun  5 23:10 corosync.conf.example

-rw-r--r--. 1 root root  767 Jun  5 23:10 corosync.conf.example.udpu

-rw-r--r--. 1 root root 3278 Jun  5 23:10 corosync.xml.example

drwxr-xr-x. 2 root root    6 Jun  5 23:10 uidgid.d



이제 cluster resource를 확인합니다.  당연히 아직 정의된 것이 없습니다.


[root@ha1 ~]# pcs resource show

NO resources configured



두 node 사이에서 failover 받을 cluster의 virtual IP를 아래와 같이 VirtualIP라는 resource ID 이름으로 등록합니다.   참고로 ha1은 10.1.1.1, ha2는 10.1.1.2이고 모두 eth1에 부여된 IP입니다.


[root@ha1 ~]# pcs resource create VirtualIP ocf:heartbeat:IPaddr2 ip=10.1.1.11 cidr_netmask=24 nic=eth1 op monitor interval=30s


[root@ha1 ~]# pcs resource enable VirtualIP


이제 다시 resource를 봅니다.


[root@ha1 ~]# pcs resource show

 VirtualIP      (ocf::heartbeat:IPaddr2):       Stopped


아직 VirtualIP가 stopped 상태인데, 이는 아직 STONITH가 enable 되어 있는 default 상태이기 때문입니다.  STONITH는 split-brain을 방지하기 위한 장치인데, 지금 당장은 disable 하겠습니다.


[root@ha1 ~]# pcs property set stonith-enabled=false


Verify를 해봅니다.  아무 메시지 없으면 통과입니다.


[root@ha1 ~]# crm_verify -L


이제 다시 status를 보면 VirtualIP가 살아 있는 것을 보실 수 있습니다.


[root@ha1 ~]# pcs status

Cluster name: tibero_cluster

Stack: corosync

Current DC: ha1 (version 1.1.23-1.el7-9acf116022) - partition with quorum

Last updated: Mon Oct 26 11:18:31 2020

Last change: Mon Oct 26 11:17:31 2020 by root via cibadmin on ha1


2 nodes configured

1 resource instance configured


Online: [ ha1 ha2 ]


Full list of resources:


 VirtualIP      (ocf::heartbeat:IPaddr2):       Started ha1


Daemon Status:

  corosync: active/enabled

  pacemaker: active/enabled

  pcsd: active/enabled



또 IP address를 보면 10.1.1.11이 eth1에 붙은 것도 보실 수 있습니다.


[root@ha1 ~]# ip a | grep 10.1.1

    inet 10.1.1.1/24 brd 10.1.1.255 scope global noprefixroute eth1

    inet 10.1.1.11/24 brd 10.1.1.255 scope global secondary eth1


다른 node에서 10.1.1.11 (havip)로 ping을 해보면 잘 됩니다.


[root@gw ~]# ping havip

PING havip (10.1.1.11) 56(84) bytes of data.

64 bytes from havip (10.1.1.11): icmp_seq=1 ttl=64 time=0.112 ms

64 bytes from havip (10.1.1.11): icmp_seq=2 ttl=64 time=0.040 ms


이 resource가 failover된 이후 죽었던 node가 되살아나면 원래의 node로 failback 하게 하려면 아래와 같이 합니다.


[root@ha1 ~]# pcs resource defaults resource-stickiness=100

Warning: Defaults do not apply to resources which override them with their own defined values


[root@ha1 ~]# pcs resource defaults

resource-stickiness=100


이제 sdb disk를 이용하여 LVM 작업을 합니다.  여기서는 /data와 /backup이 VirtualIP와 함께 ha1에 mount 되어 있다가 유사시 ha2로 failover 되도록 하고자 합니다.


[root@ha1 ~]# pvcreate /dev/sdb


[root@ha1 ~]# vgcreate datavg /dev/sdb


[root@ha1 ~]# lvcreate -L210000 -n datalv datavg


[root@ha1 ~]# lvcreate -L150000 -n backuplv datavg


[root@ha1 ~]# vgs

  VG     #PV #LV #SN Attr   VSize    VFree

  datavg   1   2   0 wz--n- <400.00g 48.43g


[root@ha1 ~]# lvs

  LV       VG     Attr       LSize    Pool Origin Data%  Meta%  Move Log Cpy%Sync Convert

  backuplv datavg -wi-a-----  146.48g                                           

  datalv   datavg -wi-a----- <205.08g         



[root@ha1 ~]# mkfs.ext4 /dev/datavg/datalv


[root@ha1 ~]# mkfs.ext4 /dev/datavg/backuplv



[root@ha1 ~]# mkdir /data

[root@ha1 ~]# mkdir /backup


[root@ha1 ~]# ssh ha2 mkdir /data

[root@ha1 ~]# ssh ha2 mkdir /backup


이제 이 VG와 filesystem들이 한쪽 node에만, 그것도 OS가 아니라 HA cluster (pacemaker)에 의해서만 mount 되도록 설정합니다.


[root@ha1 ~]# grep use_lvmetad /etc/lvm/lvm.conf

        use_lvmetad = 1


[root@ha1 ~]# lvmconf --enable-halvm --services --startstopservices


[root@ha1 ~]# grep use_lvmetad /etc/lvm/lvm.conf

    use_lvmetad = 0


[root@ha1 ~]# vgs --noheadings -o vg_name

  datavg


그리고 이 VG와 LV, filesystem을 pcs에 등록합니다.


[root@ha1 ~]# pcs resource create tibero_vg LVM volgrpname=datavg exclusive=true --group tiberogroup

Assumed agent name 'ocf:heartbeat:LVM' (deduced from 'LVM')


[root@ha1 ~]# pcs resource create tibero_data Filesystem device="/dev/datavg/datalv" directory="/data" fstype="ext4" --group tiberogroup

Assumed agent name 'ocf:heartbeat:Filesystem' (deduced from 'Filesystem')


[root@ha1 ~]# pcs resource create tibero_backup Filesystem device="/dev/datavg/backuplv" directory="/backup" fstype="ext4" --group tiberogroup

Assumed agent name 'ocf:heartbeat:Filesystem' (deduced from 'Filesystem')


[root@ha1 ~]# pcs resource update VirtualIP --group tiberogroup


그리고 VirtualIP가 항상 이 filesytem들과 함께 움직이도록 colocation constraint를 줍니다.


[root@ha1 ~]# pcs constraint colocation add tiberogroup with VirtualIP INFINITY


그리고 아래 내용은 두 node에서 모두 수행합니다.  ha2도 reboot해야 거기서 datavg 및 거기에 든 LV들이 인식됩니다.


[root@ha1 ~]# vi /etc/lvm/lvm.conf

...

 volume_list = [ ]

...


[root@ha1 ~]# dracut -H -f /boot/initramfs-$(uname -r).img $(uname -r)


[root@ha1 ~]# shutdown -r now



Reboot 이후 보면 VirtualIP나 /data, /backup filesystem이 모두 ha1에서 mount 되어 있는 것을 보실 수 있습니다.



[root@ha1 ~]# pcs status

Cluster name: tibero_cluster

Stack: corosync

Current DC: ha1 (version 1.1.23-1.el7-9acf116022) - partition with quorum

Last updated: Mon Oct 26 13:44:36 2020

Last change: Mon Oct 26 13:44:14 2020 by root via cibadmin on ha1


2 nodes configured

4 resource instances configured


Online: [ ha1 ha2 ]


Full list of resources:


 VirtualIP      (ocf::heartbeat:IPaddr2):       Started ha1

 Resource Group: tiberogroup

     tibero_vg  (ocf::heartbeat:LVM):   Started ha1

     tibero_data        (ocf::heartbeat:Filesystem):    Started ha1

     tibero_backup      (ocf::heartbeat:Filesystem):    Started ha1


Daemon Status:

  corosync: active/enabled

  pacemaker: active/enabled

  pcsd: active/enabled



ha1 노드를 죽여버리면 곧 VirtualIP와 filesystem들이 자동으로 ha2에 failover 되어 있는 것을 확인하실 수 있습니다.


[root@ha1 ~]# halt -f

Halting.


------------


[root@ha2 ~]# df -h

Filesystem                   Size  Used Avail Use% Mounted on

devtmpfs                      28G     0   28G   0% /dev

tmpfs                         28G   58M   28G   1% /dev/shm

tmpfs                         28G   14M   28G   1% /run

tmpfs                         28G     0   28G   0% /sys/fs/cgroup

/dev/sda5                     50G  2.6G   48G   6% /

/dev/sda6                    321G  8.5G  313G   3% /home

/dev/sda2                   1014M  231M  784M  23% /boot

tmpfs                        5.5G     0  5.5G   0% /run/user/0

/dev/mapper/datavg-datalv    202G   61M  192G   1% /data

/dev/mapper/datavg-backuplv  145G   61M  137G   1% /backup


[root@ha2 ~]# ip a | grep 10.1.1

    inet 10.1.1.2/24 brd 10.1.1.255 scope global noprefixroute eth1

    inet 10.1.1.11/24 brd 10.1.1.255 scope global secondary eth1



ha1이 죽은 상태에서의 status는 아래와 같이 나옵니다.


[root@ha2 ~]# pcs status

Cluster name: tibero_cluster

Stack: corosync

Current DC: ha2 (version 1.1.23-1.el7-9acf116022) - partition with quorum

Last updated: Mon Oct 26 13:48:49 2020

Last change: Mon Oct 26 13:47:18 2020 by root via cibadmin on ha1


2 nodes configured

4 resource instances configured


Online: [ ha2 ]

OFFLINE: [ ha1 ]


Full list of resources:


 VirtualIP      (ocf::heartbeat:IPaddr2):       Started ha2

 Resource Group: tiberogroup

     tibero_vg  (ocf::heartbeat:LVM):   Started ha2

     tibero_data        (ocf::heartbeat:Filesystem):    Started ha2

     tibero_backup      (ocf::heartbeat:Filesystem):    Started ha2


Daemon Status:

  corosync: active/enabled

  pacemaker: active/enabled

  pcsd: active/enabled


이 상태에서 pcs cluster를 stop 시키면 VirtualIP와 filesystem들이 모두 내려갑니다.


[root@ha2 ~]# pcs cluster stop --force

Stopping Cluster (pacemaker)...

Stopping Cluster (corosync)...



[root@ha2 ~]# df -h

Filesystem      Size  Used Avail Use% Mounted on

devtmpfs         28G     0   28G   0% /dev

tmpfs            28G     0   28G   0% /dev/shm

tmpfs            28G   14M   28G   1% /run

tmpfs            28G     0   28G   0% /sys/fs/cgroup

/dev/sda5        50G  2.6G   48G   6% /

/dev/sda6       321G  8.5G  313G   3% /home

/dev/sda2      1014M  231M  784M  23% /boot

tmpfs           5.5G     0  5.5G   0% /run/user/0



[root@ha2 ~]# ip a | grep 10.1.1

    inet 10.1.1.2/24 brd 10.1.1.255 scope global noprefixroute eth1




2020년 7월 6일 월요일

Ubuntu 18.04 (ppc64le, IBM POWER9)에서 잊어버린 root passwd reset 하는 방법


먼저 system booting할 때 petit-boot menu까지 나오면, 거기서 맨 아래줄의 'Exit to shell' 메뉴를 선택합니다.

여기서 'fdisk -l' 명령을 내리면 어떤 disk들이 있는지, 그리고 어느 disk partition에 OS가 들어있는지 보실 수 있습니다.  제가 겪은 경우에는 sda와 sdb의 2개 disk가 있었고, 그 중 sda에서 dm-0, dm-1, dm-2의 3개 device가 보였는데 그 size를 보면 dm-0는 PReP partition, dm-2는 SWAP partition이므로 아마 dm-1이 OS partition이라고 판단되었습니다.  그걸 /mnt에 mount 합니다.

# mount /dev/dm-1 /mnt

이제 이 /mnt 속을 보면 etc나 lib, usr, var 등과 같이 OS가 설치된 것이 보일 것입니다.  이제 chroot 명령으로 /mnt를 /로 바꿉니다.

# chroot /mnt

이제 dm-1 속의 OS image를 /로 mount 한 것입니다.  이제 passwd를 바꿔줍니다.

# passwd

그리고나서 Ctrl-D로 빠져나와 정상적으로 booting하면 됩니다.

2020년 4월 1일 수요일

LSF HPC Suite v10.2 설치 - CentOS 7.6 (ppc64le, IBM POWER8) 환경


LSF HPC Suite v10.2를 CentOS 7.6 (ppc64le, IBM POWER8) 환경에서 설치하는 과정을 정리했습니다.

먼저 LSF 설치에 필요한 사전 준비를 합니다.  Password를 넣지 않고 ssh가 되도록 하는 것이 관건입니다.  저는 일단 root user로도 ssh가 가능하도록 아래와 같이 설정했습니다.  여기서는 p628-kvm1~3 총 3대의 가상서버가 있습니다.

[cecuser@p628-kvm1 ~]$ sudo yum install -y openssh-clients

[cecuser@p628-kvm1 ~]$ ssh-keygen

[cecuser@p628-kvm1 ~]$ ssh-copy-id p628-kvm1

[cecuser@p628-kvm1 ~]$ ssh-copy-id p628-kvm2

[cecuser@p628-kvm1 ~]$ ssh-copy-id p628-kvm3

[cecuser@p628-kvm1 ~]$ sudo vi /etc/ssh/sshd_config
#PermitRootLogin yes
PermitRootLogin yes

[cecuser@p628-kvm1 ~]$ sudo systemctl restart sshd

[cecuser@p628-kvm1 ~]$ su -
Password:

-bash-4.2# ssh-keygen

-bash-4.2# ssh-copy-id p628-kvm1

-bash-4.2# ssh-copy-id p628-kvm2

-bash-4.2# ssh-copy-id p628-kvm3


이제 LSF HPC Suite의 installation bin file을 수행합니다.  이 파일이 수행하는 것은 LSF를 설치하는 것이 아니라, 그 설치를 위한 ansible 기반의 installer를 설치 및 구성하는 것입니다.

[cecuser@p628-kvm1 ~]$ sudo ./lsfshpc10.2.0.9-ppc64le.bin

Installing LSF Suite for HPC 10.2.0.9 deployer ...

Checking prerequisites ...
httpd is not installed.
Installing httpd ...
createrepo is not installed.
Installing createrepo ...
Installed Packages
rsync.ppc64le                      3.1.2-6.el7_6.1                      @updates
Installed Packages
curl.ppc64le                         7.29.0-51.el7                         @base
Installed Packages
python-jinja2.noarch                   2.7.2-3.el7_6                    @updates
Installed Packages
yum-utils.noarch                       1.1.31-50.el7                       @base
httpd is not running. Trying to start httpd ...
Copying LSF Suite for HPC 10.2.0.9 deployer files ...
Installing Ansible ...
Spawning worker 0 with 1 pkgs
,,,,
Gathering deployer information ...
Creating LSF Suite repository ...
LSF Suite for HPC 10.2.0.9 deployer installed

To deploy LSF Suite to a cluster:
 - Change directory to "/opt/ibm/lsf_installer/playbook"
 - Edit the "lsf-inventory" file with lists of machines and their roles in the cluster
 - Edit the "lsf-config.yml" file with required parameters
 - Run "ansible-playbook -i lsf-inventory lsf-deploy.yml"

Installing ibm-jre ...

Connecting to IBM Support: Fix Central to query the latest fix pack level ...
This is the latest available fix pack.

이제 아래와 같이 lsf-inventory 파일 속의 hostname 등을 등록합니다.  여기서는 p628-kvm1  p628-kvm2  p628-kvm3의 총 3대 중에서, p628-kvm1이 master 역할을 함과 동시에 compute host 역할까지 하는 것으로 설정했습니다.  p628-kvm1에서는 연산작업을 하지 않고 master 역할만 하기를 원한다면 [LSF_Servers]에서는 p628-kvm1이 빠져야 합니다.

[cecuser@p628-kvm1 ~]$ cd /opt/ibm/lsf_installer/playbook

[cecuser@p628-kvm1 playbook]$ sudo vi lsf-inventory
[LSF_Masters]
p628-kvm1
[LSF_Servers]
p628-kvm[1:3]     # p628-kvm1  p628-kvm2  p628-kvm3을 쉽게 표현한 것입니다

같은 directory의  lsf-config.yml도 필요시 수정합니다만, 여기서는 일단 하지 않겠습니다.

[cecuser@p628-kvm1 playbook]$ sudo vi  lsf-config.yml
(필요시)
  my_cluster_name: myCluster   # 이것이 default cluster name입니다.  여기서는 굳이 cluster 이름을 수정하지 않았습니다.
  HA_shared_dir: none  # High-Availability를 위한 directory로서, 설정하면 여기에 configuration file들과 work directory가 copy됩니다.
  NFS_install_dir: none  # 여기에 NFS mount point를 등록하면 그 directory에  LSF master, server & client binary file들과 config file들이 설치되어 cluster 내에서 공유됩니다

수정이 끝났으면 다음과 같이 문법에 오류가 없는지 테스트를 해봅니다.

[cecuser@p628-kvm1 playbook]$  sudo ansible-playbook -i lsf-inventory lsf-config-test.yml
...
TASK [Check if Private_IPv4_Range is valid] ************************************

TASK [Check if Private_IPv4_Range IP address is available] *********************

PLAY RECAP *********************************************************************
localhost                  : ok=0    changed=0    unreachable=0    failed=0
p628-kvm1                  : ok=16   changed=0    unreachable=0    failed=0
p628-kvm2                  : ok=11   changed=0    unreachable=0    failed=0
p628-kvm3                  : ok=11   changed=0    unreachable=0    failed=0

원래 manual에는 아래와 같이  lsf-predeploy-test.yml를 이용해서 추가 검증 테스트를 해보라고 되어 있는데, 해보면 아래와 같이 remote 서버 2대에서만 엉뚱한 error가 납니다.  아무리 이것저것 해봐도 계속 같은 오류가 나던데, 그냥 무시하셔도 될 것 같습니다.  저는 무시하고 그냥 진행했는데, 결과적으로는 문제가 없었습니다.

[cecuser@p628-kvm1 playbook]$ sudo ansible-playbook -i lsf-inventory lsf-predeploy-test.yml
...
TASK [debug deployer] **********************************************************
fatal: [p628-kvm2]: FAILED! => {"failed": true, "msg": "The conditional check 'LSF.Private_IPv4_Range is defined and LSF.Private_IPv4_Range != 'none' and LSF.Private_IPv4_Range is not none' failed. The error was: error while evaluating conditional (LSF.Private_IPv4_Range is defined and LSF.Private_IPv4_Range != 'none' and LSF.Private_IPv4_Range is not none): 'LSF' is undefined\n\nThe error appears to have been in '/opt/ibm/lsf_installer/playbook/roles/config/tasks/check_private_ips.yml': line 8, column 3, but may\nbe elsewhere in the file depending on the exact syntax problem.\n\nThe offending line appears to be:\n\n\n- name: debug {{ target_role }}\n  ^ here\nWe could be wrong, but this one looks like it might be an issue with\nmissing quotes.  Always quote template expression brackets when they\nstart a value. For instance:\n\n    with_items:\n      - {{ foo }}\n\nShould be written as:\n\n    with_items:\n      - \"{{ foo }}\"\n"}

이제 다음과 같이 ansible-playbook으로 실제 설치를 수행합니다.  이 명령 하나로 3대의 서버에 모두 LSF HPC Suite가 설치되는 것입니다.  위의 test 오류에도 불구하고, 설치는 전혀 이상없이 잘 됩니다.  서버 및 네트워크 사양에 따라 다르겠지만 시간은 꽤 많이 걸립니다.  저는 1시간 정도 걸렸습니다.   아래와 같이 끝 부분에 IBM Spectrum Application Center를 열 수 있는 URL이 주어집니다.  접속해보면 실제로 잘 열립니다.

[cecuser@p628-kvm1 playbook]$ sudo ansible-playbook -i lsf-inventory lsf-deploy.yml
...
PLAY [Summary] *****************************************************************

TASK [debug] *******************************************************************
ok: [p628-kvm1 -> localhost] => {
    "msg": [
        "LSF Suite for HPC 10.2.0.9 deployment is done",
        "Open IBM Spectrum Application Center with the following URL: http://p628-kvm1.cecc.ihost.com:8080"
    ]
}

PLAY RECAP *********************************************************************
localhost                  : ok=3    changed=0    unreachable=0    failed=0
p628-kvm1                  : ok=296  changed=127  unreachable=0    failed=0
p628-kvm2                  : ok=107  changed=48   unreachable=0    failed=0
p628-kvm3                  : ok=107  changed=48   unreachable=0    failed=0





** 혹시 위의 설치 과정에서 뭔가 실수 등으로 잘못되어 다시 설치하고자 한다면, 그냥 다시 하면 안되고 다음과 같이 uninstall 하고 진행해야 합니다.  그러지 않을 경우 "Check if an installation was made before" 등의 error가 나면서 fail 합니다.

[cecuser@p628-kvm1 playbook]$ sudo ansible-playbook -i lsf-inventory lsf-uninstall.yml

이제 다음과 같이 lsf.conf에서 RSH이나 RCP 대신 ssh, scp를 쓰도록 설정합니다.

[cecuser@p628-kvm1 playbook]$ sudo vi /opt/ibm/lsfsuite/lsf/conf/lsf.conf
LSF_RSH=ssh
LSF_REMOTE_COPY_CMD="scp -B -o 'StrictHostKeyChecking no'"

[cecuser@p628-kvm1 ~]$ sudo ln -s /opt/ibm/lsfsuite/lsf/conf/lsf.conf /etc/lsf.conf

이제 lsfshutdown & lsfstartup을 통해 restart 합니다.  이건 root에서 해야 합니다.

[cecuser@p628-kvm1 ~]$ su -
Password:
Last login: Tue Mar 31 05:05:59 EDT 2020

-bash-4.2# . /opt/ibm/lsfsuite/lsf/conf/profile.lsf   (이 profile.lsf를 수행해야 PATH 등의 환경변수가 자동설정됩니다.)

-bash-4.2# lsfshutdown
Shutting down all slave batch daemons ...
Shut down slave batch daemon on all the hosts? [y/n] y
Shut down slave batch daemon on <p628-kvm1> ...... done
Shut down slave batch daemon on <p628-kvm2.cecc.ihost.com> ...... done
Shut down slave batch daemon on <p628-kvm3.cecc.ihost.com> ...... done
Shutting down all RESes ...
Do you really want to shut down RES on all hosts? [y/n] y
Shut down RES on <p628-kvm1> ...... done
Shut down RES on <p628-kvm3.cecc.ihost.com> ...... done
Shut down RES on <p628-kvm2.cecc.ihost.com> ...... done
Shutting down all LIMs ...
Do you really want to shut down LIMs on all hosts? [y/n] y
Shut down LIM on <p628-kvm1> ...... done
Shut down LIM on <p628-kvm3.cecc.ihost.com> ...... done
Shut down LIM on <p628-kvm2.cecc.ihost.com> ...... done

-bash-4.2# lsfstartup
Starting up all LIMs ...
Do you really want to start up LIM on all hosts ? [y/n]y
Start up LIM on <p628-kvm1> ......
Starting up all RESes ...
Do you really want to start up RES on all hosts ? [y/n]y
Start up RES on <p628-kvm1> ......
Starting all slave daemons on LSBATCH hosts ...
Do you really want to start up slave batch daemon on all hosts ? [y/n] y
Start up slave batch daemon on <p628-kvm1> ......
Done starting up LSF daemons on the local LSF cluster ...

이제 다시 일반 user에서 각종 LSF 명령을 내려 봅니다.

[cecuser@p628-kvm1 ~]$ . /opt/ibm/lsfsuite/lsf/conf/profile.lsf

[cecuser@p628-kvm1 ~]$ lsid
IBM Spectrum LSF 10.1.0.9, Oct 16 2019
Suite Edition: IBM Spectrum LSF Suite for HPC 10.2.0.9
Copyright International Business Machines Corp. 1992, 2016.
US Government Users Restricted Rights - Use, duplication or disclosure restricted by GSA ADP Schedule Contract with IBM Corp.

My cluster name is myCluster
My master name is p628-kvm1

[cecuser@p628-kvm1 ~]$ lshosts
HOST_NAME      type    model  cpuf ncpus maxmem maxswp server RESOURCES
p628-kvm1   LINUXPP   POWER8  25.0     2  31.9G   3.9G    Yes (mg)
p628-kvm3.c LINUXPP   POWER8  25.0     2  31.9G   3.9G    Yes ()
p628-kvm2.c LINUXPP   POWER8  25.0     2  31.9G   3.9G    Yes ()


[cecuser@p628-kvm1 ~]$ bhosts
HOST_NAME          STATUS       JL/U    MAX  NJOBS    RUN  SSUSP  USUSP    RSV
p628-kvm1          ok              -      2      0      0      0      0      0
p628-kvm2.cecc.iho ok              -      2      0      0      0      0      0
p628-kvm3.cecc.iho ok              -      2      0      0      0      0      0


[cecuser@p628-kvm1 ~]$ bqueues
QUEUE_NAME      PRIO STATUS          MAX JL/U JL/P JL/H NJOBS  PEND   RUN  SUSP
admin            50  Open:Active       -    -    -    -     0     0     0     0
owners           43  Open:Active       -    -    -    -     0     0     0     0
priority         43  Open:Active       -    -    -    -     0     0     0     0
night            40  Open:Active       -    -    -    -     0     0     0     0
short            35  Open:Active       -    -    -    -     0     0     0     0
dataq            33  Open:Active       -    -    -    -     0     0     0     0
normal           30  Open:Active       -    -    -    -     0     0     0     0
interactive      30  Open:Active       -    -    -    -     0     0     0     0
idle             20  Open:Active       -    -    -    -     0     0     0     0

2020년 3월 23일 월요일

IBM POWER9 (ppc64le) 아키텍처에서의 SimpleITK 설치


SimpleITK는 Insight Segmentation과 Registration Toolkit (ITK)를 감싼 일종의 layer 또는 wrapper 소프트웨어입니다.   이를 IBM POWER9 (ppc64le) 아키텍처의 python3 환경에서 사용하시려면 그냥 pip install로 설치하시면 됩니다.

아래 내용을 통해 build 한 RHEL 7.6 Alt (ppc64le) 상에서의 python 3.6.9를 위한 SimpleITK 1.2.0의 wheel file을 아래 Google drive에 올려 놓았으니 그걸 download 받아서 다음과 같은 명령으로 설치하셔도 됩니다.

$ pip install ./SimpleITK-1.2.0-cp36-cp36m-linux_ppc64le.whl

구글 drive에서 download 받으려면 여기를 click


직접 pip 명령으로 SimpleITK을 설치하기 위해서는 아래와 같은 수순을 밟으면 됩니다.  제가 해보니 생각보다는 build하는데 CPU 사용량 및 사용 시간이 꽤 깁니다. 

먼저 gcc 및 make 등과 같은 Redhat OS의 기본 개발 tool을 설치합니다.

(base) [u0017496@vm ~]$ sudo yum groupinstall -y "Development Tools"

SimpleITK를 build하기 위해서는 scikit-build와 cmake (version 3.x)가 필요하므로 그것들도 설치합니다.

(base) [u0017496@vm ~]$ pip install scikit-build

(base) [u0017496@vm ~]$ conda install cmake

(base) [u0017496@vm ~]$ which cmake
~/anaconda3/bin/cmake

(base) [u0017496@vm ~]$ cmake --version
cmake version 3.14.0

그 다음은 그냥 pip install 명령을 사용하시면 됩니다.  그러면 internet repository에서 source를 가져와 build합니다.   저는 가상 CPU 환경에서 수행했는데, 한 30분은 걸린 것 같습니다.

(base) [u0017496@vm ~]$ pip install SimpleITK
Collecting SimpleITK
  Downloading https://files.pythonhosted.org/packages/11/f5/dfc5fe1ee82baa0bf35579ab49f0b0d318ae528a7557552579a587f9d7a3/SimpleITK-1.2.0.tar.gz (2.0MB)
     |████████████████████████████████| 2.0MB 9.5MB/s
Building wheels for collected packages: SimpleITK
...
  Created wheel for SimpleITK: filename=SimpleITK-1.2.0-cp36-cp36m-linux_ppc64le.whl size=42515246 sha256=063c44e98540f8f01ce965647b198294ec81782053b9eef960cc11f6229041eb
  Stored in directory: /home/cecuser/.cache/pip/wheels/b7/4e/7a/b7ac870691673ebd2688e1492d2ffac7b2380b6e607625baeb
Successfully built SimpleITK
Installing collected packages: SimpleITK
Successfully installed SimpleITK-1.2.0

Build되자마자 자동으로 설치까지 되며, 이떄 build된 wheel file은 아래 위치에 존재합니다.  그 크기는 45MB 정도 됩니다.

(wmlce_env3) [cecuser@vm ~]$ ls -l /home/cecuser/.cache/pip/wheels/b7/4e/7a/b7ac870691673ebd2688e1492d2ffac7b2380b6e607625baeb
total 41520
-rw-rw-r-- 1 cecuser cecuser 42515246 Mar 23 03:00 SimpleITK-1.2.0-cp36-cp36m-linux_ppc64le.whl

(wmlce_env3) [cecuser@vm ~]$ pip list | grep -i simpleitk
SimpleITK              1.2.0

(wmlce_env3) [cecuser@vm ~]$ which python
~/anaconda3/envs/wmlce_env3/bin/python

다음과 같이 import해보면 잘 되는 것을 확인할 수 있습니다.

(wmlce_env3) [cecuser@vm ~]$ python
Python 3.6.9 |Anaconda, Inc.| (default, Jul 30 2019, 19:18:58)
[GCC 7.3.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> import SimpleITK as sitk
>>>

2020년 2월 20일 목요일

POWER9 (RHEL 8.1, ppc64le)에서 MariaDB MaxScale을 source로부터 build하기



MaxScale은 MariaDB의 앞단에서 HA 및 query routing, CDC 등을 처리해주는 솔루션입니다.  MariaDB는 물론, mariadb-server-galera도 Red Hat Software Collections (RHSCL)에 포함되어 있습니다만, 현재로서는 MaxScale은 포함되어 있지 않습니다.

하지만 MaxScale은 open source이므로 source로부터 쉽게 build할 수 있습니다.

[cecuser@p606-kvm1 ~]$ cat /etc/redhat-release
Red Hat Enterprise Linux release 8.1 (Ootpa)

먼저 일반적인 개발 환경을 위해 필요한 package들을 설치합니다.

[cecuser@p606-kvm1 ~]$ sudo yum  groupinstall -y "Development Tools"

다음으로는 아래 OS package들을 설치하고, Rambbit-MQ와 Jansson 등의 open source SW를 source로부터 build합니다.  원래 이 과정들은 MaxScale source code package 중에서 포함된 install_build_deps.sh를 수행하면 자동으로 되는 것입니다만, ppc64le에서는 일부 수행 오류가 나는 것이 있어서 다음과 같이 수동으로 수행하면 됩니다.

[cecuser@p606-kvm1 ~]$ sudo yum install -y libtool openssl-devel libaio libaio-devel libedit systemtap-sdt-devel rpm-sign wget gnupg pcre-devel flex rpmdevtools git wget tcl tcl-devel openssl libuuid-devel xz-devel sqlite sqlite-devel pkgconfig lua lua-libs rpm-build createrepo yum-utils gnutls-devel libgcrypt-devel pam-devel libcurl-devel nodejs-devel


[cecuser@p606-kvm1 ~]$ git clone https://github.com/alanxz/rabbitmq-c.git

[cecuser@p606-kvm1 ~]$ cd rabbitmq-c

[cecuser@p606-kvm1 rabbitmq-c]$ git checkout v0.7.1

[cecuser@p606-kvm1 rabbitmq-c]$ sudo make install

[cecuser@p606-kvm1 rabbitmq-c]$ cd ..


[cecuser@p606-kvm1 ~]$ git clone https://github.com/akheron/jansson.git

[cecuser@p606-kvm1 ~]$ cd jansson

[cecuser@p606-kvm1 jansson]$ git checkout v2.9

[cecuser@p606-kvm1 jansson]$ mkdir build && cd build

[cecuser@p606-kvm1 build]$ cmake .. -DCMAKE_INSTALL_PREFIX=/usr -DCMAKE_C_FLAGS=-fPIC -DJANSSON_INSTALL_LIB_DIR=/usr/lib64

[cecuser@p606-kvm1 build]$ make

[cecuser@p606-kvm1 build]$ sudo make install

[cecuser@p606-kvm1 build]$ cd ../..


[cecuser@p606-kvm1 ~]$ wget -q -r -l1 -nH --cut-dirs=2 --no-parent -A.tar.gz --no-directories https://downloads.apache.org/avro/stable/c/

[cecuser@p606-kvm1 ~]$ tar -zxf avro-c-1.9.2.tar.gz

[cecuser@p606-kvm1 ~]$ cd avro-c-1.9.2

[cecuser@p606-kvm1 avro-c-1.9.2]$ mkdir build && cd build

[cecuser@p606-kvm1 build]$ cmake .. -DCMAKE_INSTALL_PREFIX=/usr -DCMAKE_C_FLAGS=-fPIC -DCMAKE_CXX_FLAGS=-fPIC

[cecuser@p606-kvm1 build]$ make && sudo make install

[cecuser@p606-kvm1 build]$ cd ../..


이제 MaxScale의 source를 download 받습니다.

[cecuser@p606-kvm1 ~]$ git clone https://github.com/mariadb-corporation/MaxScale

[cecuser@p606-kvm1 ~]$ cd MaxScale/

[cecuser@p606-kvm1 MaxScale]$ mkdir build && cd build

원래 manual에는 아래와 같이 install_build_deps.sh를 수행하라고 되어 있지만, 하지 마십시요.  위에서 이미 수동으로 다 처리했으며, 이걸 수행하면  x86용 binary를 download 받아서 멀쩡한 nodejs 관련 파일을 망쳐놓는 오작동을 합니다.

[cecuser@p606-kvm1 build]$ ../BUILD/install_build_deps.sh    --> Don't Run !


cmake를 수행합니다.

[cecuser@p606-kvm1 build]$ cmake .. -DCMAKE_INSTALL_PREFIX=/usr


그리고 다음 파일에서 ppc64만 있고 ppc64le가 없어서 발생하는 error가 있으므로, 아래와 같이 수정합니다.

[cecuser@p606-kvm1 build]$ vi ../query_classifier/qc_sqlite/sqlite-src-3110100/config.guess
...
    ppc64:Linux:*:*)
        echo powerpc64-unknown-linux-${LIBC}
        exit ;;
    ppc64le:Linux:*:*)   # 추가
        echo powerpc64le-unknown-linux-${LIBC}  # 추가
        exit ;;   # 추가
...

다음은 make를 수행하면 됩니다.

[cecuser@p606-kvm1 build]$ make && sudo make install


테스트를 해보면 모두 정상 수행되는 것을 보실 수 있습니다.

[cecuser@p606-kvm1 build]$ make test
Running tests...
Test project /home/cecuser/MaxScale/build
      Start  1: test_mxb_log
 1/62 Test  #1: test_mxb_log ........................   Passed    0.01 sec
      Start  2: test_semaphore
 2/62 Test  #2: test_semaphore ......................   Passed   12.01 sec
      Start  3: test_worker
...
61/62 Test #61: test_hintparser .....................   Passed    0.01 sec
      Start 62: test_masking_rules
62/62 Test #62: test_masking_rules ..................   Passed    0.01 sec

100% tests passed, 0 tests failed out of 62

Total Test time (real) = 108.24 sec



2020년 2월 18일 화요일

POWER9에서 sysbench를 source로부터 build하기



sysbench는 주로 DBMS의 성능 benchmark test를 할 때 사용되는 tool입니다.  IBM POWER9 즉 ppc64le 아키텍처의 Redhat에서 이를 build하는 방법은 간단합니다.

먼저 필요한 OS package들을 설치합니다.

(base) [cecuser@p663-kvm1 sysbench]$ sudo yum -y install make automake libtool pkgconfig libaio-devel

MariaDB 그리고 PostgreSQL과 연계 테스트를 위해서는 아래와 같은 OS package들도 함께 설치합니다.

(base) [cecuser@p663-kvm1 sysbench]$ sudo yum -y install mariadb-devel openssl-devel postgresql-devel

이제 source code를 download 받습니다.

(base) [cecuser@p663-kvm1 ~]$ git clone https://github.com/akopytov/sysbench.git

(base) [cecuser@p663-kvm1 ~]$ cd sysbench

여기서 travis_ppc64le branch로 checkout 합니다.  이걸 하지 않으면 "error: ‘GG_State’ {aka ‘struct GG_State’} has no member named ‘J’ "라는 error를 겪게 되는데, 이에 대해서는 https://github.com/akopytov/sysbench/pull/234 를 참조하십시요.

(base) [cecuser@p663-kvm1 sysbench]$ git checkout travis_ppc64le

다음으로 autogen,sh을 수행하여 configure script를 생성합니다.

(base) [cecuser@p663-kvm1 sysbench]$ ./autogen.sh

만약 postgresql이나 mariadb로 sysbench 테스트를 하실 거라면 아래와 같이 '--with-pgsql --with-mysql' 옵션과 함께 configure를 돌리시면 됩니다.  Default로는 mysql을 찾습니다.

(base) [cecuser@p663-kvm1 sysbench]$ ./configure --with-pgsql --with-mysql

만약 mysql이나 postgresql을 쓸 것이 아니라면 다음과 같이 하면 됩니다.

(base) [cecuser@p663-kvm1 sysbench]$ ./configure --without-mysql

그 다음으로는 make, sudo make install을 수행하면 됩니다.

(base) [cecuser@p663-kvm1 sysbench]$ make -j4

(base) [cecuser@p663-kvm1 sysbench]$ sudo make install

(base) [cecuser@p663-kvm1 sysbench]$ cd ..

sysbench는 아래 위치에 설치됩니다.

(base) [cecuser@p663-kvm1 ~]$ ls -l `which sysbench`
-rwxr-xr-x 1 root root 1384488 Feb 18 08:00 /usr/local/bin/sysbench

--without-mysql로 build된 sysbench 파일을 편의를 위해 아래의 Google drive에 올려놓았습니다.

https://drive.google.com/open?id=1tH9bbgQaipoAqxWFHAHVL3F4QPFSlcdG

혹시 몰라, 아래와 같이 위에서 "make -j4"까지 해놓은 sysbench directory 전체를 tgz로 묶어서 아래의 Google drive에 올려놓았습니다.  여기서는 --without-mysql로 build된 버전을 올렸습니다.

https://drive.google.com/open?id=1ircTWDzOKuZzglvz5cn-Wr2vEg0q0i3a

새로 build를 해야 하는 경우, 이 file을 아래와 같이 푸시고 sudo make install 만 수행하시면 됩니다. 

(base) [cecuser@p628-kvm1 ~]$ tar -zxf sysbench_ppc64le.tgz

(base) [cecuser@p628-kvm1 ~]$ cd sysbench

(base) [cecuser@p628-kvm1 sysbench]$ sudo make install

또는 postgresql 등의 옵션을 줘서 다시 build해야 한다면 맨 첫줄의 autoconf.sh부터 새로 시작하시면 됩니다.

2020년 2월 13일 목요일

IBM AC922 서버에서 CUDA-enabled HPL 수행하기


HPL (High Performance Linpack) 테스트는 수퍼컴 클러스터의 성능 측정에 널리 쓰이는 오픈소스 프로그램입니다.   GPU를 사용하여 HPL을 수행하기 위해서는 CUDA-enabled HPL이 필요한데, 그건 NVIDIA가 지적 재산권을 가진 프로그램이며 그건 오픈소스가 아닙니다.  WWW 상을 뒤져보면 CUDA-enabled HPL의 source code를 NVIDIA가 공개하기는 하는데, 그건 매우 오래된 GPU architecture인 Fermi 아키텍처의 GPU에 대한 것이라서 최신 GPU의 성능 측정에는 적절하지 않습니다.

아래에서는 NVIDIA의 협조를 받아 CUDA-enabled HPL의 executible binary file을 가지고 있다는 전제 하에 IBM AC922 서버 (POWER9 * 2, V100 SXM2 32GB GPU * 4) 1대로 CUDA-enabled HPL을 수행하는 과정만 제시합니다.

그 결과는 역시 confidential 정보라 공개하지 못하는 점 양해 부탁드립니다.

이 테스트 수행을 위해서는 서버에 먼저 CUDA 10.1이 설치되어 있어야 합니다.  또한 IBM의 XL Fortran, Spectrum MPI, ESSL 등의 library 등이 필요합니다.

[cecuser@p1235-met1 HPC]$ ls
ESSL_FOR_LINUX_ON_POWER_V6.2.0.tar.gz
hpl_cuda10.1.4gpus.tgz
IBM_SMPI_10.2_IP_GR_LINUX_PPC64LE.tgz
ibm_smpi_lic_s-10.02-p9-ppc64le.rpm
lsf10.1_lnx310-lib217-ppc64le.tar.Z
XL_FORTRAN_FOR_LINUX_V16.1.1_PRO.gz

먼저 XL Fortran을 설치합니다.

[cecuser@p1235-met1 HPC]$ mkdir xlf

[cecuser@p1235-met1 HPC]$ cd xlf

[cecuser@p1235-met1 xlf]$ tar -zxvf ../XL_FORTRAN_FOR_LINUX_V16.1.1_PRO.gz

[cecuser@p1235-met1 xlf]$ ./install
...
Press Enter to continue viewing the license agreement, or, Enter "1" to accept the agreement,
"2" to decline it or "99" to go back to the previous screen, "3" Print.
1
INFORMATIONAL: Unexpected CUDA Toolkit version detected '10.1' (9.2, 10.0 are supported), defaulting to __CUDA_API_VERSION=10000.  Re-configure with '-cudaVersion 9.2' to override.
Installation and configuration successful


이어서 Spectrum MPI를 설치합니다.  10.1이 아니라 10.2가 필요합니다.

[cecuser@p1235-met1 HPC]$ tar -zxf IBM_SMPI_10.2_IP_GR_LINUX_PPC64LE.tgz

[cecuser@p1235-met1 HPC]$ cd ibm_smpi-10.02.00.03-p9-ppc64le

[cecuser@p1235-met1 ibm_smpi-10.02.00.03-p9-ppc64le]$ sudo rpm -Uvh *.rpm ../ibm_smpi_lic_s-10.02-p9-ppc64le.rpm

[cecuser@p1235-met1 ibm_smpi-10.02.00.03-p9-ppc64le]$ su -

[root@p1235-met1 ~]# IBM_SPECTRUM_MPI_LICENSE_ACCEPT=yes /opt/ibm/spectrum_mpi/lap_se/bin/accept_spectrum_mpi_license.sh

[root@p1235-met1 ~]# exit

이어서 ESSL을 설치합니다.  이건 engineering용 library인데, GPU를 이용하도록 되어 있습니다.

[cecuser@p1235-met1 HPC]$ tar -zxvf ESSL_FOR_LINUX_ON_POWER_V6.2.0.tar.gz

[cecuser@p1235-met1 HPC]$ cd RHEL/RHEL7/

[cecuser@p1235-met1 RHEL7]$ su
Password:

[root@p1235-met1 RHEL7]# rpm -Uvh essl.license-6.2.0-0.ppc64le.rpm

[root@p1235-met1 RHEL7]# export IBM_ESSL_LICENSE_ACCEPT=yes

[root@p1235-met1 RHEL7]# /opt/ibmmath/essl/6.2/lap/accept_essl_license.sh

[root@p1235-met1 RHEL7]# rpm -Uvh essl.3264.rte-6.2.0-0.ppc64le.rpm essl.6464.rte-6.2.0-0.ppc64le.rpm essl.rte.common-6.2.0-0.ppc64le.rpm essl.man-6.2.0-0.ppc64le.rpm essl.3264.rtecuda-6.2.0-0.ppc64le.rpm essl.common-6.2.0-0.ppc64le.rpm essl.msg-6.2.0-0.ppc64le.rpm essl.rte-6.2.0-0.ppc64le.rpm


이제 CUDA-enabled HPL의 binary 및 script를 풀어냅니다.

[cecuser@p1235-met1 HPC]$ tar -zxvf hpl_cuda10.1.4gpus.tgz

[cecuser@p1235-met1 HPC]$ cd hpl

이 속에 들어있는 것은 간단합니다.  Binary 실행 파일인 xhpl과 함께, 그 수행에 필요한 HPL.dat, 기타 mpirun을 위한 script 입니다.

먼저 HPL.dat의 내용입니다.  Edit해야 하는 주요 내용은 아래 붉은 색으로 표시한 Ns (계산해야 하는 문제의 크기), NBs (한번에 어느 정도 크기로 문제를 풀 것인지 결정하는 block size), 그리고 mesh 구조를 결정하는 Ps와 Qs입니다. 

간단히 말하면 Ns는 가급적 GPU들의 메모리를 꽤 가득 채울 정도로 크게 하고, NBs는 적절한 크기를 trial & error 방식으로 찾아야 합니다.  가령 제가 해보니 아래와 같은 크기의 Ns면 32GB memory의 GPU 4장을 가득 채웁니다.  또한 256이나 768에 비해 512로 NBs를 두는 것이 가장 성능이 잘 나오는 것 같습니다.  Ps와 Qs는 서로 곱해서 GPU 갯수가 나오면 되는데, 가급적 서로 비슷하게, 그리고 가급적 Ps가 Qs보다 작게 설정하면 됩니다.

[cecuser@p1235-met1 hpl]$ cat HPL.dat
HPLinpack benchmark input file
Innovative Computing Laboratory, University of Tennessee
HPL.out      output file name (if any)
6            device out (6=stdout,7=stderr,file)
1            # of problems sizes (N)
128000       Ns
1          # of NBs
512        NBs
0            PMAP process mapping (0=Row-,1=Column-major)
1            # of process grids (P x Q)
2            Ps
2            Qs
16.0         threshold
1            # of panel fact
2            PFACTs (0=left, 1=Crout, 2=Right)
1            # of recursive stopping criterium
4            NBMINs (>= 1)
1            # of panels in recursion
2            NDIVs
1            # of recursive panel fact.
0            RFACTs (0=left, 1=Crout, 2=Right)
1            # of broadcast
3            BCASTs (0=1rg,1=1rM,2=2rg,3=2rM,4=Lng,5=LnM)
1            # of lookahead depth
0            DEPTHs (>=0)
1            SWAP (0=bin-exch,1=long,2=mix)
192          swapping threshold
1            L1 in (0=transposed,1=no-transposed) form
0            U  in (0=transposed,1=no-transposed) form
0            Equilibration (0=no,1=yes)
8            memory alignment in double (> 0)

여기서는 1대로 수행하니까 hosts 파일은 사실 필요가 없습니다만 아래와 같은 format으로 설정하면 됩니다.

[cecuser@p1235-met1 hpl]$ cat hosts
localhost  slots=4

아래는 mpirun을 수행하는 script입니다.  제가 쓴 환경처럼 infiniband가 없는 경우 "-pami_noib" 옵션을 써야 합니다.

[cecuser@p1235-met1 hpl]$ cat run_me_4_gpu_xlc_spectrum.sh
#!/bin/bash
export MPI_ROOT=/opt/ibm/spectrum_mpi
export MANPATH=$MPI_ROOT/share/man:$MANPATH
export PATH=/usr/local/cuda-10.1/bin:/opt/ibm/spectrum_mpi/bin:$PATH
export LD_LIBRARY_PATH=/opt/ibmmath/essl/6.2/lib64/:/opt/ibm/spectrum_mpi/lib:$LD_LIBRARY_PATH
sudo nvidia-smi -ac 877,1395
#echo always > /sys/kernel/mm/transparent_hugepage/enabled
TUNE="-x PAMI_IBV_DEVICE_NAME=mlx5_0:1 -x PAMI_IBV_DEVICE_NAME_1=mlx5_3:1 -x PAMI_ENABLE_STRIPING=0 -x PAMI_IBV_CQEDEPTH=4096 -x PAMI_IBV_ADAPTER_AFFINITY=1 -x PAMI_IBV_OPT_LATENCY=1 -x MLX5_SINGLE_THREADED=1 -x MLX5_CQE_SIZE=128 -x PAMI_IBV_ENABLE_DCT=1 -x PAMI_IBV_ENABLE_OOO_AR=1 -x PAMI_IBV_QP_SERVICE_LEVEL=8"
sudo ppc64_cpu --dscr=7
#mpirun -N 4 -npernode 4 --allow-run-as-root -x OMPI_MCA_common_pami_use_odp=0 -x PAMI_IBV_DEBUG_PRINT_DEVICES=1 -tag-output $TUNE --hostfile nodes -bind-to none ./run_linpack_6_gpu_xlc_spectrum_0726
mpirun -N 4 -npernode 4 --hostfile hosts -pami_noib -bind-to none ./run_linpack_4_gpu_xlc_spectrum.sh


그리고 아래가 실제 xhpl을 수행하는 script입니다.  위의 mpirun script를 수행하면 결국 아래의 script가 수행됩니다.  IBM의 Spectrum MPI에서는 내부적으로 OMPI_COMM_WORLD_LOCAL_RANK, PMIX 등의 환경 변수를 자동 생성하여 GPU를 할당하는데 사용합니다.  아래 script를 보면 case 문을 이용하여 CUDA_VISIBLE_DEVICES 환경 변수를 이용하여 GPU 1개씩마다 xhpl을 하나씩 수행합니다.


[cecuser@p1235-met1 hpl]$ cat run_linpack_4_gpu_xlc_spectrum.sh
#!/bin/bash
#location of HPL
HPL_DIR=`pwd`
# Number of CPU cores
# Total CPU cores / Total GPUs (not counting hyperthreading)
#CPU_CORES_PER_RANK=16
CPU_CORES_PER_RANK=8
export MPI_ROOT=/opt/ibm/spectrum_mpi
export OMP_NUM_THREADS=$CPU_CORES_PER_RANK
export MAX_H2D_MS=10
export MAX_D2H_MS=10
export RANKS_PER_SOCKET=2
export RANKS_PER_NODE=4
export NUM_WORK_BUF=4
export SCHUNK_SIZE=128
export GRID_STRIPE=4
export FACT_GEMM=1
export FACT_GEMM_MIN=128
export SORT_RANKS=0
export PRINT_SCALE=1.0
export TEST_SYSTEM_PARAMS=1
sudo rm -rf /dev/shm/sh_*
export LIBC_FATAL_STDERR_=1
#export PAMI_ENABLE_STRIPING=0
export CUDA_CACHE_PATH=/tmp
export OMP_NUM_THREADS=$CPU_CORES_PER_RANK
export CUDA_DEVICE_MAX_CONNECTIONS=8
export CUDA_COPY_SPLIT_THRESHOLD_MB=1
export GPU_DGEMM_SPLIT=1.0
export TRSM_CUTOFF=1000000
#export TRSM_CUTOFF=99000
export TEST_SYSTEM_PARAMS=1
export MONITOR_GPU=1
export GPU_TEMP_WARNING=70
export GPU_CLOCK_WARNING=1310
export GPU_POWER_WARNING=350
export GPU_PCIE_GEN_WARNING=3
export GPU_PCIE_WIDTH_WARNING=2
#export ICHUNK_SIZE=1536
export ICHUNK_SIZE=384
export CHUNK_SIZE=5120
APP=$HPL_DIR/xhpl
#lrank=$OMPI_COMM_WORLD_LOCAL_RANK
lrank=$(($PMIX_RANK%4))
nrank=$(($PMIX_RANK/4))
#crank=$(($nrank/89))
#neven=$(($crank%2))
neven=$(($nrank%2))
#neven=0
export CUDA_VISIBLE_DEVICES=$lrank
echo "RANK $PMIX_RANK on host $HOSTNAME PID $$ even: $neven"
if [ $neven -eq 0 ]
then
case ${lrank} in
[0])
#ldd $APP
sudo nvidia-smi -ac 877,1395 > /dev/null;
#export PAMI_IBV_DEVICE_NAME=mlx5_0:1;
#export OMPI_MCA_btl_openib_if_include=mlx5_0:1;
export CUDA_VISIBLE_DEVICES=0; numactl --physcpubind=0,4,8,12,16,20,24,28,32,36 --membind=0 $APP
  ;;
[1])
#export PAMI_IBV_DEVICE_NAME=mlx5_1:1;
#export OMPI_MCA_btl_openib_if_include=mlx5_1:1;
export CUDA_VISIBLE_DEVICES=1; numactl --physcpubind=40,44,48,52,56,60,64,68,72,76 --membind=0 $APP
  ;;
[2])
#export PAMI_IBV_DEVICE_NAME=mlx5_0:1;
#export OMPI_MCA_btl_openib_if_include=mlx5_0:1;
export CUDA_VISIBLE_DEVICES=2; numactl --physcpubind=80,84,88,92,96,100,104,108,112,116 --membind=8 $APP
  ;;
[3])
#export PAMI_IBV_DEVICE_NAME=mlx5_3:1;
#export OMPI_MCA_btl_openib_if_include=mlx5_3:1;
export CUDA_VISIBLE_DEVICES=3; numactl --physcpubind=120,124,128,132,136,140,144,148,152,156 --membind=8 $APP
  ;;
esac
exit
fi


이제 다음과 같이 run_me_4_gpu_xlc_spectrum.sh를 수행하시면 됩니다.  대략 10분 이내의 시간이 걸릴 것입니다.

[cecuser@p1235-met1 hpl]$ ./run_me_4_gpu_xlc_spectrum.sh


중간값을 빼면 결과적으로는 아래와 같은 결과물이 display 됩니다.   결과는 공개하지 못하는 점 다시 한번 양해 부탁드립니다.

...

================================================================================
T/V                N    NB     P     Q               Time                 Gflops
--------------------------------------------------------------------------------
WR03L2R4      128000   512     2     2             XXX              X.XXXe+04
--------------------------------------------------------------------------------
||Ax-b||_oo/(eps*(||A||_oo*||x||_oo+||b||_oo)*N)=        0.0005540 ...... PASSED
================================================================================