· CLUSTER BATCH SUBMISSION SLURM
§ Batch
o SLURM PARTITIONS & QUALITY OF SERVICE
§ Memory
§ Hybrid MPI+OpenMP job script
SLURM Workload Manager (or simply SLURM, which stands for “Simple Linux Utility for Resource Management”) is an open source and highly scalable job scheduling system.
SLURM has three key functions. Firstly, it allocates exclusive and/or non-exclusive access to resources (compute nodes) to users for some duration of time, so they can perform their work. Secondly, it provides a framework for starting, executing, and monitoring work (normally a parallel job) on the set of allocated nodes. Finally, it arbitrates contention for resources by managing the queue of pending jobs.
Currently, SLURM is the scheduling system of CRESCO7 and XCRESCO. A comprehensive documentation is on this portal, as well as on the original SchedMD site.
In order to monitor the status of computing nodes of the cluster, use the following command:
<iannone@xcrescox001 ~> sinfo -N -o "%N %P %t %C %E %m %d %r"
The output of XCRESCO cluster
HOSTNAMES PARTITION STATE CPUS(A/I/O/T) CPU_LOAD MEMORY TMP_DISK REASON
xcrescox003 xcresco* idle 0/128/0/128 none 257572 0 no
xcrescox004 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox005 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox006 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox007 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox008 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox009 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox010 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox011 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox012 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox013 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox014 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox015 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox016 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox017 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox018 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox019 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox020 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox021 xcresco* idle 0/128/0/128 none 257572 0 no
xcrescox022 xcresco* idle 0/128/0/128 none 257572 0 no
xcrescox023 xcresco* idle 0/128/0/128 none 257572 0 no
xcrescox024 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox025 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox026 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox027 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox028 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox029 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox030 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox031 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox032 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox033 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox034 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox035 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox036 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox037 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox038 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox039 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox040 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox041 xcresco* idle 0/128/0/128 none 257572 0 no
xcrescox042 xcresco* idle 0/128/0/128 none 257572 0 no
xcrescox043 xcresco* idle 0/128/0/128 none 257572 0 no
xcrescox044 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox045 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox046 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox047 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox048 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox049 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox050 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox051 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox052 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox053 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox054 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox055 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox056 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox057 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox058 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox059 xcresco* down* 0/0/128/128 Node unexpectedly rebooted 257572 0 no
xcrescox060 xcresco* idle 0/128/0/128 none 257572 0 no
where:
HOSTNAMES: short and full hostname
PARTITION: name of the Clustsrs partition: cresco7
and xcresco
STATE: node state as the following table state
CPUS(A/I/O/T): Allocated core/Idle core/Other core/Total core
CPU LOAD: CPU load of a node
MEMORY: Size of real memory in megabytes on these nodes
TMP_DISK:Size of temporary disk space in megabytes on these
nodes
REASON: The reason a node is unavailable
|
state |
description |
|
ALLOCATED |
The node has been allocated to one or more jobs |
|
ALLOCATED+ |
The node is allocated to one or more active jobs plus one or more jobs are in the process of COMPLETING |
|
COMPLETING |
All jobs associated with this node are in the process of COMPLETING |
|
DOWN |
The node is unavailable for use |
|
DRAINED |
The node is unavailable for use per system administrator request |
|
DRAINING |
The node is currently executing a job, but will not be allocated to additional jobs |
|
ERROR |
The node is currently in an error state and not capable of running any jobs |
|
FAIL |
The node is expected to fail soon and is unavailable for use per system administrator request |
|
FAILING |
The node is currently executing a job, but is expected to fail soon and is unavailable for use per system administrator request |
|
FUTURE |
The node is currently not fully configured, but expected to be available at some point in the indefinite future for use |
|
IDLE |
The node is not allocated to any jobs and is available for use |
|
MAINT |
The node is currently in a reservation with a flag value of “maintainence” |
|
REBOOT |
The node is currently scheduled to be rebooted |
|
MIXED |
The node has some of its CPUs ALLOCATED while others are IDLE |
|
PERFCTRS (NPC) |
Network Performance Counters associated with this node are in use, rendering this node as not usable for any other jobs |
|
POWER_DOWN |
The node is currently powered down and not capable of running any jobs |
|
POWER_UP |
The node is currently in the process of being powered up |
|
RESERVED |
The node is in an advanced reservation and not generally available |
|
UNKNOWN |
The Slurm controller has just started and the node's state has not yet been determined |
With SLURM you can specify the tasks that you want to be executed; the system takes care of running these tasks and returns the results to the user. If the resources are full, then SLURM holds your jobs and runs them when they will become available.
SLURM has two different modes to allocate resources:
Starting an interactive session using the srun command. For example, to start an interactive session using one node with cpus/cores fully allocated, execute the following command on a login node (e.g. xcrescox001) in the user home directory (e.g. ~/):
<iannone@xcrescox001 ~>srun --nodes=1 --ntasks-per-node=2 --cpus-per-task=128 -p xcresco --pty /bin/bash
In a xterm session a X11 interactive session can be run with the option –x11 and time long of 10 hours as follow:
<iannone@xcrescox001 ~>echo $DISPLAY:1101<iannone@xcrescox001 ~>export DISPLAY=localhost:1101.0<iannoned@xcrescox001 ~>srun --x11 --nodes=1 -t 10:00:00 --ntasks-per-node=2 --cpus-per-task=128 -p xcresco --pty /bin/bash
A parallel program can be executed interactively in the same way. For example to run a MPI/OpenMP program on 2 nodes of XCRESCO, with 1 MPI task per node and 128 threads per node is as follow:
<iannone@xcrescox001 ~>export OMP_NUM_THREADS=128<iannone@xcrescox001 ~>srun --mpi=pmi2 --nodes=1 --ntasks-per-node=1 --cpus-per-task=48 -p xcresco --output=myjob.out --error=myjob.err ./myprogram
With SLURM you create a batch job
which you then submit to the scheduler. A batch job is a file (a shell script
under UNIX) containing the set of commands that you want to run. It also
contains the directives that specify the characteristics (attributes) of the
job, and the resource requirements (e.g. number of processors and CPU time)
that your job needs. Once you create your job, you can reuse it if you wish.
Or, you can modify it for subsequent runs.
For example, here is a simple SLURM job script to run a user's application by
setting a limit (one hour) to the maximum wall clock time, requesting 1 full
node with 48 cores:
#!/bin/bash#SBATCH -N1 -n128 # 128 cores/threads on 1 node#SBATCH --time=1:00:00 # time limits: 1 hour#SBATCH --error=myJob.err # standard error file#SBATCH --output=myJob.out # standard output file#SBATCH --partition=<partition_name> # partition name (for XCRESCO: xcresco)#SBATCH --qos=<qos_name> # ./my_application
The table below gives a short description of the most used Slurm commands.
|
command |
description |
|
sacct |
report job accounting information about active or completed jobs |
|
salloc |
allocate resources for a job in real time (typically used to allocate resources and spawn a shell, in which the srun command is used to launch parallel tasks) |
|
sbatch |
submit a job script for later execution (the script typically contains one or more srun commands to launch parallel tasks) |
|
scancel |
cancel a pending or running job |
|
sinfo |
reports the state of partitions and nodes managed by Slurm (it has a variety of filtering, sorting, and formatting options) |
|
squeue |
reports the state of jobs (it has a variety of filtering, sorting, and formatting options), by default, reports the running jobs in priority order followed by the pending jobs in priority order |
|
srun |
used to submit a job for execution in an interactive session |
|
scontrol |
report more detailed information about nodes, partitions, jobs, job steps, and configuration |
|
sacctmgr |
used to view and modify (only admin) Slurm account information. |
All Slurm commands have extensive help through their man pages e.g.
<iannone@xcrescox001 ~>man sbatch
Will show you the help pages for the sbatch command.
Slurm partitions are essentially
different queues that point to collections of nodes.
Slurm quality of services are essentially different the resource limits set on
the partitions (e.g wall time).
The maximum number of cores and the maximum walltime depend on the chosen partition as well as the quality of service.
In the following table you can find all the main features and limits imposed on the SLURM partitions. For up-to-date information, use the sinfo -d , scontrol show partition <partition_name> or sacctmgr -p show qos or sacctmgr list qos <partition>commands on the system itself.
|
partition |
qos(quality of service) |
#cores |
max walltime (default) |
max memory |
priority |
Notes |
|
xcresco |
- |
7424 (58 nodes) |
24 h |
unlimited |
- |
main partition |
|
gw |
normal |
2048 |
24 h |
246 GB per node |
40 |
max 2048 cores per user |
The max value of walltime is the DEFAULT values
squeue is the main command for monitoring the state of systems, groups of
jobs or individual jobs.
The command squeue prints the list of current jobs. The list
looks something like this:
<fiannone@xcrescox001 ~>squeue -p xcresco
JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON)
The first column gives the job ID, the second the partition (or queue) where the job was submitted, the third the name of the job (specified by the user in the submission script) and the fourth the owner of the job. The fifth is the status of the job with codes as follow:
|
state |
description |
|
R |
RUNNING |
|
PD |
PENDING |
|
CA |
CANCELLED |
|
CF |
CONFIGURING |
|
CG |
COMPLETING |
|
CD |
COMPLETED |
|
F |
FAILED |
The sixth column gives the elapsed time for each particular job. Finally, there are the number of nodes requested and the nodelist where the job is running (or the cause that it is not running).
Some other useful squeue features include:
· -u for showing the status of all the jobs of a particular user, e.g. squeue -u bob for user bob;
· -l for showing more of the available information;
· - -start to report the expected start time of pending jobs.
Read all the options for squeue on the Linux manual using the command man squeue, including how to personalize the information to be displayed.
To query detailed information about job 93262:
<iannone@xcrescox001 ~>scontrol show job 93262
to have a different output format of the squeue command:
<g2fianno@s51 ~>squeue --format="%.18i %.11P %.20q %.8j %.8u %.8T %.10M %.5C %.10m %.9l %.6D %R" -p gwJOBID PARTITION QOS NAME USER STATE TIME CPUS MIN_MEMORY TIME_LIMI NODES NODELIST(REASON)
where MIN_MEMORY is the memory for CPU.
A real-time report of the stdout of a job running can be shown with the following command:
<g2fianno@s51 ~>sattch <jobid>.0
Use the scancel command to delete a job, e.g.
<iannone@sxcrescox001 ~>scancel 1121
to delete job with ID 1121. A user can delete his/her own jobs at any time, whether the job is pending (waiting in the queue) or running. A user cannot delete the jobs of another user. Normally, there is a (small) delay between the execution of the scancel command and the time when the job is dequeued and killed. Occasionally a job may not delete properly, in which case, the ENEA support team can delete it upon request.
The SLURM job accounting records
show job details, including job steps and memory usage. The memory usage is
accurate for jobs that use srun to launch tasks.
To show job account information for a specific job:
<iannone@xcrescox001 ~>sacct -l -j 1121
To show information for your jobs currently running:
<iannone@xcrescox001 ~>sacct -l
To Show all your job information starting from a specific date:
<iannone@xcrescox001 ~>sacct -l --starttime YYYY-MM-DD
At the time a job is launched into execution, Slurm defines multiple environment variables, which can be used from within the submission script to define the correct workflow of the job. The most useful of these environment variables are the following:
|
variable |
description |
|
SLURM_JOB_ID |
Job ID |
|
SLURM_JOB_NAME |
Job Name |
|
SLURM_SUBMIT_DIR |
Submit Directory |
|
SLURM_JOB_NODELIST |
Nodes assigned to job |
|
SLURM_SUBMIT_HOST |
Host submitted from |
|
SLURM_JOB_NUM_NODES |
Number of nodes allocated to job |
|
SLURM_CPUS_ON_NODE |
Number of cores/node |
|
SLURM_NTASKS |
Total number of cores for job |
|
SLURM_NODEID |
Index to node running on relative to nodes assigned to job |
|
SLURM_PROCID |
Index to task relative to job |
This is a list of the most common flags that any user may include on scripts to request different resources and features for jobs
|
description |
JOB Specification |
|
Script Directive |
#SBATCH |
|
Job Name |
SBATCH –job-name=My-Job_Name |
|
Wall time hours (default is the max walltime of the qos) |
#SBATCH –time=24:0:0 |
|
Number of nodes requested |
#SBATCH –nodes=1 |
|
Number of tasks |
#SBATCH –ntasks=48 |
|
Number of core per node requested |
#SBATCH –ntasks-per-node=48 |
|
Number of cpus per task requested |
–cpus-per-task=48 |
|
send mail at the end of teh job |
#SBATCH –mail-type=end Event(s) that triggers email notification (BEGIN,END,FAIL,ALL) |
|
user's email address |
#SBATCH –mail-user=userid@myaddress |
|
Copy user's environment |
#SBATCH –export=[ALL|NONE|Variables] |
|
Working Directory |
#SBATCH –workdir=dir-name |
|
Job Restart |
#SBATCH –requeue |
|
Share Nodes |
#SBATCH –shared |
|
Dedicated Nodes |
#SBATCH –exclusive |
|
Memory Size |
#SBATCH –mem=[mem |M|G|T] or –mem-per-cpu |
|
Partition Name |
#SBATCH –partitions=xcresco |
|
Quality of Service |
#SBATCH –qos=<name> |
|
Job Arrays |
#SBATCH –array=[array_spec] |
Dedicated nodes can be specified with the –exclusive flag and all CPUs and memory for each node will be allocated. Programs that rely heavily on data transfer between tasks may be suited for exclusive nodes. If exclusive nodes are not needed, whether the jobs are too small for a single node or do not leverage shared memory, the –shared flag will designate that a fraction of each node may be used.
Parallel processing may be done with either multiple processes, threads, or a combination of both. A single process may have multiple threads sharing memory. Multiple processes require some form to communicate, for example MPI. In SLURM, the number of processes is controlled by setting the number of “tasks”, while threads are controlled by the number of “cpus” (see below for relevant flags).
The number of nodes can be specified using the –nodes or -N flags and takes the form of min-max. If a single number is given, the scheduler will only allocate that number of nodes. You can also specify the resources needed by giving the number of tasks with –ntasks or -n along with the number of –cpus-per-task, in which case the scheduler will decide on the appropriate number of nodes for your job. You may also specify the number –ntasks-per-node, which will multiply –cpus-per-task if both are used. Be aware that if you ask for more CPUs than are available in a single node, the scheduler will refuse your request and throw an exception. Finally, you may also request a minimum number of CPUs with the –mincpus flag.
The –mem flag specifies the total amount of memory per node. The –mem-per-cpu specifies the amount of memory per allocated CPU. The two flags are mutually exclusive.
For a typical serial job you can
take the following script: myscript.sh as a template, modifying it
depending on your needs.
The script asks for 10 minutes wallclock time and runs a serial application
(R). The input data are in file “data”, the output file is “job.out”; job.out
will contain the std-out whilst the std-err is in the file “job.err”. The
working directory is ~/ and the job is submitted on the partition xcresco.
The job run on 1 node requiring 1 cores.
#!/bin/bash #SBATCH --job-name myjobname#SBATCH --partition=xcresco #add #SBATCH --qos=<name> to use the alternative qos #SBATCH --output=job.out#SBATCH --error=job.err#SBATCH --time=0:10:00#SBATCH --nodes=1#SBATCH --ntasks=1#SBATCH --mail-type=END # Event(s) that triggers email notification (BEGIN,END,FAIL,ALL)#SBATCH --mail-user=francesco.iannone@enea.it # Destination email address cd ~/module load R R < data
In order to submit it, digit the command:
<iannone@xcrescox001 ~/> sbatch myscript.sh
For a typical OpenMP (Multi-threaded) job you can take the following script: myscript.sh as a template, modifying it depending on your needs. The script asks for 10 minutes wallclock time and runs a multi-treaded application myprogr. The output file is “job.out”; job.out will contain the std-out whilst the std-err is in the file “job.err”. The working directory is ~/ and the job is submitted on the partition xcresco. The job run 1 task with 128 threads, 1 for each core.
#!/bin/bash #SBATCH --job-name myjobname#SBATCH --partition=xcresco #add #SBATCH --qos=<name> to use the alternative qos #SBATCH --output=job.out#SBATCH --error=job.err#SBATCH --time=0:10:00#SBATCH --ntasks=1#SBATCH --cpus-per-task=128#SBATCH --mail-type=END # Event(s) that triggers email notification (BEGIN,END,FAIL,ALL)#SBATCH --mail-user=francesco.iannone@enea.it # Destination email address export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASKcd `/ ./myprogr
For a typical MPI job you can take the following script as a template, modifying it depending on your needs.
The script asks for 48 tasks on 2 nodes and 1 hour of wallclock time, and runs a MPI application (myprogr). The job runs on 72 cores of 2 nodes. The output file is “job.out”, the working directory is where the job was submitted from.
#!/bin/bash #SBATCH --job-name myjobname#SBATCH --partition=xcrescox001 #add #SBATCH --qos=<name> to use the alternative qos #SBATCH --output=job.out#SBATCH --error=job.err#SBATCH --time=1:00:00#SBATCH --nodes=2#SBATCH --ntasks-per-node=48#SBATCH --mail-type=END # Event(s) that triggers email notification (BEGIN,END,FAIL,ALL)#SBATCH --mail-user=francesco.iannone@enea.it # Destination email address mpirun -np 4 --map-by ppr:2:node ./myprogr
The script asks for 2 nodes, 2 MPI task per node and 128 OpenMP threads (one per each core), 1 hours of wallclock time. The application (myprogr) was compiled with the intel compiler and the openmpi library. The output file is “job.out”, the working directory is where the job was submitted from.
#!/bin/bash #SBATCH --job-name myjobname#SBATCH --partition=xcresco #add #SBATCH --qos=<name> to use the alternative qos#SBATCH --output=job.out#SBATCH --error=job.err#SBATCH --time=1:00:00#SBATCH --nodes=2#SBATCH --ntasks-per-node=1#SBATCH --cpus-per-task=128#SBATCH --mail-type=END # Event(s) that triggers email notification (BEGIN,END,FAIL,ALL)#SBATCH --mail-user=francesco.iannone@enea.it # Destination email address cd ~/export KMP_AFFINITY=scatterexport OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK
mpirun -np 4 --map-by ppr:2:node --use-hwthread-cpus ./myprogr
It is sometimes necessary to submit large numbers of jobs,
especially serial (single core) jobs.
When you invoke sbatch with a Job Array, you will have access to a Job index
number stored in the $SLURM_ARRAY_TASK_ID environment variable. This allows a
single job script to be used for multiple jobs. Here is a sample script to run
100 jobs on 1 node requiring 1 cores.
#!/bin/bash #SBATCH --job-name myjobname#SBATCH --partition=xcresco #add #SBATCH --qos=<name> to use the alternative qos #SBATCH --output=job.out#SBATCH --error=job.err#SBATCH --time=0:10:00#SBATCH --nodes=1#SBATCH --ntasks=1#SBATCH --array=0-99#SBATCH --mail-type=END # Event(s) that triggers email notification (BEGIN,END,FAIL,ALL)#SBATCH --mail-user=francesco.iannone@enea.it # Destination email address cd ~/module load R R < data.$SLURM_ARRAY_TASK_ID
It is possible to use Job Arrays with an exotic numbering scheme. For example:
#SBATCH --array=0-10,50
This job submission would result in 12 jobs (indexed 0 through 10 and a job with an index of 50).
· A useful summary guide for SLURM commands is here
· The Rosetta Stone of Workload Managers (SLURM/PBS/LSF..) is here