Parallel Jobs (Slurm)
Parallel batch job submission (Slurm)
The CSF has a variety of compute node hardware, which is made available via different Slurm Partitions (think of these as queues for jobs wanting to access the different hardware types.)
Most of the parallel applications installed on the CSF will use either the “OpenMP” parallel method or the “MPI” parallel method. This determines whether you can use more than one compute node in your job, and also the settings and flags you need to use in the jobscript to run the application correctly.
If your application uses some other parallel method, please first check whether it is multi-node capable or restricted to running on a single compute node.
The following types of parallel jobs are all supported – they will all use multiple cores on the compute node(s) available to the job. Click on the link to see an example jobscript below:
- Single-node job using OpenMP parallelism (method 1) – jobscript.
- Single-node job using OpenMP parallelism (method 2) – jobscript.
- Single-node job using MPI parallelism – jobscript.
- Multi-node job using MPI parallelism – jobscript.
- Multi-node job using MPI+OpenMP (aka “mixed-mode”) parallelism (example 1, possibly less optimal) – jobscript.
- Multi-node job using MPI+OpenMP (aka “mixed-mode”) parallelism (example 2, possibly more optimal) – jobscript.
Please also consult the software page for the code / application you want to use, for advice on running that application.
A parallel job script will run in the directory (folder) from which you submit the job.
Single-node OpenMP parallel Job (method 1)
OpenMP parallel applications will only run on a single-compute node. The app use the cores of that compute node to perform parallel processing.
#!/bin/bash --login #SBATCH -p multicore # Partition is required. Runs on an AMD Genoa hardware. 2-168 cores. # You can also use 'multicore_small'. Runs on Intel hardware. 2-40 cores. #SBATCH -n numcores # (or --ntasks=) Number of cores you wish to use. #SBATCH -t 4-0 # Wallclock limit (days-hours). Required! # Max permitted is 7 days (7-0). # Load any required modulefiles. A purge is used to start with a clean environment. module purge module load apps/some/example/1.2.3 ### OpenMP jobs ### # OpenMP code will use $SLURM_NTASKS cores (-n above) export OMP_NUM_THREADS=$SLURM_NTASKS omp-app.exe
Single-node OpenMP parallel Job (method 2)
Here we request 1 “task” (process) and multiple CPUs per task.
On the CSF, this is functionally equivalent to the above jobscript, which simply uses -n numcores, so you can use either. In fact, many of our jobscript examples in the software pages use only -n to specify the number of cores. But this is because jobs in the multicore partition are single-node jobs. The distinction between -n and -c is significant in the multinode partition, shown later in this page.
If you use the Slurm srun starter to run your executable inside the batch job (optional, and in most cases we don’t use this) then the distinction between -n and -c is important.
#!/bin/bash --login #SBATCH -p multicore # Partition is required. Runs on an AMD Genoa hardware. 2-168 cores. # You can also use 'multicore_small'. Runs on Intel hardware. 2-40 cores. #SBATCH -n 1 # (or --ntasks=) The default is 1 so this line can be omitted. #SBATCH -c numcores # (or --cpus-per-task) where numcores is number of cores you wish to use. #SBATCH -t 4-0 # Wallclock limit (days-hours). Required! # Max permitted is 7 days (7-0). # Load any required modulefiles. A purge is used to start with a clean environment. module purge module load apps/some/example/1.2.3 ### OpenMP jobs ### # OpenMP code will use $SLURM_CPUS_PER_TASK cores (-c above) export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK omp-app.exe
Single-node MPI parallel Job
You can run single-node parallel MPI jobs.
#!/bin/bash --login #SBATCH -p multicore # Partition is required. Runs on an AMD Genoa hardware. # You can also use 'multicore_small'. Runs on Intel hardware. 2-40 cores. #SBATCH -n numcores # (or --ntasks=) where numcores is between 2 and 168. #SBATCH -t 4-0 # Wallclock limit (days-hours). Required! # Max permitted is 7 days (7-0). # Load any required modulefiles. A purge is used to start with a clean environment. module purge module load apps/some/example/1.2.3 ### MPI jobs ### # mpirun will run $SLURM_NTASKS (-n above) MPI processes mpirun mpi-app.exe
Multi-node MPI Job
You can run multi-node parallel MPI jobs on the CSF. The multi-node hardware uses InfiniBand EDR 100Gb/s networking between the compute nodes.
Please note: Jobs in multinode must use all cores on each node. Hence they must be a multiple of 40-cores in size. The minimum job size is 80 cores (two nodes.) If you only want to use 40 cores, say, then use the multicore partition.
This is a purely MPI job (see below for a mixed-mode MPI+OpenMP parallel job).
Note that you must provide the -N numnodes and -n numcores flags.
#!/bin/bash --login #SBATCH -p multinode # Partition is required. Runs on Intel Cascadelake hardware. #SBATCH -N numnodes # (or --nodes=) where numnodes is between 2 and 32. #SBATCH -n numcores # (or --ntasks=) where numcores is between 80 and 1280 (total number of cores) #SBATCH -t 4-0 # Wallclock limit (days-hours). Required! # Max permitted is 7 days (7-0) for jobs of 80-560 cores. # 4 days (4-0) for jobs of 600-1280 cores. # Load any required modulefiles. A purge is used to start with a clean environment. module purge module load apps/some/example/1.2.3 ### MPI jobs ### # mpirun will run $SLURM_NTASKS (-n above) MPI processes mpirun mpi-app.exe
Multi-node MPI+OpenMP mixed-mode apps (example 1)
Here we run one MPI processes on each compute node in the multi-node job, and each of those MPI processes will run 40 OpenMP threads to do multi-core processing on all of the cores in the compute node. The multicore compute nodes all have 40-cores in them, and you must use all cores on the nodes.
Note that running an MPI-process per node (when a node has two or more sockets – two in the CSF compute nodes) as shown here, can be less optimal than the MPI-process per socket (CPU) shown in example 2 below. Hence we recommend you run some timing tests of your application to determine which method is best for your app.
#!/bin/bash --login #SBATCH -p multinode # Partition is required. Runs on Intel Cascadelake hardware. #SBATCH -N numnodes # (or --nodes=) where numnodes is between 2 and 32. #SBATCH -n numcores # (or --ntasks=) where numcores is same as the number of nodes! # By using the number of nodes here too, we get one MPI process running on each node. #SBATCH -c numcores # (or --cpus-per-task) Number of OpenMP threads per MPI process (must be 40 to fill each node). # Each MPI process will use all 40 cores on the compute node it is running on. # This will use numtasks x numcores cores in total #SBATCH -t 4-0 # Wallclock limit (days-hours). Required! # Max permitted is 7 days (7-0) for jobs of 80-560 cores. # 4 days (4-0) for jobs of 600-1280 cores. # Load any required modulefiles. A purge is used to start with a clean environment. module purge module load apps/some/example/1.2.3 ### OpenMP ### # Each MPI process uses OpenMP to run on $SLURM_CPUS_PER_TASK cores (-c above) export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK ### MPI ### # The MPI application is then started with $SLURM_NTASKS processes (copies of the app) # and each MPI process will use all of the cores on the node for OpenMP processing. mpirun --map-by ppr:1:node --bind-to none mix-mode-app.exe # # The "--bind-to none" is required in OpenMPI v5 to force # mpirun to allow a single MPI process to see the cores in # more than one socket (CPU).
Example using 3 compute nodes, hence 3 MPI processes which each use 40-cores for OpenMP. A total of 120 cores used by your job in total.
+-----------------------------------+ +-----------------------------------+ +-----------------------------------+ | Node_1 | | Node_2 | | Node_3 | | Socket_0 (20-cores): MPI-proc-0 | | Socket_0 (20-cores): MPI-proc-1 | | Socket_0 (20-cores): MPI-proc-2 | | Socket_1 (20-cores): " | | Socket_1 (20-cores): " | | Socket_1 (20-cores): " | +-----------------------------------+ +-----------------------------------+ +-----------------------------------+
The jobscript is:
#!/bin/bash --login #SBATCH -N 3 # 3 compute nodes #SBATCH -n 3 # 3 MPI processes in total (hence one per compute node) #SBATCH -c 40 # 40 cores per MPI process for OpenMP threads. module purge module load ...... export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK mpirun --map-by ppr:1:node --bind-to none mix-mode-app.exe
Multi-node MPI+OpenMP mixed-mode apps (example 2)
Here we run two MPI processes on each compute node in the multi-node job, and each of those MPI processes will run 20 OpenMP threads to do multi-core processing on all of the cores in each socket (aka CPU) in the compute node. The multicore compute nodes all have two sockets (CPUs) in them, and each socket has 20-cores.
Running an MPI-process per socket can be more optimal than the MPI-process per node shown in example 1 above.
#!/bin/bash --login #SBATCH -p multinode # Partition is required. Runs on Intel Cascadelake hardware. #SBATCH -N numnodes # (or --nodes=) where numnodes is between 2 and 32. #SBATCH -n numcores # (or --ntasks=) where numcores is twice the number of nodes! # By using twice the number of nodes here, we get two MPI process running on each node. #SBATCH -c numcores # (or --cpus-per-task) Number of OpenMP threads per MPI process (must be 20 to fill each socket). # Each MPI process will use all 20 cores on the socket it is running on. # This will use numtasks x numcores cores in total #SBATCH -t 4-0 # Wallclock limit (days-hours). Required! # Max permitted is 7 days (7-0) for jobs of 80-560 cores. # 4 days (4-0) for jobs of 600-1280 cores. # Load any required modulefiles. A purge is used to start with a clean environment. module purge module load apps/some/example/1.2.3 ### OpenMP ### # Each MPI process uses OpenMP to run on $SLURM_CPUS_PER_TASK cores (-c above) export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK ### MPI ### # The MPI application is then started with $SLURM_NTASKS processes (copies of the app) # and each MPI process will use all of the cores on the node for OpenMP processing. mpirun --map-by ppr:1:socket mix-mode-app.exe # # Do NOT use "--bind-to none" in this example. Instead, mpirun will automatically # bind an MPI process to the cores found in the socket it's running on.
+-----------------------------------+ +-----------------------------------+ +-----------------------------------+ | Node_1 | | Node_2 | | Node_3 | | Socket_0 (20-cores): MPI-proc-0 | | Socket_0 (20-cores): MPI-proc-2 | | Socket_0 (20-cores): MPI-proc-4 | | Socket_1 (20-cores): MPI-proc-1 | | Socket_1 (20-cores): MPI-proc-3 | | Socket_1 (20-cores): MPI-proc-5 | +-----------------------------------+ +-----------------------------------+ +-----------------------------------+
The jobscript is:
#!/bin/bash --login #SBATCH -N 3 # 3 compute nodes #SBATCH -n 6 # 6 MPI processes in total (hence two per compute node) #SBATCH -c 20 # 20 cores per MPI process for OpenMP threads. module purge module load ...... export OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK mpirun --map-by ppr:1:socket mix-mode-app.exe
Available Hardware and Resources
Please see the Partitions page for details on available compute resources.
