{"id":9233,"date":"2025-04-02T15:36:16","date_gmt":"2025-04-02T14:36:16","guid":{"rendered":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/?page_id=9233"},"modified":"2026-10-08T10:14:34","modified_gmt":"2026-10-08T09:14:34","slug":"parallel-jobs-slurm","status":"publish","type":"page","link":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/batch-slurm\/parallel-jobs-slurm\/","title":{"rendered":"Parallel Jobs (Slurm)"},"content":{"rendered":"<h2>Parallel batch job submission (Slurm)<\/h2>\n<p>The CSF has a variety of compute node hardware, which is made available via different Slurm <a href=\"\/csf3\/batch-slurm\/partitions\">Partitions<\/a> (think of these as queues for jobs wanting to access the different hardware types.) <\/p>\n<p>Most of the parallel applications installed on the CSF will use either the &#8220;OpenMP&#8221; parallel method or the &#8220;MPI&#8221; parallel method. This determines whether you can use more than one compute node in your job, and also the settings and flags you need to use in the jobscript to run the application correctly. <\/p>\n<p>If your application uses some other parallel method, please first check whether it is multi-node capable or restricted to running on a single compute node. <\/p>\n<p>The following types of parallel jobs are all supported &#8211; they will all use multiple cores on the compute node(s) available to the job. Click on the link to see an example jobscript below:<\/p>\n<ul>\n<li>Single-node job using OpenMP parallelism (method 1) &#8211; <a href=\"#omp1\">jobscript<\/a>.<\/li>\n<li>Single-node job using OpenMP parallelism (method 2) &#8211; <a href=\"#omp2\">jobscript<\/a>.<\/li>\n<li>Single-node job using MPI parallelism &#8211; <a href=\"#mpi1\">jobscript<\/a>.<\/li>\n<\/ul>\n<ul>\n<li>Multi-node job using MPI parallelism &#8211; <a href=\"#mpi2\">jobscript<\/a>.<\/li>\n<li>Multi-node job using MPI+OpenMP (aka &#8220;mixed-mode&#8221;) parallelism (example 1, possibly less optimal) &#8211; <a href=\"#mpi3\">jobscript<\/a>.<\/li>\n<li>Multi-node job using MPI+OpenMP (aka &#8220;mixed-mode&#8221;) parallelism (example 2, possibly more optimal) &#8211; <a href=\"#mpi4\">jobscript<\/a>.<\/li>\n<\/ul>\n<p><em><strong>Please also consult the <a href=\"\/csf3\/software\/applications\">software page<\/a> for the code \/ application you want to use, for advice on running that application<\/strong><\/em>.<\/p>\n<p>A parallel job script will run in the directory (folder) from which you submit the job.<\/p>\n<h3 id=\"omp1\">Single-node OpenMP parallel Job (method 1)<\/h3>\n<p>OpenMP parallel applications will only run on a single-compute node. The app use the cores of that compute node to perform parallel processing.<\/p>\n<pre>\r\n#!\/bin\/bash --login\r\n#SBATCH -p multicore  # Partition is <strong>required<\/strong>. Runs on an <strong>AMD Genoa<\/strong> hardware. 2-168 cores.\r\n                      # You can also use 'multicore_small'. Runs on <strong>Intel<\/strong> hardware. 2-40 cores.\r\n#SBATCH -n <em>numcores<\/em>   # (or --ntasks=) Number of cores you wish to use.\r\n#SBATCH -t 4-0        # Wallclock limit (days-hours). <strong>Required!<\/strong>\r\n                      # Max permitted is 7 days (7-0).\r\n\r\n# Load any required modulefiles. A purge is used to start with a clean environment.\r\nmodule purge\r\nmodule load <em>apps\/some\/example\/1.2.3<\/em>\r\n\r\n<strong>### OpenMP jobs ###<\/strong>\r\n# OpenMP code will use $SLURM_NTASKS cores (-n above)\r\nexport OMP_NUM_THREADS=$SLURM_NTASKS\r\nomp-app.exe\r\n<\/pre>\n<h3 id=\"omp2\">Single-node OpenMP parallel Job (method 2)<\/h3>\n<p>Here we request 1 &#8220;task&#8221; (process) and multiple CPUs per task. <\/p>\n<p>On the CSF, this is functionally equivalent to the above jobscript, which simply uses <code>-n <em>numcores<\/em><\/code>, so you can use either. In fact, many of our jobscript examples in the <a href=\"\/csf3\/software\/applications\">software pages<\/a> use only <code>-n<\/code> to specify the number of cores. But this is because jobs in the <code>multicore<\/code> partition are single-node jobs. The distinction between <code>-n<\/code> and <code>-c<\/code> is significant in the <code>multinode<\/code> partition, shown later in this page.<\/p>\n<p>If you use the Slurm <code>srun<\/code> starter to run your executable inside the batch job (optional, and in most cases we don&#8217;t use this) then the distinction between <code>-n<\/code> and <code>-c<\/code> is important.<\/p>\n<pre>\r\n#!\/bin\/bash --login\r\n#SBATCH -p multicore  # Partition is <strong>required<\/strong>. Runs on an <strong>AMD Genoa<\/strong> hardware. 2-168 cores.\r\n                      # You can also use 'multicore_small'. Runs on <strong>Intel<\/strong> hardware. 2-40 cores.\r\n#SBATCH -n 1          # (or --ntasks=) The default is 1 so this line can be omitted.\r\n#SBATCH -c <em>numcores<\/em>   # (or --cpus-per-task) where <em>numcores<\/em> is number of cores you wish to use.\r\n#SBATCH -t 4-0        # Wallclock limit (days-hours). <strong>Required!<\/strong>\r\n                      # Max permitted is 7 days (7-0).\r\n\r\n# Load any required modulefiles. A purge is used to start with a clean environment.\r\nmodule purge\r\nmodule load <em>apps\/some\/example\/1.2.3<\/em>\r\n\r\n<strong>### OpenMP jobs ###<\/strong>\r\n# OpenMP code will use $SLURM_CPUS_PER_TASK cores (-c above)\r\nexport OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK\r\nomp-app.exe\r\n<\/pre>\n<h3 id=\"mpi1\">Single-node MPI parallel Job<\/h3>\n<p>You can run single-node parallel MPI jobs.<\/p>\n<pre>\r\n#!\/bin\/bash --login\r\n#SBATCH -p multicore  # Partition is <strong>required<\/strong>. Runs on an <strong>AMD Genoa<\/strong> hardware.\r\n                      # You can also use 'multicore_small'. Runs on <strong>Intel<\/strong> hardware. 2-40 cores.\r\n#SBATCH -n <em>numcores<\/em>   # (or --ntasks=) where <em>numcores<\/em> is between 2 and 168.\r\n#SBATCH -t 4-0        # Wallclock limit (days-hours). <strong>Required!<\/strong>\r\n                      # Max permitted is 7 days (7-0).\r\n\r\n# Load any required modulefiles. A purge is used to start with a clean environment.\r\nmodule purge\r\nmodule load <em>apps\/some\/example\/1.2.3<\/em>\r\n\r\n<strong>### MPI jobs ###<\/strong>\r\n# mpirun will run $SLURM_NTASKS (-n above) MPI processes\r\nmpirun <em>mpi-app.exe<\/em>\r\n<\/pre>\n<h3 id=\"mpi2\">Multi-node MPI Job<\/h3>\n<p>You can run multi-node parallel MPI jobs on the CSF. The multi-node hardware uses InfiniBand EDR 100Gb\/s networking between the compute nodes.<\/p>\n<p><strong>Please note:<\/strong> Jobs in <code>multinode<\/code> <em>must<\/em> use all cores on each node. Hence they must be a multiple of 40-cores in size. The minimum job size is 80 cores (two nodes.) If you only want to use 40 cores, say, then use the <code>multicore<\/code> partition.<\/p>\n<p>This is a purely MPI job (see <a href=\"#mpi3\">below<\/a> for a <em>mixed-mode<\/em> MPI+OpenMP parallel job).<\/p>\n<p>Note that you <strong>must<\/strong> provide the <code>-N <em>numnodes<\/em><\/code> and <code>-n <em>numcores<\/em><\/code> flags.<\/p>\n<pre>\r\n#!\/bin\/bash --login\r\n#SBATCH -p multinode  # Partition is <strong>required<\/strong>. Runs on <strong>Intel Cascadelake<\/strong> hardware.\r\n#SBATCH -N <em>numnodes<\/em>   # (or --nodes=) where <em>numnodes<\/em> is between 2 and 32.\r\n#SBATCH -n <em>numcores<\/em>   # (or --ntasks=) where <em>numcores<\/em> is between 80 and 1280 (<strong>total number of cores<\/strong>)\r\n#SBATCH -t 4-0        # Wallclock limit (days-hours). <strong>Required!<\/strong>\r\n                      # Max permitted is 7 days (7-0) for jobs of 80-560 cores.\r\n                      #                  4 days (4-0) for jobs of 600-1280 cores.\r\n\r\n# Load any required modulefiles. A purge is used to start with a clean environment.\r\nmodule purge\r\nmodule load <em>apps\/some\/example\/1.2.3<\/em>\r\n\r\n<strong>### MPI jobs ###<\/strong>\r\n# mpirun will run $SLURM_NTASKS (-n above) MPI processes\r\nmpirun <em>mpi-app.exe<\/em>\r\n<\/pre>\n<h3 id=\"mpi3\">Multi-node MPI+OpenMP mixed-mode apps (example 1)<\/h3>\n<p>Here we run <strong>one MPI processes<\/strong> on <em>each<\/em> compute node in the multi-node job, and each of those MPI processes will run 40 OpenMP threads to do multi-core processing on all of the cores in the compute node. The <code>multicore<\/code> compute nodes all have 40-cores in them, and you must use all cores on the nodes.<\/p>\n<p>Note that running an MPI-process per <em>node<\/em> (when a node has two or more sockets &#8211; two in the CSF compute nodes) as shown here, can be <strong>less optimal<\/strong> than the MPI-process per <em>socket<\/em> (CPU) shown in example 2 below. Hence we recommend you run some timing tests of your application to determine which method is best for your app.<\/p>\n<pre>\r\n#!\/bin\/bash --login\r\n#SBATCH -p multinode  # Partition is <strong>required<\/strong>. Runs on <strong>Intel Cascadelake<\/strong> hardware.\r\n#SBATCH -N <em>numnodes<\/em>   # (or --nodes=) where <em>numnodes<\/em> is between 2 and 32.\r\n#SBATCH -n <em>numcores<\/em>   # (or --ntasks=) where <em>numcores<\/em> is same as the number of nodes!\r\n                      # By using the number of nodes here too, we get one MPI process running on each node.\r\n#SBATCH -c <em>numcores<\/em>   # (or --cpus-per-task) Number of OpenMP threads per MPI process (must be 40 to fill each node).\r\n                      # Each MPI process will use all 40 cores on the compute node it is running on.\r\n                      # This will use <em>numtasks<\/em> x <em>numcores<\/em> cores in total\r\n#SBATCH -t 4-0        # Wallclock limit (days-hours). <strong>Required!<\/strong>\r\n                      # Max permitted is 7 days (7-0) for jobs of 80-560 cores.\r\n                      #                  4 days (4-0) for jobs of 600-1280 cores.\r\n\r\n# Load any required modulefiles. A purge is used to start with a clean environment.\r\nmodule purge\r\nmodule load <em>apps\/some\/example\/1.2.3<\/em>\r\n\r\n<strong>### OpenMP ###<\/strong>\r\n# Each MPI process uses OpenMP to run on $SLURM_CPUS_PER_TASK cores (-c above)\r\nexport OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK\r\n\r\n<strong>### MPI ###<\/strong>\r\n# The MPI application is then started with $SLURM_NTASKS processes (copies of the app)\r\n# and each MPI process will use all of the cores on the node for OpenMP processing.\r\nmpirun --map-by ppr:1:<strong>node<\/strong> <strong>--bind-to none<\/strong> mix-mode-app.exe\r\n                               #\r\n                               # The \"--bind-to none\" is required in OpenMPI v5 to force\r\n                               # mpirun to allow a single MPI process to see the cores in\r\n                               # more than one <em>socket<\/em> (CPU).\r\n<\/pre>\n<p>Example using 3 compute nodes, hence 3 MPI processes which each use 40-cores for OpenMP. A total of 120 cores used by your job in total.<\/p>\n<pre>\r\n+-----------------------------------+   +-----------------------------------+   +-----------------------------------+\r\n| Node_1                            |   | Node_2                            |   | Node_3                            |\r\n|   Socket_0 (20-cores): MPI-proc-0 |   |   Socket_0 (20-cores): MPI-proc-1 |   |   Socket_0 (20-cores): MPI-proc-2 |\r\n|   Socket_1 (20-cores):     \"      |   |   Socket_1 (20-cores):     \"      |   |   Socket_1 (20-cores):     \"      |\r\n+-----------------------------------+   +-----------------------------------+   +-----------------------------------+\r\n<\/pre>\n<p>The jobscript is:<\/p>\n<pre>\r\n#!\/bin\/bash --login\r\n#SBATCH -N 3      # 3 compute nodes\r\n#SBATCH -n 3      # 3 MPI processes in total (hence one per compute node)\r\n#SBATCH -c 40     # 40 cores per MPI process for OpenMP threads.\r\nmodule purge\r\nmodule load ......\r\nexport OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK\r\nmpirun --map-by ppr:1:<strong>node<\/strong> <strong>--bind-to none<\/strong> mix-mode-app.exe\r\n<\/pre>\n<h3 id=\"mpi4\">Multi-node MPI+OpenMP mixed-mode apps (example 2)<\/h3>\n<p>Here we run <strong>two MPI processes<\/strong> on <em>each<\/em> compute node in the multi-node job, and each of those MPI processes will run 20 OpenMP threads to do multi-core processing on all of the cores in each <em>socket<\/em> (aka CPU) in the compute node. The <code>multicore<\/code> compute nodes all have two <em>sockets<\/em> (CPUs) in them, and each socket has 20-cores.<\/p>\n<p>Running an MPI-process per <em>socket<\/em> can be <strong>more optimal<\/strong> than the MPI-process per <em>node<\/em> shown in example 1 above.<\/p>\n<pre>\r\n#!\/bin\/bash --login\r\n#SBATCH -p multinode  # Partition is <strong>required<\/strong>. Runs on <strong>Intel Cascadelake<\/strong> hardware.\r\n#SBATCH -N <em>numnodes<\/em>   # (or --nodes=) where <em>numnodes<\/em> is between 2 and 32.\r\n#SBATCH -n <em>numcores<\/em>   # (or --ntasks=) where <em>numcores<\/em> is <em>twice<\/em> the number of nodes!\r\n                      # By using <em>twice<\/em> the number of nodes here, we get two MPI process running on each node.\r\n#SBATCH -c <em>numcores<\/em>   # (or --cpus-per-task) Number of OpenMP threads per MPI process (must be 20 to fill each socket).\r\n                      # Each MPI process will use all 20 cores on the socket it is running on.\r\n                      # This will use <em>numtasks<\/em> x <em>numcores<\/em> cores in total\r\n#SBATCH -t 4-0        # Wallclock limit (days-hours). <strong>Required!<\/strong>\r\n                      # Max permitted is 7 days (7-0) for jobs of 80-560 cores.\r\n                      #                  4 days (4-0) for jobs of 600-1280 cores.\r\n\r\n# Load any required modulefiles. A purge is used to start with a clean environment.\r\nmodule purge\r\nmodule load <em>apps\/some\/example\/1.2.3<\/em>\r\n\r\n<strong>### OpenMP ###<\/strong>\r\n# Each MPI process uses OpenMP to run on $SLURM_CPUS_PER_TASK cores (-c above)\r\nexport OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK\r\n\r\n<strong>### MPI ###<\/strong>\r\n# The MPI application is then started with $SLURM_NTASKS processes (copies of the app)\r\n# and each MPI process will use all of the cores on the node for OpenMP processing.\r\nmpirun --map-by ppr:1:<strong>socket<\/strong> mix-mode-app.exe\r\n  #\r\n  # Do <strong>NOT<\/strong> use \"--bind-to none\" in this example. Instead, mpirun will automatically\r\n  # bind an MPI process to the cores found in the <em>socket<\/em> it's running on.\r\n<\/pre>\n<pre>\r\n+-----------------------------------+   +-----------------------------------+   +-----------------------------------+\r\n| Node_1                            |   | Node_2                            |   | Node_3                            |\r\n|   Socket_0 (20-cores): MPI-proc-0 |   |   Socket_0 (20-cores): MPI-proc-2 |   |   Socket_0 (20-cores): MPI-proc-4 |\r\n|   Socket_1 (20-cores): MPI-proc-1 |   |   Socket_1 (20-cores): MPI-proc-3 |   |   Socket_1 (20-cores): MPI-proc-5 |\r\n+-----------------------------------+   +-----------------------------------+   +-----------------------------------+\r\n<\/pre>\n<p>The jobscript is:<\/p>\n<pre>\r\n#!\/bin\/bash --login\r\n#SBATCH -N 3      # 3 compute nodes\r\n#SBATCH -n 6      # 6 MPI processes in total (hence two per compute node)\r\n#SBATCH -c 20     # 20 cores per MPI process for OpenMP threads.\r\nmodule purge\r\nmodule load ......\r\nexport OMP_NUM_THREADS=$SLURM_CPUS_PER_TASK\r\nmpirun --map-by ppr:1:<strong>socket<\/strong> mix-mode-app.exe\r\n<\/pre>\n<h2>Available Hardware and Resources<\/h2>\n<p>Please see the <a href=\"\/csf3\/batch-slurm\/partitions\">Partitions<\/a> page for details on available compute resources.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Parallel batch job submission (Slurm) The CSF has a variety of compute node hardware, which is made available via different Slurm Partitions (think of these as queues for jobs wanting to access the different hardware types.) Most of the parallel applications installed on the CSF will use either the &#8220;OpenMP&#8221; parallel method or the &#8220;MPI&#8221; parallel method. This determines whether you can use more than one compute node in your job, and also the settings.. <a href=\"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/batch-slurm\/parallel-jobs-slurm\/\">Read more &raquo;<\/a><\/p>\n","protected":false},"author":2,"featured_media":0,"parent":9105,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-9233","page","type-page","status-publish","hentry"],"_links":{"self":[{"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/pages\/9233","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/comments?post=9233"}],"version-history":[{"count":20,"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/pages\/9233\/revisions"}],"predecessor-version":[{"id":12840,"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/pages\/9233\/revisions\/12840"}],"up":[{"embeddable":true,"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/pages\/9105"}],"wp:attachment":[{"href":"https:\/\/ri.itservices.manchester.ac.uk\/csf3\/wp-json\/wp\/v2\/media?parent=9233"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}