Overview¶
Partitions¶
intel partition¶
This is the default partition selected if the -p/--partition argument is not provided to srun or sbatch, and is currently where the majority of our CPU-only nodes reside. There are no nodes in the intel partition with GPUs.
The CPUs in this partition vary. Some are E5 2660's, some are E5 2660v2's, and some are E5 2670v2's. All of the E5 2660v2 and E5 2670v2 nodes are connected with a low-latency high-throughput FDR InfiniBand fabric. The openmpi modules on the cluster are compiled to communicate over the InfiniBand fabric when available through UCX.
Some of these CPUs are older, but workable throughput can still be achieved with them through the use of Job Arrays or openmpi to achieve some degree of parallelism.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| intel | 25 | ACF R1 | Intel(R) Xeon(R) CPU E5-2670 v2 @ 2.50GHz | 2 | 20 | 125GiB | none/0 | 1 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
| intel | 20 | ACF R29 | Intel(R) Xeon(R) CPU E5-2670 v2 @ 2.50GHz | 2 | 20 | 125GiB | none/0 | 1 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
| intel | 16 | ACF R4 | Intel(R) Xeon(R) CPU E5-2660 v2 @ 2.20GHz | 2 | 20 | 125GiB | none/0 | QDR Infiniband (40Gb/s), 10 Gigabit Ethernet |
| intel | 11 | ACF R3 | Intel(R) Xeon(R) CPU E5-2660 0 @ 2.20GHz | 2 | 16 | 125GiB | none/0 | 1 Gigabit Ethernet |
| intel | 6 | Upstairs R1 | Intel(R) Xeon(R) CPU E5-2660 0 @ 2.20GHz | 2 | 16 | 62GiB | none/0 | 10 Gigabit Ethernet |
| intel | 6 | Upstairs R1 | Intel(R) Xeon(R) CPU E5-2660 0 @ 2.20GHz | 2 | 16 | 31GiB | none/0 | 10 Gigabit Ethernet |
| intel | 4 | ACF R3 | Intel(R) Xeon(R) CPU E5-2660 0 @ 2.20GHz | 2 | 16 | 251GiB | none/0 | 1 Gigabit Ethernet |
| intel | 4 | ACF R29 | Intel(R) Xeon(R) CPU E5-2670 v2 @ 2.50GHz | 2 | 20 | 251GiB | none/0 | 1 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
| intel | 3 | ACF R1 | Intel(R) Xeon(R) CPU E5-2670 v2 @ 2.50GHz | 2 | 20 | 251GiB | none/0 | 1 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
| intel | 2 | Upstairs R1 | Intel(R) Xeon(R) CPU E5-2660 0 @ 2.20GHz | 2 | 16 | 251GiB | none/0 | 10 Gigabit Ethernet |
| intel | 1 | Upstairs R1 | Intel(R) Xeon(R) CPU E5-2660 v2 @ 2.20GHz | 2 | 20 | 125GiB | none/0 | 10 Gigabit Ethernet |
| intel | 3 | ACF R25 | Intel(R) Xeon(R) Platinum 8268 CPU @ 2.90GHz | 2 | 48 | 251GiB | none/0 | 10 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
gpu-a100 partition¶
Visit Slurm GPU Jobs for more information on how to use GPU nodes. If using OnDemand, visit OnDemand Desktop first.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| gpu | 1 | ACF R26 | Intel(R) Xeon(R) Platinum 8260 CPU @ 2.40GHz | 4 | 96 | 1510GiB | NVIDIA A100-PCIE-40GB/2 | 10 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
gpu-p100 partition¶
Visit Slurm GPU Jobs for more information on how to use GPU nodes. If using OnDemand, visit OnDemand Desktop first.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| gpu | 2 | ACF R28 | Intel(R) Xeon(R) CPU E5-2650 v4 @ 2.20GHz | 2 | 24 | 125GiB | Tesla P100-PCIE-16GB/4 | 1 Gigabit Ethernet |
gpu-v100 partition¶
Visit Slurm GPU Jobs for more information on how to use GPU nodes. If using OnDemand, visit OnDemand Desktop first.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| gpu | 1 | ACF R28 | Intel(R) Xeon(R) CPU E5-2620 v4 @ 2.10GHz | 2 | 16 | 125GiB | Tesla V100S-PCIE-32GB/2 | 10 Gigabit Ethernet |
gpu-l40s partition¶
Visit Slurm GPU Jobs for more information on how to use GPU nodes. If using OnDemand, visit OnDemand Desktop first.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| gpu | 1 | ACF R25 | INTEL(R) XEON(R) GOLD 6526Y | 2 | 32 | 251GiB | NVIDIA L40S/2 | 10 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
gpu-pro6000 partition¶
Visit Slurm GPU Jobs for more information on how to use GPU nodes. If using OnDemand, visit OnDemand Desktop first.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| gpu | 1 | ACF R28 | AMD EPYC 9135 | 2 | 16 | 503GiB | NVIDIA RTX Pro 6000/2 | 10 Gigabit Ethernet |
gpu-small partition¶
Visit Slurm GPU Jobs for more information on how to use GPU nodes. If using OnDemand, visit OnDemand Desktop first.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| gpu | 2 | ACF R28 | Intel(R) Core(TM) i7-6850K CPU @ 3.60GHz | 1 | 6 | 62GiB | NVIDIA TITAN Xp/2 | 1 Gigabit Ethernet |
| gpu | 1 | ACF R28 | Intel(R) Core(TM) i7-6850K CPU @ 3.60GHz | 1 | 6 | 94GiB | NVIDIA TITAN Xp/1, NVIDIA TITAN RTX/1 | 1 Gigabit Ethernet |
bigm partition¶
This partition contains nodes with 512GB of memory or more.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| bigm | 4 | ACF R1 | Intel(R) Xeon(R) CPU E5-2670 v2 @ 2.50GHz | 2 | 20 | 503GiB | none/0 | 1 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
neoverse partition¶
This partition contains two Ampere Altra servers with 1 80-core Neoverse N1 CPU each.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| ampere | 2 | ACF R26 | Neoverse-N1 | 1 | 80 | 125GiB | none/0 | 1 Gigabit Ethernet |
mmicc partition¶
This partition contains two nodes with two NVIDIA L40S GPUs, and is exclusive to researchers associated with MMICC.
| Partition | Nodes | Location | CPU | Sockets | Cores | RAM | GPUs/Count | Connectivity |
|---|---|---|---|---|---|---|---|---|
| mmicc | 2 | ACF R25 | INTEL(R) XEON(R) GOLD 6526Y | 2 | 32 | 251GiB | NVIDIA L40S/2 | 10 Gigabit Ethernet, FDR InfiniBand (56Gb/s) |
cpu-debug partition¶
Debug partitions have short walltime limits but have the highest priority tier relative to other partitions to facilitate rapid debugging/testing. This partition includes all the CPU nodes in the cluster and allows users to submit short, high priority jobs to them. Each user can only have one running or pending job in this partition at a time.
You can check current walltime limits and other information with scontrol show partitions or scontrol show partition <specified partition>.
gpu-debug partition¶
Visit Slurm GPU Jobs for more information on how to use GPU nodes. If using OnDemand, visit OnDemand Desktop first.
Debug partitions have short walltime limits but have the highest priority tier relative to other partitions to facilitate rapid debugging/testing. This partition includes all the GPU nodes in the cluster and allows users to submit short, high priority jobs to them. Each user can only have one running or pending job in this partition at a time.
You can check current walltime limits and other information with scontrol show partitions or scontrol show partition <specified partition>.
scavenger partition¶
This partition includes every node in the cluster, meaning it overlaps with all other partitions, including hidden ones.
Jobs submitted to this partition have almost no restrictions, but are preemptible by jobs submitted to other partitions.
"Preemptible" means that jobs submitted to this partition can be stopped and requeued at any point to make way for higher priority jobs. Jobs submitted to this partition should checkpoint periodically (and be able to resume from checkpoint) to minimize lost progress.
You can use this partition to "scavenge" unused cycles/time on hardware you otherwise would not have access to. Since walltime limits are less strict, you can also use this partition for long-running jobs so long as your jobs can handle being stopped/requeued without losing all their progress.
This is a good place to submit long-running jobs that are capable of checkpoints/resumes.
Using the nodes¶
Visit Slurm or OnDemand Desktop to learn how to submit a job.