TIL hwloc
I’ve seen hwloc mentioned here and there whenever I deal with building and using MPI. You can
build MPI with its own internal bundled version of hwloc (among other libraries that it needs) or it
can make use of a version of hwloc that is already installed on your system. But what even is it?
Why does MPI need it?
hwloc is a library that can gather info about your computers hardware - memory, CPUs, cores, sockets, caches, all of it, and make it visible to your program in a hardware and OS agnostic way. Every OS is going to have a different way of providing the information about the hardware it is running on, but it would be nice to have some portable code that can identify that information (and takes care of the system specific details of gathering that information for you). Without hwloc, if you had code that needed to identify the specifics of the hardware it was running on, and you needed that code to run on different OSes, you’d have to write code specific to each OS to do that identification. hwloc takes that headache away from you. It also provides a way for you to bind your process to a specific hardware heirarchy level (e.g. you could bind your process to a specific socket so that your process uses one of the cores in that socket, or you could get granular and bind at the core level so that your process uses a spefic core). You can see how something like this is useful for MPI, which would need to know things like how many cores are available on a particular node so it can know how many processes it can start on a specific node when the user asks to start N number of processes for their MPI application. This is also useful to keep a process bound to a specific core to maintain cache locality because the process doesn’t get migrated by the OS, or if MPI processes need to be bound to locations close together because of their communication patterns.
hwloc provides a heirarchical view. The lstopo command comes bundled with installing hwloc and the
default output shows a tree like view of the CPU at the top level, the L3 cache as its child, and
the L2 cache levels as the children of the L3 cache, and then its L1 caches and cores under the L2
cache. it also shows the PCIe connections like the GPU. For example, here is what my lstopo output
looks like:
Machine (15GB total)
Package L#0
NUMANode L#0 (P#0 15GB)
L3 L#0 (32MB)
L2 L#0 (1024KB) + L1d L#0 (32KB) + L1i L#0 (32KB) + Core L#0
PU L#0 (P#0)
PU L#1 (P#8)
L2 L#1 (1024KB) + L1d L#1 (32KB) + L1i L#1 (32KB) + Core L#1
PU L#2 (P#1)
PU L#3 (P#9)
L2 L#2 (1024KB) + L1d L#2 (32KB) + L1i L#2 (32KB) + Core L#2
PU L#4 (P#2)
PU L#5 (P#10)
L2 L#3 (1024KB) + L1d L#3 (32KB) + L1i L#3 (32KB) + Core L#3
PU L#6 (P#3)
PU L#7 (P#11)
L2 L#4 (1024KB) + L1d L#4 (32KB) + L1i L#4 (32KB) + Core L#4
PU L#8 (P#4)
PU L#9 (P#12)
L2 L#5 (1024KB) + L1d L#5 (32KB) + L1i L#5 (32KB) + Core L#5
PU L#10 (P#5)
PU L#11 (P#13)
L2 L#6 (1024KB) + L1d L#6 (32KB) + L1i L#6 (32KB) + Core L#6
PU L#12 (P#6)
PU L#13 (P#14)
L2 L#7 (1024KB) + L1d L#7 (32KB) + L1i L#7 (32KB) + Core L#7
PU L#14 (P#7)
PU L#15 (P#15)
HostBridge
PCIBridge
PCI 01:00.0 (VGA)
CoProc(OpenCL) "opencl0d0"
GPU(Display) ":0.0"
PCIBridge
PCI 02:00.0 (NVMExp)
Block(Disk) "nvme0n1"
PCIBridge
PCIBridge
PCIBridge
PCI 05:00.0 (NVMExp)
Block(Disk) "nvme1n1"
PCIBridge
PCI 08:00.0 (Ethernet)
Net "eno1"
PCIBridge
PCI 09:00.0 (Network)
Net "wlp9s0"
PCIBridge
PCI 0b:00.0 (SATA)
PCIBridge
PCI 0c:00.0 (SATA)
PCIBridge
PCI 0d:00.0 (SATA)
Block(Disk) "sda"
PCIBridge
PCI 0e:00.0 (VGA)lstopo can also output in other formats, including more visual ones like ASCII art or SVG.
That’s a neat tool, but the real power of hwloc is the library that you can use to programmatically navigate these hardware levels and bind your processes to them. For example, if you wanted to bind your process to a specific core, say core 2, you could do that with the following code:
#include "hwloc.h"
int ackermann(int m, int n) {
if (m == 0) {
return n + 1;
} else if (m > 0 && n == 0) {
return ackermann(m - 1, 1);
} else {
return ackermann(m - 1, ackermann(m, n - 1));
}
}
int main() {
hwloc_topology_t topology;
hwloc_obj_t obj;
hwloc_cpuset_t cpuset;
hwloc_topology_init(&topology);
hwloc_topology_load(topology);
// getting total depth in the hwloc heirarchical tree
int topodepth = hwloc_topology_get_depth(topology);
// getting the depth in the tree at which the cores are
int depth = hwloc_get_type_depth(topology, HWLOC_OBJ_CORE);
// get core #2
obj = hwloc_get_obj_by_depth(topology, depth, 2);
if (obj) {
// make copy so we can modify it
cpuset = hwloc_bitmap_dup(obj->cpuset);
hwloc_bitmap_singlify(cpuset);
// bind current thread to specified cpu
if (hwloc_set_cpubind(topology, cpuset, 0)) {
char *str;
int error = errno;
hwloc_bitmap_asprintf(&str, obj->cpuset);
printf("Couldn't bind to cpuset %s: %s\n", str, strerror(error));
free(str);
} else {
int a = ackermann(4, 4);
printf("Ackermann result: %d\n", a);
}
}
// unset binding, or rather bind the thread back to the whole cpu
obj = hwloc_get_obj_by_depth(topology, 0, 0);
if (hwloc_set_cpubind(topology, obj->cpuset, 0)) {
char *str;
int error = errno;
hwloc_bitmap_asprintf(&str, obj->cpuset);
printf("Couldn't bind to cpuset %s: %s\n", str, strerror(error));
free(str);
}
}The above code gets the ‘depth’ at which the cores reside (as in the depth in the heirarchical tree that hwloc arranges the hardware components) and picks the core 2. It then binds the current thread to that core before running the Ackermann fucnction.
You can compile and run the above code with
gcc -o hello hello.c -lhwloc
./helloIf you run the above code and in a separate terminal run htop, you will see the usage of core 2
shoot up as it runs the Ackermann function (you should Ctrl+C out of it since that function will
take a while).
This of course can get a lot more useful when you are working with powerful CPUs with
dozens of cores each, as is the case on HPC clusters. Being able to bind to different NUMA nodes or
L3 cache levels (if your CPU has multiple) has a lot of uses in spreading out your work across a CPU
and keeping them there for cache locality. So you can see why this is useful for MPI and for HPC job
schedulers scheduling jobs for different users on a shared node. You can see a more in depth ‘hello
world’ style example showing the capabilities of hwloc here.
I mentioned that hwloc abstracts away the OS specific methods to gather the hardware information and
presents them with a common API to navigate them and bind to them. On Linux, for example, the
hardware information lives in the sysfs filesystem mounted on /sys. Information about the CPU as a
whole lives in /sys/devices/system/node/ and information about individual cores live in
/sys/devices/system/cpu/cpuN where you replace N with your core number. hwloc gathers that info
for you and presents an easier API to access that info. And if you want to bind to a specific core,
on Linux this would be the sched_setaffinity syscall. Here’s an example using the glibc function for the syscall to
bind your current thread to core 2.
#define _GNU_SOURCE
#include <stdio.h>
#include <sched.h>
#include <string.h>
#include <errno.h>
int ack(int m, int n)
{
if (m == 0){
return n+1;
}
else if((m > 0) && (n == 0)){
return ack(m-1, 1);
}
else if((m > 0) && (n > 0)){
return ack(m-1, ack(m, n-1));
}
}
int main(){
int A;
// setting thread affinity with sched_setaffinity
cpu_set_t mask;
CPU_ZERO(&mask);
// set cpu 2 as the core to pin
CPU_SET(2, &mask);
if(sched_setaffinity(0, sizeof(mask), &mask) != 0) {
fprintf(stderr, "pin to core 2 failed: %s\n", strerror(errno) );
return -1;
}
A = ack(4, 4);
printf("%d", A);
return 0;
}But again, this would be different on a different OS, as would gathering the hardware info. hwloc
saves you that trouble of writing OS specific code! And provides other quality of life features when
you’re having to deal with hardware info. So overall, choose hwloc if need to do stuff with
collecting and using hardware info and don’t want to deal with the OS and hardware nitty gritties.