System Monitoring¶
SystemMonitor (include/pulsatrix/system_monitor.hpp)
records CPU and GPU utilization, memory and temperatures while your code runs. It samples on a
background thread and writes each sample to a log file as soon as it's taken, so you can follow
a training run live with tail -f.
Usage¶
#include "pulsatrix/system_monitor.hpp"
pulsatrix::SystemMonitor::Options opts;
opts.interval = std::chrono::milliseconds(1000);
opts.log_path = "run.jsonl"; // CSV: opts.format = pulsatrix::LogFormat::Csv
pulsatrix::SystemMonitor monitor(opts);
monitor.start();
monitor.mark("epoch 1"); // labels the samples that follow and logs a "mark" record
// ... train ...
auto last = monitor.latest(); // std::optional<SystemSample>
auto all = monitor.history(); // bounded ring buffer (Options::history_capacity)
monitor.stop(); // returns within milliseconds
Other methods:
sample_now(): take one reading synchronously, without starting the thread.on_sample(callback): run a callback for every sample. It runs on the monitor thread; if it throws, the exception is caught and reported once.capabilities(): which metrics are available on this machine, where each one comes from, or why it's missing.
From Python, samples come back as dicts:
from pulsatrix_py import SystemMonitor
with SystemMonitor(interval_ms=1000, log_path="run.jsonl", format="jsonl") as m:
... # train
To see it run, build and run the system_monitor_demo example
(examples/system_monitor_demo.cpp).
Log format¶
JSON Lines, one object per line, with ISO-8601 UTC timestamps to the millisecond:
{"event":"mark","timestamp":"2026-10-02T18:52:48.529Z","elapsed_seconds":0.000,"label":"idle"}
{"event":"sample","timestamp":"2026-10-02T18:52:49.530Z","elapsed_seconds":1.001,"label":"idle","cpu_utilization_percent":4.698,"process_cpu_percent":0,"memory_used_bytes":11615903744,"memory_total_bytes":32723628032,"process_rss_bytes":199770112,"cpu_temperature_c":86.625,"gpus":[{"index":0,"name":"AMD GPU 0x1586 (card1)","vendor":"AMD","utilization_percent":0,"memory_used_bytes":52692738048,"memory_total_bytes":103079215104,"memory_gtt_used_bytes":94412800,"temperature_c":42,"power_watts":34.032}]}
CSV is also available (pulsatrix::LogFormat::Csv in C++, format="csv" in Python).
Missing metrics are never reported as zero¶
If the platform can't read a metric, it's reported as missing: std::nullopt in C++, null in
JSON, an empty cell in CSV and None in Python. It's never 0 or an estimate.
capabilities() tells you why, for example
"Windows: CPU temperature not available without WMI/admin".
A few things to know when reading the numbers:
- CPU percentages are computed from the change in counters between samples.
start()takes a baseline, so the first background sample arrives one interval later. process_cpu_percentis relative to one logical CPU, so it can go above 100 on a multi-core machine.- On an AMD APU, the amdgpu power reading covers the whole chip package, not just the GPU.
Where each metric comes from¶
| Metric | Linux | Windows | Other OSes |
|---|---|---|---|
| System CPU % | /proc/stat |
GetSystemTimes |
-- |
| Process CPU %, RSS | /proc/self/stat, /proc/self/status |
GetProcessTimes, GetProcessMemoryInfo |
-- |
| Memory used/total | /proc/meminfo (MemTotal - MemAvailable) |
GlobalMemoryStatusEx |
-- |
| CPU temperature | hwmon k10temp/zenpower/coretemp/cpu_thermal (Tctl/Package preferred), acpitz last |
-- (needs WMI/admin) | -- |
| AMD GPU util/VRAM/GTT/temp/power | amdgpu sysfs (/sys/class/drm/card<N>/device) |
-- | -- |
| NVIDIA GPU util/memory/temp/power | NVML, dlopen("libnvidia-ml.so.1") |
NVML, LoadLibrary("nvml.dll") |
-- |
Everything is detected at runtime. You don't need the NVIDIA or AMD SDK to build. If a driver is
missing, you simply get no GPU entries, and capabilities() says why.