· 11 min read · Gaia Lab

What really runs on our cluster, read from the scripts themselves

Four months of the compute cluster of the ANTS group at the University of Murcia, reconstructed from the Slurm jobs themselves and anonymised: what is really computed and what engineering is built around it to coexist on a shared machine.

Rows of six-digit job identifiers in monospaced type, with a red stripe
Illustration generated for the series: the wall of identifiers.

A cluster’s job history is usually a wall of six-digit identifiers and cryptic names —mcpfw_cc, eval_dynamic_pid, boltz2— that say everything to whoever launched them and nothing to anyone else. But if you keep the scripts, the wall turns into a document: people write down why they do things, especially when a batch costs money and a night of waiting.

We have gone through four months of the compute cluster of the ANTS group, from the Department of Information and Communications Engineering of the Faculty of Computer Science at the University of Murcia —around 7,700 jobs from 19 users, reconstructed from the sbatch files themselves— and we tell the story anonymised: no names, no internal paths, looking at patterns rather than people. This is what a small research cluster looks like when you read it backwards.

7,656jobs19users2,989 hGPU hours29,598 hCPU hours
Figures for the period (19 May to 21 September 2026). GPU-hours and CPU-hours as reserved capacity.

Two meters, two stories #

If it had to be summed up in one sentence: the cluster today is mostly about language models, and mostly about testing them, not training them from scratch. Of the seven lines of work that share the machine, most touch LLMs: safety fine-tuning, code and function-calling benchmarks, automated attacks, a model for phones and inference serving. Of serious pretraining, nothing; the closest thing is LoRA fine-tunes and a great deal of evaluation. That is what a university GPU cluster looks like in 2026: not a training furnace but a measurement laboratory for models trained elsewhere.

But the counters tell two stories depending on which one you look at:

  • GPU-hours (~3,000): dominated by the benchmarking of quantized models (~37%); then proteins (~15%), self-alignment, forests and attacks (10 to 12% each) and the model for phones (~9%). The GPUs belong to the LLMs, but above all to measuring them.
  • CPU-hours (~30,000): split differently: LiDAR forests (~27%) and benchmarking (~25%) in front, followed by the mobile classifier (~19%) and the attacks (~12%). Whoever reserves the most GPU is not whoever asks for the most CPU.
GPU hours reserved by line of workGGUF benchmarkingGGUF benchmarking: 1,104 h · 37 %1,104 h · 37 %Proteins (Boltz-2)Proteins (Boltz-2): 442 h · 15 %442 h · 15 %LLM self-alignmentLLM self-alignment: 343 h · 11 %343 h · 11 %LiDAR forestsLiDAR forests: 320 h · 11 %320 h · 11 %Multi-agent jailbreakMulti-agent jailbreak: 299 h · 10 %299 h · 10 %Mobile classifierMobile classifier: 257 h · 9 %257 h · 9 %Environmental GNNs / AutoMLEnvironmental GNNs / AutoML: 162 h · 5 %162 h · 5 %OperationsOperations: 22 h · 1 %22 h · 1 %Dental imagingDental imaging: 21 h · 1 %21 h · 1 %GenomicsGenomics: 14 h · 0 %14 h · 0 %Shared inference and operationsShared inference and operations: 5 h · 0 %5 h · 0 %
GPU-hours reserved per line of work.
CPU hours reserved by line of workLiDAR forestsLiDAR forests: 7,867 h · 27 %7,867 h · 27 %GGUF benchmarkingGGUF benchmarking: 7,328 h · 25 %7,328 h · 25 %Mobile classifierMobile classifier: 5,621 h · 19 %5,621 h · 19 %Multi-agent jailbreakMulti-agent jailbreak: 3,409 h · 12 %3,409 h · 12 %Environmental GNNs / AutoMLEnvironmental GNNs / AutoML: 2,198 h · 7 %2,198 h · 7 %LLM self-alignmentLLM self-alignment: 1,421 h · 5 %1,421 h · 5 %OperationsOperations: 496 h · 2 %496 h · 2 %Shared inference and operationsShared inference and operations: 443 h · 1 %443 h · 1 %Proteins (Boltz-2)Proteins (Boltz-2): 442 h · 1 %442 h · 1 %Dental imagingDental imaging: 259 h · 1 %259 h · 1 %GenomicsGenomics: 114 h · 0 %114 h · 0 %
CPU-hours reserved per line of work.

The GPUs: defending and attacking the same model #

The GPU side is almost a self-contained debate on LLM safety, with different users in different roles.

The defence. An SFT + LoRA study on “self-alignment”, evaluating each checkpoint with HarmBench, the standard benchmark of harmful behaviours. All wired up as a dependency graph: train, and on finishing fire off an array of evaluations with --dependency=afterok, saving the outputs separately so they can be re-judged later without retraining.

The attack. Much smaller, but striking: an automated, multi-agent jailbreak framework that attacks a 120B open model across dozens of targets and turns, with the official HarmBench classifier as judge. The whole set —target, agents and judge— is stacked onto a single 141 GB H200, because each GPU node has exactly one card.

The measurement. Serving quantized GGUF models with llama.cpp and scoring them on BFCL (function calling), BigCodeBench and RULER (long context). Client and server on the same node, after discarding SSH tunnels because they kept failing.

So one user fine-tunes models to make them safer and scores them with HarmBench; another automatically attacks a large model and scores the attacks with the same HarmBench classifier. The two halves of the same question, on the same nodes.

The long tail: proteins, trees, teeth and lagoons #

The remaining lines are smaller, but each one is a clean, single-topic tenant:

  • Protein folding: Boltz-2 , a diffusion model from the AlphaFold family, over ~2,500 protein pairs. Only a handful of jobs, but huge and resumable: a five-day reservation that skips the pairs already done.
  • Forests from point clouds: above-ground biomass estimation from LiDAR with Pointcept (PTv3, OACNN), segmenting individual trees.
  • Medical imaging: segmentation with UNet++, DeepLabV3+ and SegFormer, including a 100-seed array to measure run-to-run variance.
  • A model for phones: a risk classifier for minors’ conversations, a small LLM exported to LiteRT-LM int4/int8 for phones, plus a “main” model on an H100.
  • And, around the edges, spatiotemporal GNNs predicting chlorophyll in a lagoon from satellite imagery, an AutoML for intrusion detection pilot, and single-cell genomics.
Jobs submitted per week and lineGGUF benchmarkingLLM self-alignmentMobile classifierDental imagingOperationsLiDAR forestsOther05001,0001,500Week of 2026-05-18 · LLM self-alignment: 56Week of 2026-05-18 · Dental imaging: 1Week of 2026-05-18 · Operations: 5Week of 2026-05-18 · LiDAR forests: 3Week of 2026-05-18 · Other: 1118 MayWeek of 2026-05-25 · LLM self-alignment: 32Week of 2026-05-25 · Dental imaging: 2Week of 2026-05-25 · Other: 18Week of 2026-06-01 · LLM self-alignment: 20Week of 2026-06-01 · Mobile classifier: 35Week of 2026-06-01 · Operations: 7Week of 2026-06-01 · Other: 11Week of 2026-06-08 · GGUF benchmarking: 12Week of 2026-06-08 · LLM self-alignment: 73Week of 2026-06-08 · Mobile classifier: 31Week of 2026-06-08 · Dental imaging: 6Week of 2026-06-08 · Operations: 5Week of 2026-06-08 · Other: 13Week of 2026-06-15 · GGUF benchmarking: 2Week of 2026-06-15 · LLM self-alignment: 983Week of 2026-06-15 · Mobile classifier: 46Week of 2026-06-15 · Operations: 34Week of 2026-06-15 · LiDAR forests: 4015 JunWeek of 2026-06-22 · LLM self-alignment: 1,407Week of 2026-06-22 · Dental imaging: 6Week of 2026-06-22 · LiDAR forests: 78Week of 2026-06-29 · LLM self-alignment: 71Week of 2026-06-29 · Mobile classifier: 33Week of 2026-06-29 · LiDAR forests: 8Week of 2026-07-06 · GGUF benchmarking: 7Week of 2026-07-06 · Mobile classifier: 17Week of 2026-07-06 · Dental imaging: 23Week of 2026-07-06 · Operations: 21Week of 2026-07-06 · LiDAR forests: 14Week of 2026-07-13 · Mobile classifier: 8Week of 2026-07-13 · Dental imaging: 6Week of 2026-07-13 · Operations: 5Week of 2026-07-13 · LiDAR forests: 6Week of 2026-07-13 · Other: 913 JulWeek of 2026-07-20 · LiDAR forests: 26Week of 2026-07-20 · Other: 1Week of 2026-07-27 · Mobile classifier: 29Week of 2026-07-27 · Operations: 4Week of 2026-07-27 · LiDAR forests: 1Week of 2026-07-27 · Other: 2Week of 2026-08-03 · GGUF benchmarking: 32Week of 2026-08-03 · Dental imaging: 100Week of 2026-08-03 · Other: 1Week of 2026-08-10 · GGUF benchmarking: 58Week of 2026-08-10 · Other: 2110 AugWeek of 2026-08-17 · GGUF benchmarking: 662Week of 2026-08-17 · Operations: 21Week of 2026-08-17 · Other: 8Week of 2026-08-24 · GGUF benchmarking: 1,213Week of 2026-08-24 · Other: 39Week of 2026-08-31 · GGUF benchmarking: 97Week of 2026-09-07 · GGUF benchmarking: 471Week of 2026-09-07 · Mobile classifier: 12Week of 2026-09-07 · Dental imaging: 1Week of 2026-09-07 · Operations: 5Week of 2026-09-07 · Other: 57 SepWeek of 2026-09-14 · GGUF benchmarking: 1,361Week of 2026-09-14 · Mobile classifier: 35Week of 2026-09-14 · Dental imaging: 76Week of 2026-09-14 · Operations: 67Week of 2026-09-14 · Other: 41Week of 2026-09-21 · GGUF benchmarking: 26Week of 2026-09-21 · Mobile classifier: 7Week of 2026-09-21 · Operations: 44Week of 2026-09-21 · Other: 25week (Monday), 2026
Jobs submitted per week and line: each line has its season.
GPU hours by card (whole cluster)L40S: 1,250 (41.8 %)L40S · 1,250 (42 %)H200 NVL: 731 (24.5 %)H200 NVL · 731 (24 %)L4: 622 (20.8 %)L4 · 622 (21 %)H100 NVL: 386 (12.9 %)H100 NVL · 386 (13 %)Thor (ARM): 0 (0.0 %)Thor (ARM) · 0 (0 %)2,989 hGPU reserved
GPU-hours per card across the whole cluster. One GPU per node: the partition picks the card.

What the scripts teach about sharing a machine #

Beyond what gets computed, the scripts are a quiet guide to being a good tenant. Several patterns repeat across users who have nothing to do with each other:

  • Ask for less, start sooner. A comment dated the day someone discovered it: lowering the request from 4 CPUs / 8 GB to 1 CPU / 2 GB took a job from “wait for days” to “start in seconds”, because the small one fitted into a backfill gap on a node with 28 of 30 cores busy.
  • Stage on local scratch and sync home at the end, inside a cleanup trap, so as not to hammer the NFS during the run and so that a dead job still collects its results.
  • No preemption: people use a high --nice and honest walltimes. As one script sums it up, what really frees a GPU for your neighbour is not priority, it is finishing early.
  • Notifications via ntfy and email: you enqueue and forget.

And my favourite, a security decision hidden in a benchmarking script. The team measuring code models needs the model to write code; their scripts are blunt: the cluster only generates text, it never executes the code the model writes. Grading —which does run that code— happens afterwards, in an isolated container on a local machine. One comment even corrects an earlier comment that had accepted the unsafe route, warning that misreading it “would have put the isolation in the wrong place”. Thinking clearly about where untrusted code is allowed to run, on a shared machine, is exactly the right instinct.

Conclusion #

A cluster’s job history, read calmly, is a portrait of what a community cares about. That of the ANTS group, in mid-2026, cares about language models —making them safer, breaking them on purpose, measuring what they can do and shrinking them down to a phone—, and proteins, forests and teeth around the edges.

None of that shows in the job names. It is in the scripts. Keep the scripts.

The series, line by line #

  1. Self-alignment of a 3B LLM with LoRA and HarmBench : 2,600 jobs to fine-tune and measure safety.
  2. A council of small models against a 120B one : automated jailbreaking on a single H200.
  3. Measuring quantized LLMs : 300 GGUFs, three benchmarks and one security rule. Part of the results is published at Yardstick .
  4. 4,400 protein complexes with Boltz-2 : fifteen jobs and a factor of ten.
  5. Trees one by one from LiDAR point clouds : three architectures, ten seeds and two nodes.
  6. Dental segmentation and 100 seeds : measuring variance before comparing.
  7. A risk classifier for minors that fits in a phone : synthetic data, LoRA and LiteRT-LM. Code on GitHub .
  8. GNNs for a lagoon’s chlorophyll , and an AutoML pilot.
  9. Serving models and looking after the cluster : shared inference, burn-in and probes.

Figures reconstructed from Slurm accounting and the recovered sbatch files; GPU/CPU hours reflect reserved capacity (elapsed × allocated), not measured usage, and the intent of each job is inferred from the script bodies and their comments. One orphan record that appeared as RUNNING since May and inflated the GPU total by some 2,500 hours is excluded, and for five jobs that hung some 18 hours past their limit (Slurm closed them all at once the next day) the reservation is counted up to the requested limit, not to the close; the figures in the first version of this post included both. At the request of one of the group’s members, their line of work and all their jobs have been removed from this post, from the figures and from the charts. Anonymised post: no identifiable users, paths, emails or project names.